Mixing Sean Goedecke’s take with the full-contact roast the Hacker News wizards gave it (957 points, 390 comments), you start to see that being a software engineer means trying to please everyone at once.

Here’s a savage-but-affectionate recap of the internet’s opinions on system design — full credit to Sean for the original ideas, the roasting is all Hacker News, and the part at the end where I get petty is entirely my own doing.

🧘 1. The “Chill” Doctrine (from the author)

  • The more boring, the better it’s working: Good system design is a system running so smoothly nobody even notices it exists. Your reward for doing this well is that nobody will ever congratulate you for it. Starting a project with a mountain of complex architecture is a recipe for disaster — a complex system that actually works almost always grew out of a simple one that already worked, it never starts complex.
  • Worship at the altar of “stateless”: Avoid stateful components like they’re cursed, because a stateless service that crashes just gets automatically restarted and forgives itself, while a stateful one with one bad row in the table needs an actual human to show up and fix it. Practical translation: pick exactly one service that’s allowed to write to a given table, and make everyone else send it an API request like a civilized adult. Reads are allowed to be a little more relaxed about this — sometimes a quick direct read beats a slower round-trip through an internal service.
  • Schemas: flexible, but not too flexible: Design your table so a human can skim it and roughly guess what the app is storing and why. Yes, dumping everything into one giant JSON “value” column is more flexible — it’s also how you smuggle a truckload of complexity straight into your application code and regret it later.
  • Indexes: aim, don’t spray: Put indexes where your actual queries are, put the highest-cardinality column first (nobody wants an index that still has to scan every user before finding the one with the right email), and don’t index literally everything you can think of — each index you add is a tax on every future write.
  • Let the database do the database’s job: If you need data from two tables, JOIN them — don’t fetch both and staple them together in application code like a caveman. And if you’re using an ORM, watch out for it quietly firing off a query inside a loop, which is the classic, silent way to turn one clean query into a hundred embarrassing ones.
  • Reads go to the replica, not the poor overworked writer node: Send as much read traffic as you can to read replicas and leave the primary node alone to do its one job: writing. The exception is when you truly cannot tolerate even a few milliseconds of replication lag — in which case, just keep the freshly-written value in memory instead of immediately reading it back like some kind of trust issue.
  • Watch out for query spikes, especially write ones: A database that’s getting overloaded gets slower, which makes it more overloaded, which is a delightful death spiral. If you’re building something that could cause a burst of writes (bulk import, we’re looking at you), throttle it before it throttles you.
  • Split fast work from slow work: If a user is waiting on it, it needs to answer in a few hundred milliseconds — everything else belongs in a background job. This is such a universal pattern that literally every company reinvents the exact same “queue + worker” wheel and nobody’s mad about it. Bonus tip from the trenches: if you need a job to run a month from now, don’t put it on your Redis queue — Redis wasn’t built to babysit something for that long. Stick it in a plain database table with a scheduled_at column instead, like the adults who came before Redis existed.
  • Cache like a senior, not like a junior: Juniors discover caching and want to cache everything; seniors want to cache almost nothing. That’s because a cache is just another stateful thing that can go stale, get out of sync, or serve confidently wrong answers. Never cache a slow thing before you’ve actually tried to make it fast — caching an unindexed query is not an optimization, it’s a confession.
  • Events are a queue with main character energy: An event says “this happened,” instead of a job saying “please go do this.” Great when the sender genuinely doesn’t care what happens next, or when the volume is too high and too unimportant to justify a direct call. Bad as a default — most of the time, one service just calling another with a plain API request is easier to trace, log, and reason about at 2 a.m.
  • Push vs. pull is really a “how many people are asking” question: Pulling (classic website, refresh to see new stuff) is simple and fine at small scale. Pushing (Gmail-style, new mail just appears) scales better once you’ve got a lot of listeners who all want the same thing at the same time. At truly enormous scale, both options basically turn into “throw more infrastructure at it” — an event-fanout army if you’re pushing, or a battalion of read-replica caches if you’re pulling.
  • Hot paths deserve outsized paranoia: A handful of code paths carry way more risk than everything else combined — the stuff that touches every single user action, or decides whether someone gets billed. A busted settings page annoys one person. A busted hot path takes the whole product down with it.
  • Log like you’re already being blamed for something: Log every rejection, every weird 422, every “we didn’t charge this because of X.” It feels like ugly boilerplate right up until a big customer is furious about something and you need the receipts. And don’t just watch your averages — watch p95/p99, because the handful of painfully slow requests you’re not seeing are disproportionately your biggest, angriest customers.
  • Retries, idempotency keys, and choosing your failure mode on purpose: Retries without circuit breakers just politely DDoS a service that’s already struggling. And if a “bill this user” request times out, you genuinely don’t know if it succeeded — that’s what idempotency keys are for. Also: decide in advance whether each system fails open or fails closed. Rate limiting going down should probably let traffic through; auth going down should never, ever let anyone in.

🍿 2. Life Isn’t a Fairy Tale (the internet claps back)

  • Try that in an interview and you’re getting “next”: A lot of engineers are out here warning: walk into a system design interview with “I’d just use Postgres” energy and you will get rejected. Interviewers want to see you conjure up Kubernetes and draw a diagram that looks like a spider built it. That said, the seniors chime in too: pick the simple answer, sure, but you’d better be able to talk your way through the tradeoffs. Something like: “At this QPS I’d just reach for SQL, but once we hit a million users I’d bolt Kafka on.”
  • Resume-Driven Development Syndrome: Everyone agrees simple is good, but if you don’t cram some heavyweight FAANG-grade tech into your side project, what exactly are you supposed to put on your resume to land the big paycheck?
  • Civil war over the shared-database question: The author says write an API so services talk to each other instead of poking the same table directly. The comments section splits into two camps: one side says an API is expensive to write and just adds a “microservices tax,” the other says share a table and the day someone needs to change a column, the whole company goes down with the ship.
  • The ORM vs. raw SQL holy war: The moment the author suggests letting the database do the work (i.e., use a JOIN), the thread turns into a bloodbath over ORMs (Entity Framework, ActiveRecord, pick your poison). One camp swears ORMs quietly generate the dumb queries that take servers down; the other snaps back that it’s not the ORM’s fault, it’s developers who don’t understand what their ORM is doing — meanwhile go ahead and try writing raw SQL by hand at scale and see how that goes.

😅 3. My Two Cents (extra spice, no extra charge)

Having sat through this entire roast session, here’s my own petty little addition to the pile:

  • The big incidents I’ve seen almost always seem to happen on Kubernetes — I’ve genuinely never seen one on a system running plain ECS (AWS’s managed container service)… heh. Not that K8s is bad, it’s just that more knobs means more ways to shoot yourself in the foot at 3 a.m.
  • Honestly, a lot of the K8s-everywhere obsession isn’t really about the workload at all — it’s about wanting to look like you know your stuff. Plenty of companies (and plenty of engineers) reach for Kubernetes the same way someone buys a Ferrari to drive to the grocery store: not because they need the horsepower, but because “Kubernetes” just sounds like the correct answer. I’ve personally watched a senior engineer stand up an entire K8s cluster to host exactly two containers — a backend API and a batch job — for a workload whose I/O pattern was about as thrilling as a spreadsheet. Not underpowered, not over-scaled, just… completely ordinary. The cluster added nothing except a control plane to babysit and a very impressive-looking architecture diagram nobody asked for.
  • On the shared-database civil war, I’m firmly Team Author — yes, writing an extra API costs you time, but it’s cheap compared to the price you pay once five services are writing to one table and nobody dares touch the schema anymore for fear of breaking someone else’s code.
  • As for ORM vs. raw SQL, I’ll happily sit in the crossfire on this one: the ORM isn’t killing anyone — it’s the ORM’s own convenience that’s the trap. It makes writing a query so effortless that a lot of engineers quietly take the whole system down without ever realizing it, because they never see how many hidden queries it’s actually firing off underneath.
Kubernetes vs Serverless, illustrated

🎬 The Verdict

Being a software engineer really is a contradiction. At work, you’re lighting incense and praying for the simplest system possible so you don’t get paged at 3 a.m. But walk into an interview, and suddenly you’d better look dangerous and all-knowing again.

Go read the actual full original, it’s a much better use of your time than this remix: Everything I know about good system design.

📚 References