Engineering

Boring infrastructure is a feature

Every piece of novel technology you add is a thing you will one day debug at three in the morning, with no documentation, while eleven people ask whether it is DNS. Here is how I decide what earns a place in a production system.

I once spent a Saturday night on a video call with three engineers, a database that refused to accept writes, and a queue that had silently stopped draining forty minutes earlier. Nobody had written that queue. It had arrived, as these things do, with the best of intentions: it was faster than the thing it replaced, it had a nicer API, and the blog post announcing it had a lot of diagrams.

We fixed it at 4am by restarting a process none of us could explain. The incident review took eleven minutes, because there was exactly one honest thing to say: we had chosen a tool we could not debug.

I have thought about that night for years, and it has quietly shaped almost every technical decision I have made since. Not into conservatism for its own sake — I like new tools, I read the release notes, I have migrated more systems than is strictly healthy — but into a habit of asking what a technology costs me on my worst day rather than on my best one.

The bill for novelty arrives later

New technology is almost always evaluated at the moment of adoption, which is precisely the moment when it looks best. You are reading about it because someone solved a hard problem with it. The demo is clean. The happy path is genuinely nicer than what you have now. And the cost — the part that decides whether you sleep through the night in two years — has not been invoiced yet.

That invoice has a few line items, and they are remarkably consistent:

  • Familiarity debt. Your team's ability to reason about a system is built from thousands of small exposures. A new tool resets that clock to zero for everyone, including the person who introduced it.
  • Failure-mode debt. You know how Postgres behaves when the disk fills. You do not yet know how this year's favourite datastore behaves, and the answer is not in the documentation.
  • Ecosystem debt. Backups, metrics, tracing, migrations, the emergency path when the primary is unreachable — each is a separate integration that somebody has to write and then maintain forever.
  • Abstraction debt. The nicer API usually means a larger surface between you and the machine. That is a fine trade until the day you need to see underneath it.

None of these are arguments against change. They are arguments for pricing it properly. My rule of thumb is that a new component has to be at least twice as good as the boring alternative before it is worth introducing, because the second half of that margin is what the invoice will eat.

The systems I am proudest of are the ones that have not needed a decision in two years. That is not luck. It is a series of unglamorous choices made on purpose.

Something I wrote in an internal design doc, and have since stolen from myself repeatedly

What boring actually buys you

Boring is not the absence of ambition. It is a specific, spendable resource: the ability to hold the whole system in your head at once. When you can do that, everything else gets easier.

Debugging becomes reading instead of guessing

With a stack you know, an incident is an investigation. With a stack you do not, an incident is a séance. You restart things and watch. You search a Discord server. You find a GitHub issue from 2021 with the same stack trace and no resolution. Every minute spent there is a minute not spent on the actual problem, and the two are not equivalent: guessing produces changes nobody can justify later.

Onboarding stops being archaeology

A new engineer can read a well-worn stack's documentation and be useful in a week. They can read your custom operator's README and be useful in a month, if the person who wrote it is still around. The difference compounds every time you hire.

You can say no with evidence

This one is underrated. When you run the obvious thing, you have years of other people's production experience behind every objection you raise. "We should not do this because it will page us at 3am" is a much stronger sentence when you can point at a decade of incident reports from a thousand other companies.

Your deploy becomes a command, not a project

The deployment procedure for the system I am happiest with is nine lines of shell. It has been nine lines for three years. Nobody has ever asked for a rollback plan, because the rollback is line four.

#!/usr/bin/env bash
set -euo pipefail

git pull --ff-only
pnpm install --frozen-lockfile
pnpm build
rsync -a --delete ./dist/ server:/var/www/site/
ssh server 'systemctl reload nginx'
curl -fsS https://example.com/health | grep -q '"ok":true'
echo "deployed in $(($SECONDS))s"

There is no orchestration layer here, and that is the point. If this script fails, the failure is legible. If it half-succeeds, the person fixing it can see exactly which line stopped.

A useful framing

Ask what the on-call engineer sees at 3am: a stack trace in a language they know, a dashboard with metrics they understand, and a runbook with steps they have run before. If any of those three is missing, you have not finished adopting the technology — you have only finished installing it.

A test I run before adopting anything

I keep a short questionnaire in a note. It takes ten minutes and it has talked me out of roughly half the things I was excited about, which is the correct hit rate.

  1. Who do I call? Not a support contract — a person, or at least a community where questions get answered by someone other than the maintainer within a day.
  2. What does the failure look like? I want a specific, boring answer. "It panics and the supervisor restarts it" is a good answer. "It depends on the failure mode" is a research project.
  3. How do I get the data out? If the export path is not documented, the import path was a sales pitch.
  4. Has it been in production for three years somewhere? Somewhere specific. Version 2.0 of anything is a different product from the one the case studies describe.
  5. Would I bet a weekend on it? Because that is the actual price. If the answer is no, the answer is no.

Where novelty is worth it

I do not want to talk myself, or you, into a world where nothing changes. That world is worse, and it is also not what I practise. There are places where the new thing is genuinely a step change, and I have taken those bets.

SituationBoring choiceWhen I break the rule
Primary datastorePostgresNever, so far. Replicas and a read cache have covered every case.
Background jobsA table and a workerPast roughly a million jobs a day, where the table becomes the bottleneck.
Static hostingnginx on one boxWhen traffic is genuinely global and the CDN bill is smaller than the engineering time.
Frontend frameworkWhatever the team already knowsWhen the page's interactivity is the product, not a garnish.

The pattern is that I break the rule when the workload has changed shape, not when a better tool has appeared. A tool getting better is not a reason to migrate. A workload outgrowing its tool is.

The maintenance question

There is one more thing I ask, and it is the one that has saved me the most time: who maintains this in eighteen months?

If the answer is "the person who introduced it", and that person is me, then I have just taken on a permanent obligation in exchange for a temporary convenience. That is sometimes a good trade. It is almost never a good trade when nobody wrote it down.

So when I do adopt something new, I write the runbook first. One page: what it is, why it is here, what breaks, how to tell, how to recover, how to remove it. If I cannot fill that page in, I do not understand the system well enough to be responsible for it — and that is a much stronger signal than any benchmark.

The queue that woke us up on that Saturday night is gone now. It was replaced by a table in Postgres and a worker with a retry limit. It handles about a fifth of the throughput the original was rated for, and nobody has been paged about it since.

That is the whole argument, really. It was slower, and it was better.

Portrait of Elliot Vance

Elliot Vance

Software engineer in Rotterdam. I help teams make their systems fast, observable and boring, usually by removing more than I add. Eleven years in, mostly on web performance, databases and platform work.

More about me