Performance

A performance budget you will actually keep

Most performance budgets die inside a spreadsheet, because nobody owns the number and nothing blocks the merge. This one has been enforced on every pull request for two years at a product with real traffic. Here is what it measures, what it costs, and what broke.

The first performance budget I wrote lived in a spreadsheet with eleven tabs, including one called benchmarks nobody could explain. I updated it twice, and by the third month the numbers were stale.

The budget I am describing has run on every pull request for two years, on a product taking four million sessions a month. It has blocked sixty or seventy merges and been raised twice, both times in a diff with a reason and a name attached. It is barely a hundred lines of Node and a JSON file, with the two things the spreadsheet lacked: a named human accountable for each number, and a red build when one of them slips.

Why most budgets die

Nobody owns the number. The budget becomes a document rather than a constraint, sitting in a wiki page under "Engineering standards", which is where numbers go to be forgotten. Everyone agrees the LCP target is good. Nobody is the person who has to explain why it moved 380 milliseconds last month. Without that person it is a preference, and preferences lose to deadlines every time.

Nothing blocks the merge. This is the fatal one. A number that is reported is not a number that is enforced, and the difference shows up in about six weeks. The dashboard has a graph, the graph has a line climbing to the right, and everyone reads it the way you read a weather forecast for a city you do not live in. The regression is never one event; it is forty defensible ones, and undoing the total means reopening forty decisions.

So we built it in the other order: first the thing that can fail a build, then the person who has to look at it, and only then the number.

Pick four numbers, no more

Our first draft had eleven metrics, because eleven is what a good audit gives you. It lasted three weeks: eleven numbers cannot all be somebody's job, and a budget that is nobody's job is the spreadsheet again, only with a YAML file. Four survived.

Largest Contentful Paint at p75, mid-tier Android over 4G. This is what a user feels, and the profile is not negotiable: a laptop on office wifi reports an LCP half the real one. We throttle the lab run to mimic a Moto G Power on a congested connection, and keep the field number for the Monday conversation.

Total JavaScript transferred on the two key routes. Transferred, after compression, exactly as the CDN serves it — not parsed, and not the number the build tool prints. The two routes are the home page and the signed-in dashboard. Choosing a small set of routes we cared about mattered enormously.

Hero image bytes. The largest above-the-fold image, including the responsive variants selected at common breakpoints. Image budgets are unglamorous and that is where the wins are: our hero went from 620 KB to 180 KB in a fortnight, four times the total saving of every JavaScript change that quarter.

One server-side number on the critical path. I would take TTFB at p75, but we already instrumented query counts, so ours became the query count on the critical rendering path, capped at nine. The cap is what stopped the tenth appearing.

Why not a fifth? The fifth is what stopped the team caring. We tried adding total CSS bytes in month seven. It never failed a build on its own, and its presence taught people that most of the report was noise.

The number is somebody's job, on a rotation

Every metric has exactly one owner, and the owner is a rotating engineer: one per quarter, named in the repository, about an hour a week. A rotation rather than a permanent role, because a permanent role produces a person who cares about performance instead of a team that does.

Ownership works only if the numbers reach people where they already are. A dashboard is somewhere nobody goes on a Tuesday, so the budget lives in two unavoidable places: the pull request template, and the deploy log, which prints one line per metric after every release. "JS: 176.2 KB (+12.4)" is seen by everyone who deploys.

Metric Target How it is measured Owner
LCP p75, mid-tier Android over 4G ≤ 2.5 s Lighthouse CI, throttled, median of five runs per key route Rotating performance owner
JS transferred, home page ≤ 170 KB Brotli bytes served for the route's chunks plus shared vendor code Feature team owning the route
JS transferred, dashboard ≤ 260 KB Same script, same compression, budget attributed per route Feature team owning the route
Hero image bytes ≤ 180 KB Variants referenced by the selected srcset at 360, 768 and 1280 px Design systems engineer
Critical-path query count ≤ 9 Server timing spans on first render, p75 over 24 hours Backend owner for the route

The first number to break was a rule rather than a metric. Someone changed the home page ceiling from 170 to 175 KB in a pull request that also refactored the navigation, so the raise and the reason arrived in one diff. Uncomfortable, and correct: a spreadsheet records what you decided, a diff records who decided it.

The CI job that does the arguing

The enforcing part is small: a JSON file of thresholds, a script that walks the build output, an exit code. It runs in eleven seconds, which matters more than accuracy. A check that takes four minutes gets skipped under deadline pressure; one that takes ten seconds gets trusted.

{
  "routes": {
    "/":    { "js": 174080, "css": 40960, "lcp": 2500 },
    "/app": { "js": 266240, "css": 61440, "lcp": 3000 }
  },
  "hero": 184320,
  "exceptions": [
    {
      "route": "/app",
      "metric": "js",
      "delta": 18432,
      "owner": "@priya",
      "until": "2025-03-01",
      "reason": "campaign reporting widget, temporary"
    }
  ]
}

The script sums the bytes the route requests, adds any live exception delta, compares against the threshold, and fails with a message written for someone who has never read the budget.

#!/usr/bin/env node
// tools/check-budget.mjs — runs after `pnpm build`
import { readFile } from "node:fs/promises";
import { brotliCompressSync } from "node:zlib";

const budget = JSON.parse(await readFile("budget.json", "utf8"));
const today = new Date().toISOString().slice(0, 10);
const failures = [];

const bytesOnWire = async (file) =>
  brotliCompressSync(await readFile(file, "utf8")).length;

for (const [route, limits] of Object.entries(budget.routes)) {
  const manifest = JSON.parse(
    await readFile(`.next/routes/${encodeURIComponent(route)}.json`, "utf8"),
  );
  const js = (
    await Promise.all(manifest.chunks.map((c) => bytesOnWire(`.next/${c}`)))
  ).reduce((a, b) => a + b, 0);

  const exception = budget.exceptions.find(
    (e) => e.route === route && e.metric === "js" && e.until >= today,
  );
  const ceiling = limits.js + (exception ? exception.delta : 0);

  if (js > ceiling) {
    failures.push(
      `${route} JS is ${(js / 1024).toFixed(1)} KB, ceiling ${(ceiling / 1024).toFixed(1)} KB ` +
        `(${((js - limits.js) / 1024).toFixed(1)} KB over the standing budget). ` +
        `Largest chunk: ${manifest.chunks.at(-1)}. ` +
        `Delete, defer or downgrade — see docs/performance-budget.md#what-to-do.`,
    );
  }
}

if (failures.length) {
  console.error("\nPerformance budget exceeded:\n");
  for (const f of failures) console.error(`  x ${f}\n`);
  console.error("To request an exception, add an entry to budget.json with an");
  console.error("owner and an expiry date, then tag the performance owner.\n");
  process.exit(1);
}
console.log("Performance budget: all routes within budget.");

The exit code is the enforcing half. The message decides whether people resent the check or use it. Three things in it earn their place: the overage is stated against the standing budget, so an expiring exception cannot quietly become permanent; the largest chunk is named, which turns a minute of guesswork into a glance; and the failure points at documentation with three concrete options rather than an instruction to be faster. Our first version printed only a number and an exit code, and the first response was somebody raising the threshold in the same pull request.

One thing went wrong here. The original script measured gzipped output, because that was the size the build tool reported. The CDN serves Brotli, so our numbers came out roughly nine per cent smaller than what users downloaded, and we believed we had headroom we did not have right up to the week a real regression ate it.

Field data beats lab data, but lab data blocks merges

Field data — real users, real devices, real networks — is the only measurement that describes the product, and it is far too slow and noisy to gate anything. Our p75 LCP moves by up to 300 milliseconds a day on volume alone, and proving a change in a p75 rollup needs five to seven days of comparable traffic, which puts the offending pull request forty merges behind you. It also cannot name the cause: it tells you mid-tier Android in one region got slower, precisely, three days after that stopped being anybody's context.

Lab data is the mirror image: deterministic, fast, attributable to one diff, and wrong in a specific way — usually optimistic, and silent on whether anyone cared.

So we use both. The blocking budget is lab, tuned slightly generous so it fires on real regressions rather than measurement jitter. The pager-worthy signal is field: a weekly review of the 28-day rollup segmented by device class, with one alert that fires when the mid-tier Android p75 LCP moves 200 milliseconds and stays there for three days.

What each one is bad at

The lab will fail you for a 3 KB regression no human can perceive, and it will pass a change that halves the bytes while pushing the hero image behind an interaction. Field data tells you something is wrong with total confidence and no attribution. Lab is the gate. Field is the verdict.

When you blow the budget

The CI job is the easy part. The part that took two years is the agreed response to a red build, because a budget with no response is an obstacle people route around. There are three legitimate ones, in descending order of preference.

The test I use now

If enforcing your budget would need a conversation, you do not have a budget. You have a wish with a unit attached.

  • Delete it. A surprising amount of JavaScript is there because it arrived on a Thursday and nobody has looked since. Twice the answer to "who uses this" was "a dashboard we stopped promoting", and the fix removed four kilobytes.
  • Defer it. If the feature is wanted but not needed for the first interaction, move it off the critical path rather than shrinking it. A date picker that appears on focus does not belong in the initial bundle.
  • Downgrade it. A lighter library, a smaller image format, a native control. The user-visible cost is almost always smaller than the engineer's estimate of it.

The illegitimate response is raising the number. Raising a budget is sometimes correct — twice, for us, it has been — but it is not a response to a failing build. It is a change to the contract: deliberate, by a named person, in a diff, with a reason written down. The failure mode is the silent raise, a threshold nudged five kilobytes inside a large refactor twelve times in eighteen months, until the budget describes what the code already does.

Which is why the escape hatch is time-boxed. A legitimate exception carries an owner, a reason, a byte amount and an expiry date, and the owner is whoever benefits: the engineer who wants the merge, not a manager. Twenty-one days by default, ninety for something with a contract behind it. When the expiry passes, the build goes red on its own, which is the only reminder mechanism I trust. We have used four exceptions in two years; three were resolved by changing code rather than dates.

The regression that taught us the most

In month fourteen we shipped a booking form. It needed date constraints, so the engineer installed a date-picker library. It added 180 KB to the route: its own locale data plus an internationalisation runtime, for a form that exists in one language.

The part that still bothers me is that CI stayed green. Our budget covered the home page and the dashboard, and the home page had just shed sixty kilobytes from an unrelated image change, so the aggregate looked healthy while the booking route became the slowest page on the site for three weeks. Nobody noticed until a customer called the calendar slow on her phone.

A budget on the whole bundle is a budget on an average, and nobody experiences an average. User number four hundred thousand loads one route, on one device, at one moment, and that route is the only one that exists.

From the incident notes after the booking form, and since quoted at me twice by colleagues

We replaced it with a native date input and forty lines of constraint validation, four kilobytes in total. I did learn that a native picker on an older Android ROM is genuinely worse than a custom one; that is the trade we took.

The lasting change was to attribution, not to the number. The script now maps each chunk to the route that imports it, so the cost of a dependency follows the dependency instead of being diluted by everything else on the site.

Two years in, nothing here is clever. It counts bytes, checks the clock, names a person, and fails a build in eleven seconds. The home page is 214 KB today, up from 168 KB when we started, and every kilobyte of that has a commit behind it explaining itself.

Portrait of Elliot Vance

Elliot Vance

Software engineer in Rotterdam. I help teams make their systems fast, observable and boring, usually by removing more than I add. Eleven years in, mostly on web performance, databases and platform work.

More about me