Design systems

Shipping a design system in six weeks

Four product teams, four competing button components and three brand skins. We built the token layer first and let the components fall out of it — which is the only reason the system outlived the project that funded it.

I spent six weeks last spring building a design system at a fintech with two other people. Four competing button components and three hand-maintained brand skins became sixty-four documented components generated from one token pipeline. Nearly everything that went right was decided in the first four days; nearly everything that went wrong was decided in the second week.

The brief, and why it was six weeks

In February the accessibility audit landed and the numbers were bad enough to be politically useful. Four product teams — Payments, Onboarding, Cards and the Ops console — each shipped their own button: Button, PrimaryButton, Btn and ButtonBase. Between them, four line heights, three border radii, two disabled treatments, and one that had forgotten type="button" and quietly submitted the surrounding form whenever a customer clicked "Add payee".

The audit ran axe-core over 87 screens and returned 212 colour-contrast failures, mostly the same three grey-on-grey pairs copied from a mockup produced in 2022. Design also maintained three brand skins — retail, business, and a white-label partner whose palette changed every quarter — as three parallel Figma libraries plus an unofficial spreadsheet. Applying a brand change cost about three weeks of engineering.

There were three of us. Priya owned design and had already written the component inventory, which saved a fortnight of discovery. Tomas owned build tooling and CI, which mattered more than either of us expected. I wrote components and, as it turned out, the token contract, and we borrowed a Payments engineer one day a week for review.

Six weeks was not our estimate. It was the gap between the partner's next delivery and the quarterly review that would decide the budget. So we defined done in terms a machine could check: one button in the repository, with four teams migrated off the old three; sixty-four components documented and a11y-clean in Storybook; three brands generated from one token source; zero contrast failures in CI. Not design approved, not most teams on board.

Tokens first, components second

The one architectural decision from those first four days was to finish the token layer before writing a single component. It was not obvious then: the obvious move is to start with the button, because a button is fun, it unblocks people, and it produces a screenshot for a slide. Tokens are a JSON file nobody can see.

Naming that survives a rebrand

We split tokens into two tiers and never let them touch. Primitives are what you expect: --grey-200, --blue-600, --space-4. They exist, they are useful for tooling, and no component may reference them. Everything else is semantic: --colour-surface-raised, --colour-text-subtle, --colour-focus-ring, --radius-control.

--colour-surface-raised beats --grey-200 for one boring reason: it survives a rebrand. A grey ramp is a claim about what a colour looks like; a surface is a claim about what it is for. The business brand has no use for grey-200 — its raised surface is a very slightly warmed white with a one-pixel border — and the partner renders it as a tinted panel with no border. When we renamed a dozen semantic tokens in week five, no component moved.

Semantic names may not contain a colour, a size or a brand, and state goes last: --colour-surface-raised-hover, never --hover-colour-surface-raised. A token that exactly one component needs is usually a prop, or a mistake. And because the rules were short enough to lint, a sixty-line stylelint plugin now fails the build on any primitive reference inside component CSS.

One pipeline, three outputs

Everything downstream came from one Style Dictionary config. Source tokens live in primitive.json, semantic.json, and one file per brand that may only override semantic values. The pipeline emits three artefacts: CSS custom properties, a TypeScript union type plus a flat object for JavaScript consumers, and a Figma variables JSON that a plugin imports into the design file.

{
  "colour": {
    "surface": {
      "canvas": { "value": "{colour.neutral.0}" },
      "raised": { "value": "{colour.neutral.0}" }
    },
    "text": {
      "default": { "value": "{colour.neutral.900}" },
      "subtle":  { "value": "{colour.neutral.600}" }
    }
  }
}
import StyleDictionary from 'style-dictionary';

const brand = process.env.BRAND ?? 'retail';

export default {
  source: ['tokens/primitive.json', 'tokens/semantic.json', `tokens/brand/${brand}.json`],
  platforms: {
    css: {
      transformGroup: 'css',
      prefix: 'ds',
      buildPath: `dist/${brand}/`,
      files: [{
        destination: 'tokens.css',
        format: 'css/variables',
        options: { outputReferences: true }
      }]
    },
    ts: {
      transformGroup: 'js',
      buildPath: 'dist/types/',
      files: [{ destination: `tokens.${brand}.ts`, format: 'typescript/es6-declarations' }]
    },
    figma: {
      transformGroup: 'js',
      buildPath: 'dist/figma/',
      files: [{ destination: `${brand}.json`, format: 'json/nested' }]
    }
  }
};

The first version of that config did not set outputReferences, which resolves every alias to a literal at build time. That is fine until you switch brands, because the alias chain is gone. A component inheriting a colour through semantic.json kept the retail blue when we compiled the business brand; we shipped it to staging, saw a blue primary button in a green brand, and lost most of a day to it.

Semantic token Retail Business Partner
--colour-surface-canvas#FFFFFF#F7F8FA#FFFFFF
--colour-surface-raised#FFFFFF#FFFFFF#F2F4F7
--colour-text-default#12161C#0E1B2A#16161A
--colour-accent-default#0B57E0#0F6B4F#5B21B6
--colour-focus-ring#0B57E0#0F6B4F#7C3AED
--radius-control8px4px12px

The full override surface is 47 semantic tokens and the three brands between them touch 19. Everything else — spacing, the type scale, motion, z-index — is shared. A brand is a palette and two radii, not a fork of the library.

A token is a promise about a role

If a name would be wrong in a different brand, theme or time of day, it is not a semantic token; it is a primitive wearing a costume. The test we use in review: swap the brand file and read the name aloud. If the name becomes a lie, it belongs a tier down.

Week two, where we nearly lost it

On the Monday of week two we had a good token pipeline and no components, and the pressure to show something demonstrable was considerable. So we did the thing I now tell people not to do: we started building components before the token contract was agreed.

Thirty-one props and a mutiny

The first Button merged on the Wednesday with 31 props. Nobody wanted 31 props; every reviewer wanted one more. Payments needed loading and loadingText, Onboarding needed iconLeft and iconRight, Cards had a pill variant, the Ops console wanted a dense mode, and the partner brand an uppercase treatment. Each request was reasonable and cheaper to grant than to argue about, so the component grew into the union of four teams' requirements: eight variants, four sizes, five tones, nine booleans. That is 160 combinations before the booleans, with stories for 22.

interface ButtonProps {
  variant: 'primary' | 'secondary' | 'ghost' | 'outline' | 'danger'
         | 'payments' | 'onboarding' | 'cards' | 'ops';
  tone?: 'neutral' | 'accent' | 'positive' | 'caution' | 'critical';
  size?: 'xs' | 'sm' | 'md' | 'lg';
  dense?: boolean;
  pill?: boolean;
  uppercase?: boolean;
  loading?: boolean;
  iconLeft?: React.ReactNode;
  iconRight?: React.ReactNode;
  iconOnly?: boolean;
  /* ...and twenty-one more, which is rather the point */
}

The pushback arrived in the pull request, which ran to 140 comments. Priya's objection was the sharpest and correct: nine of those variants were product or brand names, which meant the component had absorbed four organisational charts. A Payments engineer put it in a sentence I have repeated since — if I have to read the source to work out which of the eight primary buttons to use, I will keep my own.

We stopped treating the token file as an implementation detail and started treating it as the interface. The Button was never the interface. It was the first consumer.

From the token contract review, week two, after 140 comments and one very long Thursday

The three-day freeze

We froze component work for three days, Wednesday to Friday of week two. No new components, no new props, no merges except fixes to the token pipeline. In those three days we wrote the token contract as a document that could be reviewed, argued with and versioned.

It is nine pages and contains almost no code. It names every semantic token, says what it means, lists which pairs may be composed — text on surface, border on surface, focus ring on surface — records the primitive each resolves to per brand, and ends with the rules: components may not reference primitives, and a new semantic token needs a written justification and one non-button consumer.

The following Monday we deleted 22 props, moved the variance into tokens, and merged in a day. It shipped with nine. The four teams took the migration pull requests without much complaint, because a written contract let their engineers verify the change themselves instead of trusting us. Losing four days of component work is why the other 63 landed in the remaining four.

Accessibility as a build step, and what six weeks does not buy

Contrast and focus enforced in the token layer

The 212 failures were never a components problem; they were an expressiveness problem. The old system made inaccessible combinations easy to type — anyone could write background: #E5E7EB with color: #FFFFFF and nothing in the toolchain objected. So we moved the check to the place where the combination is really declared: the token layer.

We listed the pairs the system permits in tokens/contrast.pairs.json. A Node script resolves each pair for each brand, computes contrast with the WCAG 2.2 formula, and exits non-zero with the token paths and ratios if anything falls below 4.5:1 for body text, 3:1 for large text and non-text UI, or 3:1 for the focus ring against every adjacent surface. The constraint matters more than the check: a brand file cannot introduce a failing pair without failing the build.

Focus works the same way with fewer moving parts. Every focusable component takes its ring from --colour-focus-ring and --focus-ring-offset through one generated focus.css layer, and a stylelint rule bans outline: none as well as any :focus rule without a matching :focus-visible. The one documented exception is the Ops data grid's cell selection outline.

Storybook in CI, and the parts it cannot see

Storybook runs the a11y addon with test: 'error', and the test runner executes 192 stories on every pull request: 64 components across three brand contexts. It takes four minutes eleven seconds, down from nineteen before Tomas split the brand contexts into a build matrix. Half of the original 212 failures were those repeated grey-on-grey pairs, so the token fix cleared more than a hundred at once.

The suite cannot tell you whether a combobox announces its options in a sensible order, whether a date picker returns focus somewhere useful when it closes, or whether an error message is announced before or after its field. Four manual sessions — NVDA for the Ops console, VoiceOver for Onboarding, a keyboard-only pass over the payment form, 200% text zoom across all three brands — produced eleven issues, none of which the suite found. One was a modal returning focus to document.body instead of the button that opened it.

The disabled state resisted automation too. WCAG exempts disabled controls from contrast requirements, and design had leaned on that exemption to ship 2.9:1 grey-on-grey labels in the Cards flow. We set disabled text at 4.5:1 anyway, on the grounds that exempt is not the same as usable.

Documentation rot

Sixty-four components arrived with sixty-four MDX pages, and for about four months those pages were accurate. Then a rename landed — variant="secondary" became variant="outline" — and eleven pages went on describing props that no longer existed. Nobody noticed for a fortnight, which tells you how often the documentation is read. A CI step now runs react-docgen over the built types and fails when a props table disagrees with the interface.

The contribution model nobody wanted to own

We wrote a contribution ladder — fix a bug, own a component, join the review rota — and almost nobody climbed it. Contributing to a shared library is nobody's roadmap item. In the three months after launch, 78% of merged pull requests came from the original three of us, and our review queue averaged three and a half days.

We fixed it structurally rather than with exhortation. Each product team nominates one steward for one day a week; that day is written into the team's quarterly plan by their manager rather than requested by us; and we hold a 48-hour first-response SLA. Adoption went from 40% of screens at week six to about 82% at month five.

The first component we had to deprecate

IconButton was the first casualty, five weeks after launch. It existed separately only because the Figma library did, and it duplicated about 60% of the button's loading and disabled logic. A jscodeshift codemod touched 214 call sites across 63 files in one pull request, the old export gained a @deprecated tag and a migration guide, and we kept it for two minor versions. Nineteen of those call sites relied on behaviour the new API could not express, which is how we learned the old component had three undocumented features.

The numbers, and the honest caveats

Measure Before After
Button components in the repository41, with 9 props
Documented components064
Brand skins3 hand-maintained libraries3 brand files, 1 pipeline
Contrast failures2120
Theming effort per brand per quarterabout 3 weeks1 config file, 15-minute build
Shared components on product screens0%82% at month five

That zero is measured on our story set and the 87 screens we audited, not on every screen in the product. In month three we found 14 fresh failures in two teams' local overrides: hard-coded hex values in CSS Modules that no token could reach, precisely because those files were not using tokens. Zero within the system is not yet zero everywhere.

Six weeks did not include the six months. For three months afterwards, roughly 40% of my week went to systems work: reviews, migration help, and one silent Figma plugin failure that let the design library drift for three weeks. And sixty-four components counts things that exist, not things that are used.

If I ran it again the order would be nearly the same and so would the mistakes. Tokens first. Freeze component work until the contract is written and reviewed, even when that means four days of visible nothing. Put the accessibility checks where the failure is expressed rather than where it is rendered. Budget as much time for the review SLA and the contribution model as for the components.

Portrait of Elliot Vance

Elliot Vance

Software engineer in Rotterdam. I help teams make their systems fast, observable and boring, usually by removing more than I add. Eleven years in, mostly on web performance, databases and platform work.

More about me