Skip to content

How the catalogue is made

Built by one agent. Argued over by another.

Every component starts as a written brief, gets built by a coding agent, then goes through rounds of adversarial design review against real screenshots before it ships. This page shows the whole process, the numbers behind it and where a person still makes the call.

Made by
  • Claude Code
  • Codex
  • A person
Checked at
  • 360 px
  • 768 px
  • 1280 px

The process in numbers

Components published
64
Median score, first review to last
8.7 to 9.9
Review rounds, median
3
First commit to last, median
1.5 h

The count is today’s catalogue. The other three come from the 65 component pull requests merged between 17 September and 1 October 2026, refinements included, scored out of 10 in each one’s design review.

01 / The pipeline

One issue in. One pull request out.

Each component is one run of the same five phases. The run is automated from the issue to the pull request, and each pull request gets one more review before it merges.

  1. Brief

    Written ahead

    Every component starts as a GitHub issue for one pattern from a plan of 784. It says what to build, how it differs from its neighbours, the variations, props, states, accessibility, fixtures, what is out of scope and how we will know it is done.

    Leaves behind An issue with acceptance criteria

  2. Build

    Claude Code

    The builder researches what a standard version of the pattern looks like, names the one structural idea the component is built around, then follows the playbook: real content first, then structure, foundations, every state, other languages, 360 px and tests.

    Leaves behind A published draft and its first commit

  3. Verify

    Claude Code, Codex

    Type check, lint and tests, then the export is copied into a fresh SvelteKit project with nothing from ours and built there. Codex reviews the code against the export contract, accessibility and the acceptance criteria, and the fixes are committed.

    Leaves behind A green build and a code review

  4. Design review

    Codex reviews, Claude Code fixes

    Rounds of capture, review, verification and fixes against real screenshots, scored out of ten against our design standard. This is where a good component becomes a crafted one, and it is most of the work.

    Leaves behind A score table, round by round

  5. Ship

    Claude Code

    Thumbnails and preview sizes are regenerated, the full build runs, and a pull request goes up with the score table, the decisions made along the way and the checks that ran. It gets one more review before it merges.

    Leaves behind A merged pull request and a new catalogue entry

02 / The starting point

Nobody starts from a blank prompt.

Before the builder writes a line it has three documents: the issue that says what this component is, the standard that says what good looks like, and the playbook that says what order to build it in. The last two are rewritten whenever a review finds something that will come up again.

It also researches the pattern itself: the parts a competent version always has, the states that carry weight, the accessibility contract, and the mistakes that make one look amateur. Then it names the single structural idea the component will be built around, in one sentence.

New component: Background particles (C783)

  1. What to build
  2. How this differs from similar patterns
  3. Proposed identity
  4. Variations
  5. Anatomy
  6. Props sketch
  7. States
  8. Behaviour
  9. Accessibility
  10. Dependencies and services
  11. Out of scope
  12. Preview fixtures
  13. Acceptance criteria
  14. Definition of done
The headings of a real, open issue. There are 14 of them, and the builder answers to every one.

The standard

Eight hundred lines on what a PageSugar component should look like, and the rubric it is scored against. Its bar, in one sentence: a senior product designer would look at the component and not find a change worth making.

  • Three text tones, one type scale, tracking that tightens with size
  • A 4 px grid with at most five spacing values
  • One radius scale, nested corners that share a centre
  • Hairlines before shadows, and one elevated surface at most
  • Motion only on a change of state, with a reduced-motion path
  • Copy that reads like a real product, never lorem ipsum or invented proof

The playbook

The order to build in, one step for each thing reviews kept catching. Doing it in the build is what lets the review spend its rounds on craft instead of on things we already know.

  1. Write the content before the markup
  2. Name the structure and the refinement
  3. Start from the foundations
  4. Design every state
  5. Make it work in other languages
  6. Make it work at 360 px
  7. Write the tests with the component
  8. Lint, capture, then review

03 / The adversarial review

A second model, looking for what is wrong.

The builder is good at making things work and bad at seeing its own work. So the review is done by a model from a different company, on real screenshots, against a written standard. It has to commit to a number, and it has to rank what it found by how much each fix would move that number. Then the builder checks every claim before it changes anything.

  1. Capture

    Every fixture, palette and state at 360, 768 and 1280 px through the real preview route, plus hover, focus, pressed and reduced motion.

  2. Measure

    The harness records overflow at 360 px, targets under 44 px and how far the layout moves between states.

  3. Describe

    A separate agent transcribes each screenshot, text and all, and lists anything that renders wrong.

  4. Review

    Codex sees the pixels, reads the code, commits to a score and ranks its findings by what each would move.

  5. Verify

    Factual claims are checked against the source and the measurements before anything is conceded.

  6. Fix

    The builder fixes what is real, records what it declines and why, and pins each fix with a test.

  7. Score again

    The next round reviews the fixed component from scratch, with notes on what changed, until the score stops moving.

Three rounds by default. A fourth runs only if the third scored below 9.8, a fifth only if the fourth scored below 9.5. Every component review runs the reviewer at maximum effort.

Who does what

Builder Claude Code
Writes the component, reads every screenshot itself and decides which findings are right. The reviewer proposes; the builder rules.
Reviewer Codex
A model from a different company, for a second view that does not start from the builder’s assumptions. Scores from pixels and source, at maximum effort.
Eyes A vision agent
Transcribes each capture without judging it. Its defect list is cross-checked against the reviewer’s.
Verifiers Read-only agents
Test the reviewer’s claims against the real source. A cropped capture once produced a false “cards touch the edge”.
Research Google
What the pattern is called, its accessibility contract and the mistakes it is known for, before the brief is written.
Taste Us
We watch reviews as they run and step in. When our eye and the score disagree, our eye wins and the rule gets written down.

04 / The score

Correct is a 6. Crafted starts at 8.

Eight dimensions, each scored from 0 to 10 and averaged. Accessibility and the export contract are not among them: they are gates, and a component does not pass a round with a contrast failure however good it looks.

1 to 3 Broken No hierarchy, failing contrast, overflow.
4 to 5 Placeholder Works, but could be any template.
6 to 7 Competent Correct, with no idea and no refinement.
8 to 9 Crafted A clear idea, refined details. Medians: first review 8.7, last review 9.9
10 Nothing to change A senior designer would not touch it.
First review 8.7 Last review 9.9
The bands are categories, drawn at equal widths. The median first and last review scores sit inside the band they fall in.
  1. Hierarchy and focus

    Blur the capture. Is exactly one thing dominant, and is it the right one?

  2. Composition

    Can you name its one structural idea in a sentence, and does it survive at 360 px?

  3. Typography

    A real scale, weight contrast, tracking that tightens with size, a bounded measure.

  4. Spacing and rhythm

    Five spacing values or fewer, grouped by proximity, on one left edge.

  5. Colour and surfaces

    One neutral ramp, three text tones, hairlines over shadows, one elevation at most.

  6. Detail and states

    Concentric radii, a focus ring that fits, hover only where a click lands, every state designed.

  7. Responsiveness

    Reflowed rather than scaled. No overflow at 360 px, and targets clear 44 px.

  8. Content and honesty

    Copy that reads like a real product, no invented proof, and the default as the best state.

Applied before every score

  • Correct is a 6. No technical flaw and no idea scores at most 6. Correctness is the floor.
  • Copy vacuum caps at 5. Placeholder headlines or meaningless numbers in any fixture cap the score, whatever the layout.
  • The one-change test. Before a 10, the reviewer has to try to name one change worth making. If it can, the score is 9.
  • Deltas are earned by pixels. A round that changed code but not the rendered result does not move the score.

05 / Real reviews

Three reviews, round by round.

Quoted from the pull requests that shipped them. Most runs are quieter than these; we picked the ones where the process had to argue with itself.

  • Section wrapper Layout · 4 rounds

    Round two came back a perfect 10 that did not square with what its fixes were predicted to be worth, so round three re‑scored the same code adversarially.

    1. R1 8.7 Buttons moved under their copy on the shared edge, facts separated by space instead of a rule, the lede held under 70 characters, Arabic eyebrows given a size step instead of capitals.
    2. R2 10.0 (thrown out) Nothing named. Treated as inflated and re‑scored.
    3. R3 9.3 Fact values set at 600 over 400 labels; the header’s last link lined up with the content edge.
    4. R4 9.6 Nothing left above the threshold.
    5. Gate The consistency survey found the wrapper put its gutters inside the container where every other section put them outside. Moved to match.
  • Product screenshot hero Hero · 3 rounds

    At 360 px the product screenshot’s own text could not be read, and the first round capped the score at 6 for it.

    1. R1 6.0 Capped at 6: the screenshot’s text was unreadable at 360 px.
    2. R2 Not recorded Phone and tablet renditions of the screenshot, drawn at the size they display.
    3. R3 9.9 Fixture copy that did not match its pictures rewritten. The last fix re‑scored at 10.
  • Site header Navigation · 10 rounds

    The longest review so far, and the one that taught us the most. In round four the reviewer capped the header for being conventional, and we disagreed.

    1. R1, 2 5.0 Held at 5 by the copy-vacuum rule.
    2. R3 8.6 Copy fixed, and the cap lifted.
    3. R4 6.0 (thrown out) Capped for being a conventional header. A full-width sheet was built to beat the cap, looked wrong, and was reverted.
    4. R5 to 10 9.5 8.4, 9.3, 9.0, 9.4, 9.4, 9.5, judged on the craft of a conventional structure: press states, close timing, tones, the drawer.

06 / How long it takes

About an hour and a half, most of it review.

The median component took 1.5 h from its first commit to its last, and none of the 65 took more than 2.8 h. A few early ones were reworked a day or two later; only their first run of work is counted here. The clock starts at the first commit, so the reading, research and build before it are not counted either.

  • Build and code review 13 min
  • Design review 1 h 29 min
  • Ship 5 min
  • Waiting to merge 42 min
  1. start Built Component, mock screenshots and tests
  2. +13 min Code review Hairline and alt text fixes from Codex
  3. +54 min Round 1 fixes Phone crops drawn at the size they show
  4. +1 h 17 min Round 2 fixes A tablet view between phone and desktop
  5. +1 h 33 min Round 3 fixes The board lists every card its count promises
  6. +1 h 42 min Captured Thumbnail and preview sizes
  7. +1 h 47 min Pull request Opened with the score table
  8. +2 h 29 min Merged Reviewed once more, then merged
The commits behind the product screenshot hero, 30 September. Design review was 1 h 29 min of the 2 h 29 min.
Under two hours, first commit to last
51 of 65
Commits per component, median
7
Pull request open to merged, median
44 min

07 / Getting better

Every review teaches the next build.

When a review finds something that will come up again, it goes into the standard or the playbook in the same pull request, so later builds start closer to the bar. The first ten pull requests averaged 7.5 in their first round; the last ten averaged 8.5.

  • First review round
  • Last review round
First and last review scores for all 65 component pull requests, in the order they merged. These are one model’s review scores, not user testing. Scroll the chart sideways to see every one.
All 65 review runs

Scroll the table sideways for every column.

Every component pull request in the snapshot, in merge order
Merge orderFirst reviewLast reviewRoundsFirst burst of commitsCommitsOpen to merged
15.79.040.7 h1338.3 h
25.09.5101.0 h3043.1 h
36.08.631.2 h1119.3 h
48.59.930.9 h78.3 h
58.38.520.7 h715 min
67.99.031.1 h811 min
77.98.731.1 h811 min
88.310.030.4 h71.9 h
98.610.020.9 h53.6 h
108.49.931.3 h63.1 h
118.99.631.5 h72.8 h
129.010.032.0 h52.7 h
139.010.032.0 h52.6 h
147.59.931.8 h82.4 h
159.410.020.8 h72.3 h
168.110.052.2 h72.1 h
178.79.642.4 h81.9 h
188.810.021.0 h558 min
199.39.931.4 h545 min
208.99.931.7 h640 min
218.510.031.8 h822 min
228.49.541.7 h78 min
239.010.021.2 h54 min
248.99.531.5 h631 min
258.810.032.3 h715 min
268.910.031.5 h840 min
278.89.432.0 h726 min
289.010.031.4 h928 min
299.110.032.1 h53 min
308.810.032.3 h116 min
318.69.831.5 h71.0 h
328.89.831.9 h652 min
338.010.032.0 h741 min
348.89.431.2 h620 min
358.99.931.7 h51.3 h
368.09.341.6 h855 min
379.510.031.0 h544 min
386.09.931.7 h643 min
398.810.031.6 h540 min
408.59.541.7 h914 min
418.69.452.8 h95 min
428.49.342.4 h95.3 h
438.610.031.5 h65.2 h
449.410.021.5 h45.2 h
459.09.321.4 h45.0 h
468.910.022.0 h44.7 h
478.69.942.6 h84.6 h
489.410.032.2 h84.5 h
499.09.930.8 h643 min
507.910.031.1 h535 min
518.810.031.2 h528 min
528.39.931.3 h726 min
538.08.931.4 h617 min
548.89.131.8 h751 min
559.39.931.0 h613 min
568.89.931.5 h555 min
578.410.041.4 h755 min
587.59.731.4 h652 min
597.89.241.8 h716 min
609.39.831.7 h125 min
618.410.031.3 h531 min
628.99.931.0 h619 min
638.69.931.3 h64 min
648.39.931.3 h84 min
658.910.031.0 h759 min

What the first seven reviews kept finding

Each of these is now handled before the review starts, so a round is no longer spent on it.

FindingSeenHandled by
Weak default copy in the preview7 of 7Content is step one
Spacing over budget for the section6 of 7Linted
Right-to-left tracking and CJK wrapping6 of 7Step five, other languages
Reviewer pushing an unusual structure4 of 7Written into house taste
Layout moving when state changes4 of 7Measured every capture
Coverage not recorded3 of 7Validation fails without it

Your agent can join the loop

When an agent connected over MCP finds a problem with a component, it can send it to us with submit_component_feedback. Feedback is screened as it arrives, triaged every week, and anything worth acting on becomes a refine issue that runs through the same pipeline as a new component.

It is new. The four refinements so far came from our own audits, not from feedback.

08 / Where a person decides

The score drives the loop. It does not get the last word.

We watch reviews as they run and step in. In the site header’s fourth round the reviewer capped it at 6 for being a conventional header, and suggested a full-width sheet instead. We built the sheet. It scored higher and it looked wrong, so it was reverted, and the rule went into the standard: conventional structures, polished by craft.

Every rule in that section was learned the same way, by shipping the opposite and rejecting it. The reviewer is told them up front, as constraints to work around rather than findings to score against.

  • Conventional structures, polished by craft.

    When a reviewer asks for a more distinctive structure, we decline and improve type, spacing, states and copy instead.

  • Keep presence.

    Deleting an element to raise a score reads as losing personality. Fix the badge; do not remove it.

  • Open with air.

    Heroes and introductions keep their padding, even when a reviewer prices the space.

  • Hover only where a click lands.

    Proved with a real hover and a real click in a browser, not by reading the CSS.

09 / What a score means

A 9.9 is a careful opinion, not a guarantee.

What it tells you

  • One model’s judgement of real screenshots against a written standard, with its reasoning recorded.
  • Its factual claims were checked against the source before anything changed.
  • Every fixture and state the component has was captured at three widths and looked at.
  • Accessibility and the export contract passed as gates, not traded for looks.

What it does not

  • User testing. Nobody watched a visitor use it.
  • Perfection. A 10 means the reviewer could not name a change, and reviewers inflate: we have thrown out a 10 that re‑scored 9.3.
  • That it fits your site as is. It is a starting point your agent adapts.
  • That it has no bugs. If you find one, your agent can tell us through the MCP server.

Found something wrong? Here is how to report it.

Judge the result yourself.

Every component in the catalogue went through this. Open one, read its source, and see whether the review was right.