How the catalogue is made
Built by one agent. Argued over by another.
Every component starts as a written brief, gets built by a coding agent, then goes through rounds of adversarial design review against real screenshots before it ships. This page shows the whole process, the numbers behind it and where a person still makes the call.
- Made by
- Claude Code
- Codex
- A person
- Checked at
- 360 px
- 768 px
- 1280 px
The process in numbers
- Components published
- 64
- Median score, first review to last
- 8.7 to 9.9
- Review rounds, median
- 3
- First commit to last, median
- 1.5 h
The count is today’s catalogue. The other three come from the 65 component pull requests merged between 17 September and 1 October 2026, refinements included, scored out of 10 in each one’s design review.
01 / The pipeline
One issue in. One pull request out.
Each component is one run of the same five phases. The run is automated from the issue to the pull request, and each pull request gets one more review before it merges.
Brief
Written ahead
Every component starts as a GitHub issue for one pattern from a plan of 784. It says what to build, how it differs from its neighbours, the variations, props, states, accessibility, fixtures, what is out of scope and how we will know it is done.
Leaves behind An issue with acceptance criteria
Build
Claude Code
The builder researches what a standard version of the pattern looks like, names the one structural idea the component is built around, then follows the playbook: real content first, then structure, foundations, every state, other languages, 360 px and tests.
Leaves behind A published draft and its first commit
Verify
Claude Code, Codex
Type check, lint and tests, then the export is copied into a fresh SvelteKit project with nothing from ours and built there. Codex reviews the code against the export contract, accessibility and the acceptance criteria, and the fixes are committed.
Leaves behind A green build and a code review
Design review
Codex reviews, Claude Code fixes
Rounds of capture, review, verification and fixes against real screenshots, scored out of ten against our design standard. This is where a good component becomes a crafted one, and it is most of the work.
Leaves behind A score table, round by round
Ship
Claude Code
Thumbnails and preview sizes are regenerated, the full build runs, and a pull request goes up with the score table, the decisions made along the way and the checks that ran. It gets one more review before it merges.
Leaves behind A merged pull request and a new catalogue entry
02 / The starting point
Nobody starts from a blank prompt.
Before the builder writes a line it has three documents: the issue that says what this component is, the standard that says what good looks like, and the playbook that says what order to build it in. The last two are rewritten whenever a review finds something that will come up again.
It also researches the pattern itself: the parts a competent version always has, the states that carry weight, the accessibility contract, and the mistakes that make one look amateur. Then it names the single structural idea the component will be built around, in one sentence.
New component: Background particles (C783)
- What to build
- How this differs from similar patterns
- Proposed identity
- Variations
- Anatomy
- Props sketch
- States
- Behaviour
- Accessibility
- Dependencies and services
- Out of scope
- Preview fixtures
- Acceptance criteria
- Definition of done
The standard
Eight hundred lines on what a PageSugar component should look like, and the rubric it is scored against. Its bar, in one sentence: a senior product designer would look at the component and not find a change worth making.
- Three text tones, one type scale, tracking that tightens with size
- A 4 px grid with at most five spacing values
- One radius scale, nested corners that share a centre
- Hairlines before shadows, and one elevated surface at most
- Motion only on a change of state, with a reduced-motion path
- Copy that reads like a real product, never lorem ipsum or invented proof
The playbook
The order to build in, one step for each thing reviews kept catching. Doing it in the build is what lets the review spend its rounds on craft instead of on things we already know.
- Write the content before the markup
- Name the structure and the refinement
- Start from the foundations
- Design every state
- Make it work in other languages
- Make it work at 360 px
- Write the tests with the component
- Lint, capture, then review
03 / The adversarial review
A second model, looking for what is wrong.
The builder is good at making things work and bad at seeing its own work. So the review is done by a model from a different company, on real screenshots, against a written standard. It has to commit to a number, and it has to rank what it found by how much each fix would move that number. Then the builder checks every claim before it changes anything.
-
Capture
Every fixture, palette and state at 360, 768 and 1280 px through the real preview route, plus hover, focus, pressed and reduced motion.
-
Measure
The harness records overflow at 360 px, targets under 44 px and how far the layout moves between states.
-
Describe
A separate agent transcribes each screenshot, text and all, and lists anything that renders wrong.
-
Review
Codex sees the pixels, reads the code, commits to a score and ranks its findings by what each would move.
-
Verify
Factual claims are checked against the source and the measurements before anything is conceded.
-
Fix
The builder fixes what is real, records what it declines and why, and pins each fix with a test.
-
Score again
The next round reviews the fixed component from scratch, with notes on what changed, until the score stops moving.
Three rounds by default. A fourth runs only if the third scored below 9.8, a fifth only if the fourth scored below 9.5. Every component review runs the reviewer at maximum effort.
Who does what
- Builder Claude Code
- Writes the component, reads every screenshot itself and decides which findings are right. The reviewer proposes; the builder rules.
- Reviewer Codex
- A model from a different company, for a second view that does not start from the builder’s assumptions. Scores from pixels and source, at maximum effort.
- Eyes A vision agent
- Transcribes each capture without judging it. Its defect list is cross-checked against the reviewer’s.
- Verifiers Read-only agents
- Test the reviewer’s claims against the real source. A cropped capture once produced a false “cards touch the edge”.
- Research Google
- What the pattern is called, its accessibility contract and the mistakes it is known for, before the brief is written.
- Taste Us
- We watch reviews as they run and step in. When our eye and the score disagree, our eye wins and the rule gets written down.
04 / The score
Correct is a 6. Crafted starts at 8.
Eight dimensions, each scored from 0 to 10 and averaged. Accessibility and the export contract are not among them: they are gates, and a component does not pass a round with a contrast failure however good it looks.
-
Hierarchy and focus
Blur the capture. Is exactly one thing dominant, and is it the right one?
-
Composition
Can you name its one structural idea in a sentence, and does it survive at 360 px?
-
Typography
A real scale, weight contrast, tracking that tightens with size, a bounded measure.
-
Spacing and rhythm
Five spacing values or fewer, grouped by proximity, on one left edge.
-
Colour and surfaces
One neutral ramp, three text tones, hairlines over shadows, one elevation at most.
-
Detail and states
Concentric radii, a focus ring that fits, hover only where a click lands, every state designed.
-
Responsiveness
Reflowed rather than scaled. No overflow at 360 px, and targets clear 44 px.
-
Content and honesty
Copy that reads like a real product, no invented proof, and the default as the best state.
Applied before every score
- Correct is a 6. No technical flaw and no idea scores at most 6. Correctness is the floor.
- Copy vacuum caps at 5. Placeholder headlines or meaningless numbers in any fixture cap the score, whatever the layout.
- The one-change test. Before a 10, the reviewer has to try to name one change worth making. If it can, the score is 9.
- Deltas are earned by pixels. A round that changed code but not the rendered result does not move the score.
05 / Real reviews
Three reviews, round by round.
Quoted from the pull requests that shipped them. Most runs are quieter than these; we picked the ones where the process had to argue with itself.
- Section wrapper Layout · 4 rounds
Round two came back a perfect 10 that did not square with what its fixes were predicted to be worth, so round three re‑scored the same code adversarially.
- R1 8.7 Buttons moved under their copy on the shared edge, facts separated by space instead of a rule, the lede held under 70 characters, Arabic eyebrows given a size step instead of capitals.
- R2 10.0 (thrown out) Nothing named. Treated as inflated and re‑scored.
- R3 9.3 Fact values set at 600 over 400 labels; the header’s last link lined up with the content edge.
- R4 9.6 Nothing left above the threshold.
- Gate The consistency survey found the wrapper put its gutters inside the container where every other section put them outside. Moved to match.
- Product screenshot hero Hero · 3 rounds
At 360 px the product screenshot’s own text could not be read, and the first round capped the score at 6 for it.
- R1 6.0 Capped at 6: the screenshot’s text was unreadable at 360 px.
- R2 Not recorded Phone and tablet renditions of the screenshot, drawn at the size they display.
- R3 9.9 Fixture copy that did not match its pictures rewritten. The last fix re‑scored at 10.
- Site header Navigation · 10 rounds
The longest review so far, and the one that taught us the most. In round four the reviewer capped the header for being conventional, and we disagreed.
- R1, 2 5.0 Held at 5 by the copy-vacuum rule.
- R3 8.6 Copy fixed, and the cap lifted.
- R4 6.0 (thrown out) Capped for being a conventional header. A full-width sheet was built to beat the cap, looked wrong, and was reverted.
- R5 to 10 9.5 8.4, 9.3, 9.0, 9.4, 9.4, 9.5, judged on the craft of a conventional structure: press states, close timing, tones, the drawer.
06 / How long it takes
About an hour and a half, most of it review.
The median component took 1.5 h from its first commit to its last, and none of the 65 took more than 2.8 h. A few early ones were reworked a day or two later; only their first run of work is counted here. The clock starts at the first commit, so the reading, research and build before it are not counted either.
- Build and code review 13 min
- Design review 1 h 29 min
- Ship 5 min
- Waiting to merge 42 min
- start Built Component, mock screenshots and tests
- +13 min Code review Hairline and alt text fixes from Codex
- +54 min Round 1 fixes Phone crops drawn at the size they show
- +1 h 17 min Round 2 fixes A tablet view between phone and desktop
- +1 h 33 min Round 3 fixes The board lists every card its count promises
- +1 h 42 min Captured Thumbnail and preview sizes
- +1 h 47 min Pull request Opened with the score table
- +2 h 29 min Merged Reviewed once more, then merged
- Under two hours, first commit to last
- 51 of 65
- Commits per component, median
- 7
- Pull request open to merged, median
- 44 min
07 / Getting better
Every review teaches the next build.
When a review finds something that will come up again, it goes into the standard or the playbook in the same pull request, so later builds start closer to the bar. The first ten pull requests averaged 7.5 in their first round; the last ten averaged 8.5.
- First review round
- Last review round
All 65 review runs
Scroll the table sideways for every column.
| Merge order | First review | Last review | Rounds | First burst of commits | Commits | Open to merged |
|---|---|---|---|---|---|---|
| 1 | 5.7 | 9.0 | 4 | 0.7 h | 13 | 38.3 h |
| 2 | 5.0 | 9.5 | 10 | 1.0 h | 30 | 43.1 h |
| 3 | 6.0 | 8.6 | 3 | 1.2 h | 11 | 19.3 h |
| 4 | 8.5 | 9.9 | 3 | 0.9 h | 7 | 8.3 h |
| 5 | 8.3 | 8.5 | 2 | 0.7 h | 7 | 15 min |
| 6 | 7.9 | 9.0 | 3 | 1.1 h | 8 | 11 min |
| 7 | 7.9 | 8.7 | 3 | 1.1 h | 8 | 11 min |
| 8 | 8.3 | 10.0 | 3 | 0.4 h | 7 | 1.9 h |
| 9 | 8.6 | 10.0 | 2 | 0.9 h | 5 | 3.6 h |
| 10 | 8.4 | 9.9 | 3 | 1.3 h | 6 | 3.1 h |
| 11 | 8.9 | 9.6 | 3 | 1.5 h | 7 | 2.8 h |
| 12 | 9.0 | 10.0 | 3 | 2.0 h | 5 | 2.7 h |
| 13 | 9.0 | 10.0 | 3 | 2.0 h | 5 | 2.6 h |
| 14 | 7.5 | 9.9 | 3 | 1.8 h | 8 | 2.4 h |
| 15 | 9.4 | 10.0 | 2 | 0.8 h | 7 | 2.3 h |
| 16 | 8.1 | 10.0 | 5 | 2.2 h | 7 | 2.1 h |
| 17 | 8.7 | 9.6 | 4 | 2.4 h | 8 | 1.9 h |
| 18 | 8.8 | 10.0 | 2 | 1.0 h | 5 | 58 min |
| 19 | 9.3 | 9.9 | 3 | 1.4 h | 5 | 45 min |
| 20 | 8.9 | 9.9 | 3 | 1.7 h | 6 | 40 min |
| 21 | 8.5 | 10.0 | 3 | 1.8 h | 8 | 22 min |
| 22 | 8.4 | 9.5 | 4 | 1.7 h | 7 | 8 min |
| 23 | 9.0 | 10.0 | 2 | 1.2 h | 5 | 4 min |
| 24 | 8.9 | 9.5 | 3 | 1.5 h | 6 | 31 min |
| 25 | 8.8 | 10.0 | 3 | 2.3 h | 7 | 15 min |
| 26 | 8.9 | 10.0 | 3 | 1.5 h | 8 | 40 min |
| 27 | 8.8 | 9.4 | 3 | 2.0 h | 7 | 26 min |
| 28 | 9.0 | 10.0 | 3 | 1.4 h | 9 | 28 min |
| 29 | 9.1 | 10.0 | 3 | 2.1 h | 5 | 3 min |
| 30 | 8.8 | 10.0 | 3 | 2.3 h | 11 | 6 min |
| 31 | 8.6 | 9.8 | 3 | 1.5 h | 7 | 1.0 h |
| 32 | 8.8 | 9.8 | 3 | 1.9 h | 6 | 52 min |
| 33 | 8.0 | 10.0 | 3 | 2.0 h | 7 | 41 min |
| 34 | 8.8 | 9.4 | 3 | 1.2 h | 6 | 20 min |
| 35 | 8.9 | 9.9 | 3 | 1.7 h | 5 | 1.3 h |
| 36 | 8.0 | 9.3 | 4 | 1.6 h | 8 | 55 min |
| 37 | 9.5 | 10.0 | 3 | 1.0 h | 5 | 44 min |
| 38 | 6.0 | 9.9 | 3 | 1.7 h | 6 | 43 min |
| 39 | 8.8 | 10.0 | 3 | 1.6 h | 5 | 40 min |
| 40 | 8.5 | 9.5 | 4 | 1.7 h | 9 | 14 min |
| 41 | 8.6 | 9.4 | 5 | 2.8 h | 9 | 5 min |
| 42 | 8.4 | 9.3 | 4 | 2.4 h | 9 | 5.3 h |
| 43 | 8.6 | 10.0 | 3 | 1.5 h | 6 | 5.2 h |
| 44 | 9.4 | 10.0 | 2 | 1.5 h | 4 | 5.2 h |
| 45 | 9.0 | 9.3 | 2 | 1.4 h | 4 | 5.0 h |
| 46 | 8.9 | 10.0 | 2 | 2.0 h | 4 | 4.7 h |
| 47 | 8.6 | 9.9 | 4 | 2.6 h | 8 | 4.6 h |
| 48 | 9.4 | 10.0 | 3 | 2.2 h | 8 | 4.5 h |
| 49 | 9.0 | 9.9 | 3 | 0.8 h | 6 | 43 min |
| 50 | 7.9 | 10.0 | 3 | 1.1 h | 5 | 35 min |
| 51 | 8.8 | 10.0 | 3 | 1.2 h | 5 | 28 min |
| 52 | 8.3 | 9.9 | 3 | 1.3 h | 7 | 26 min |
| 53 | 8.0 | 8.9 | 3 | 1.4 h | 6 | 17 min |
| 54 | 8.8 | 9.1 | 3 | 1.8 h | 7 | 51 min |
| 55 | 9.3 | 9.9 | 3 | 1.0 h | 6 | 13 min |
| 56 | 8.8 | 9.9 | 3 | 1.5 h | 5 | 55 min |
| 57 | 8.4 | 10.0 | 4 | 1.4 h | 7 | 55 min |
| 58 | 7.5 | 9.7 | 3 | 1.4 h | 6 | 52 min |
| 59 | 7.8 | 9.2 | 4 | 1.8 h | 7 | 16 min |
| 60 | 9.3 | 9.8 | 3 | 1.7 h | 12 | 5 min |
| 61 | 8.4 | 10.0 | 3 | 1.3 h | 5 | 31 min |
| 62 | 8.9 | 9.9 | 3 | 1.0 h | 6 | 19 min |
| 63 | 8.6 | 9.9 | 3 | 1.3 h | 6 | 4 min |
| 64 | 8.3 | 9.9 | 3 | 1.3 h | 8 | 4 min |
| 65 | 8.9 | 10.0 | 3 | 1.0 h | 7 | 59 min |
What the first seven reviews kept finding
Each of these is now handled before the review starts, so a round is no longer spent on it.
| Finding | Seen | Handled by |
|---|---|---|
| Weak default copy in the preview | 7 of 7 | Content is step one |
| Spacing over budget for the section | 6 of 7 | Linted |
| Right-to-left tracking and CJK wrapping | 6 of 7 | Step five, other languages |
| Reviewer pushing an unusual structure | 4 of 7 | Written into house taste |
| Layout moving when state changes | 4 of 7 | Measured every capture |
| Coverage not recorded | 3 of 7 | Validation fails without it |
Your agent can join the loop
When an agent connected over MCP finds a problem with a component, it can send it to us
with submit_component_feedback. Feedback is screened as it arrives, triaged
every week, and anything worth acting on becomes a refine issue that runs through the
same pipeline as a new component.
- Your agent sends a note
- Screened for relevance
- Weekly triage into an issue
- Built, reviewed and shipped
It is new. The four refinements so far came from our own audits, not from feedback.
08 / Where a person decides
The score drives the loop. It does not get the last word.
We watch reviews as they run and step in. In the site header’s fourth round the reviewer capped it at 6 for being a conventional header, and suggested a full-width sheet instead. We built the sheet. It scored higher and it looked wrong, so it was reverted, and the rule went into the standard: conventional structures, polished by craft.
Every rule in that section was learned the same way, by shipping the opposite and rejecting it. The reviewer is told them up front, as constraints to work around rather than findings to score against.
Conventional structures, polished by craft.
When a reviewer asks for a more distinctive structure, we decline and improve type, spacing, states and copy instead.
Keep presence.
Deleting an element to raise a score reads as losing personality. Fix the badge; do not remove it.
Open with air.
Heroes and introductions keep their padding, even when a reviewer prices the space.
Hover only where a click lands.
Proved with a real hover and a real click in a browser, not by reading the CSS.
09 / What a score means
A 9.9 is a careful opinion, not a guarantee.
What it tells you
- One model’s judgement of real screenshots against a written standard, with its reasoning recorded.
- Its factual claims were checked against the source before anything changed.
- Every fixture and state the component has was captured at three widths and looked at.
- Accessibility and the export contract passed as gates, not traded for looks.
What it does not
- User testing. Nobody watched a visitor use it.
- Perfection. A 10 means the reviewer could not name a change, and reviewers inflate: we have thrown out a 10 that re‑scored 9.3.
- That it fits your site as is. It is a starting point your agent adapts.
- That it has no bugs. If you find one, your agent can tell us through the MCP server.
Found something wrong? Here is how to report it.
Judge the result yourself.
Every component in the catalogue went through this. Open one, read its source, and see whether the review was right.


