# How components are made

Every component starts as a written brief, gets built by a coding agent (Claude Code), then goes through rounds of adversarial design review by a second model (Codex) against real screenshots before it ships. Each pull request gets one more review before it merges.

- Components published: 64
- Median score, first review to last: 8.7 to 9.9 (out of 10)
- Review rounds, median: 3
- First commit to last, median: 1.5 h

Figures from the 65 component pull requests merged between 17 September and 1 October 2026.

## The pipeline

1. **Brief** (Written ahead). Every component starts as a GitHub issue for one pattern from a plan of 784. It says what to build, how it differs from its neighbours, the variations, props, states, accessibility, fixtures, what is out of scope and how we will know it is done. Leaves behind: an issue with acceptance criteria.
2. **Build** (Claude Code). The builder researches what a standard version of the pattern looks like, names the one structural idea the component is built around, then follows the playbook: real content first, then structure, foundations, every state, other languages, 360 px and tests. Leaves behind: a published draft and its first commit.
3. **Verify** (Claude Code, Codex). Type check, lint and tests, then the export is copied into a fresh SvelteKit project with nothing from ours and built there. Codex reviews the code against the export contract, accessibility and the acceptance criteria, and the fixes are committed. Leaves behind: a green build and a code review.
4. **Design review** (Codex reviews, Claude Code fixes). Rounds of capture, review, verification and fixes against real screenshots, scored out of ten against our design standard. This is where a good component becomes a crafted one, and it is most of the work. Leaves behind: a score table, round by round.
5. **Ship** (Claude Code). Thumbnails and preview sizes are regenerated, the full build runs, and a pull request goes up with the score table, the decisions made along the way and the checks that ran. It gets one more review before it merges. Leaves behind: a merged pull request and a new catalogue entry.

## The starting point

Before the builder writes a line it has the issue, the design standard and the build playbook, and it has researched the pattern itself. Every new component issue carries these sections:

- What to build
- How this differs from similar patterns
- Proposed identity
- Variations
- Anatomy
- Props sketch
- States
- Behaviour
- Accessibility
- Dependencies and services
- Out of scope
- Preview fixtures
- Acceptance criteria
- Definition of done

The standard fixes, among other things:

- Three text tones, one type scale, tracking that tightens with size
- A 4 px grid with at most five spacing values
- One radius scale, nested corners that share a centre
- Hairlines before shadows, and one elevated surface at most
- Motion only on a change of state, with a reduced-motion path
- Copy that reads like a real product, never lorem ipsum or invented proof

The playbook, in order:

1. Write the content before the markup
2. Name the structure and the refinement
3. Start from the foundations
4. Design every state
5. Make it work in other languages
6. Make it work at 360 px
7. Write the tests with the component
8. Lint, capture, then review

## The adversarial review

One round:

1. **Capture.** Every fixture, palette and state at 360, 768 and 1280 px through the real preview route, plus hover, focus, pressed and reduced motion.
2. **Measure.** The harness records overflow at 360 px, targets under 44 px and how far the layout moves between states.
3. **Describe.** A separate agent transcribes each screenshot, text and all, and lists anything that renders wrong.
4. **Review.** Codex sees the pixels, reads the code, commits to a score and ranks its findings by what each would move.
5. **Verify.** Factual claims are checked against the source and the measurements before anything is conceded.
6. **Fix.** The builder fixes what is real, records what it declines and why, and pins each fix with a test.
7. **Score again.** The next round reviews the fixed component from scratch, with notes on what changed, until the score stops moving.

Three rounds by default. A fourth runs only if the third scored below 9.8, a fifth only if the fourth scored below 9.5. Every component review runs the reviewer at maximum effort.

- **Builder** (Claude Code): Writes the component, reads every screenshot itself and decides which findings are right. The reviewer proposes; the builder rules.
- **Reviewer** (Codex): A model from a different company, for a second view that does not start from the builder’s assumptions. Scores from pixels and source, at maximum effort.
- **Eyes** (A vision agent): Transcribes each capture without judging it. Its defect list is cross-checked against the reviewer’s.
- **Verifiers** (Read-only agents): Test the reviewer’s claims against the real source. A cropped capture once produced a false “cards touch the edge”.
- **Research** (Google): What the pattern is called, its accessibility contract and the mistakes it is known for, before the brief is written.
- **Taste** (Us): We watch reviews as they run and step in. When our eye and the score disagree, our eye wins and the rule gets written down.

## The score

Eight dimensions, each scored from 0 to 10 and averaged. Accessibility and the export contract are gates, not dimensions.

- **Hierarchy and focus**: Blur the capture. Is exactly one thing dominant, and is it the right one?
- **Composition**: Can you name its one structural idea in a sentence, and does it survive at 360 px?
- **Typography**: A real scale, weight contrast, tracking that tightens with size, a bounded measure.
- **Spacing and rhythm**: Five spacing values or fewer, grouped by proximity, on one left edge.
- **Colour and surfaces**: One neutral ramp, three text tones, hairlines over shadows, one elevation at most.
- **Detail and states**: Concentric radii, a focus ring that fits, hover only where a click lands, every state designed.
- **Responsiveness**: Reflowed rather than scaled. No overflow at 360 px, and targets clear 44 px.
- **Content and honesty**: Copy that reads like a real product, no invented proof, and the default as the best state.

| Score | Band | Meaning |
| --- | --- | --- |
| 1 to 3 | Broken | No hierarchy, failing contrast, overflow. |
| 4 to 5 | Placeholder | Works, but could be any template. |
| 6 to 7 | Competent | Correct, with no idea and no refinement. |
| 8 to 9 | Crafted | A clear idea, refined details. |
| 10 | Nothing to change | A senior designer would not touch it. |

- **Correct is a 6.** No technical flaw and no idea scores at most 6. Correctness is the floor.
- **Copy vacuum caps at 5.** Placeholder headlines or meaningless numbers in any fixture cap the score, whatever the layout.
- **The one-change test.** Before a 10, the reviewer has to try to name one change worth making. If it can, the score is 9.
- **Deltas are earned by pixels.** A round that changed code but not the rendered result does not move the score.

## Real reviews

### [Section wrapper](https://pagesugar.com/components/section-wrapper-01) (Layout · 4 rounds)

Round two came back a perfect 10 that did not square with what its fixes were predicted to be worth, so round three re‑scored the same code adversarially.

| Round | Score | What landed |
| --- | --- | --- |
| 1 | 8.7 | Buttons moved under their copy on the shared edge, facts separated by space instead of a rule, the lede held under 70 characters, Arabic eyebrows given a size step instead of capitals. |
| 2 | 10.0 (thrown out) | Nothing named. Treated as inflated and re‑scored. |
| 3 | 9.3 | Fact values set at 600 over 400 labels; the header’s last link lined up with the content edge. |
| 4 | 9.6 | Nothing left above the threshold. |
| Gate |  | The consistency survey found the wrapper put its gutters inside the container where every other section put them outside. Moved to match. |

### [Product screenshot hero](https://pagesugar.com/components/hero-product-screenshot-01) (Hero · 3 rounds)

At 360 px the product screenshot’s own text could not be read, and the first round capped the score at 6 for it.

| Round | Score | What landed |
| --- | --- | --- |
| 1 | 6.0 | Capped at 6: the screenshot’s text was unreadable at 360 px. |
| 2 |  | Phone and tablet renditions of the screenshot, drawn at the size they display. |
| 3 | 9.9 | Fixture copy that did not match its pictures rewritten. The last fix re‑scored at 10. |

### [Site header](https://pagesugar.com/components/site-header-01) (Navigation · 10 rounds)

The longest review so far, and the one that taught us the most. In round four the reviewer capped the header for being conventional, and we disagreed.

| Round | Score | What landed |
| --- | --- | --- |
| 1, 2 | 5.0 | Held at 5 by the copy-vacuum rule. |
| 3 | 8.6 | Copy fixed, and the cap lifted. |
| 4 | 6.0 (thrown out) | Capped for being a conventional header. A full-width sheet was built to beat the cap, looked wrong, and was reverted. |
| 5 to 10 | 9.5 | 8.4, 9.3, 9.0, 9.4, 9.4, 9.5, judged on the craft of a conventional structure: press states, close timing, tones, the drawer. |

## How long it takes

The median component took 1.5 h from its first commit to its last, 51 of 65 took under two hours and none more than 2.8 h. Where an early component was reworked a day or two later, only its first run of work is counted. Median 7 commits; pull requests were open for a median 44 minutes. The clock starts at the first commit, so the research and build before it are not counted.

Product screenshot hero, commit by commit (minutes after the first):

- +0 min: Built. Component, mock screenshots and tests.
- +13 min: Code review. Hairline and alt text fixes from Codex.
- +54 min: Round 1 fixes. Phone crops drawn at the size they show.
- +77 min: Round 2 fixes. A tablet view between phone and desktop.
- +93 min: Round 3 fixes. The board lists every card its count promises.
- +102 min: Captured. Thumbnail and preview sizes.
- +107 min: Pull request. Opened with the score table.
- +149 min: Merged. Reviewed once more, then merged.

## Getting better

The first ten component pull requests averaged 7.5 in their first round; the last ten averaged 8.5.

| Pull request | First review | Last review | Rounds |
| --- | --- | --- | --- |
| 1 | 5.7 | 9.0 | 4 |
| 2 | 5.0 | 9.5 | 10 |
| 3 | 6.0 | 8.6 | 3 |
| 4 | 8.5 | 9.9 | 3 |
| 5 | 8.3 | 8.5 | 2 |
| 6 | 7.9 | 9.0 | 3 |
| 7 | 7.9 | 8.7 | 3 |
| 8 | 8.3 | 10.0 | 3 |
| 9 | 8.6 | 10.0 | 2 |
| 10 | 8.4 | 9.9 | 3 |
| 11 | 8.9 | 9.6 | 3 |
| 12 | 9.0 | 10.0 | 3 |
| 13 | 9.0 | 10.0 | 3 |
| 14 | 7.5 | 9.9 | 3 |
| 15 | 9.4 | 10.0 | 2 |
| 16 | 8.1 | 10.0 | 5 |
| 17 | 8.7 | 9.6 | 4 |
| 18 | 8.8 | 10.0 | 2 |
| 19 | 9.3 | 9.9 | 3 |
| 20 | 8.9 | 9.9 | 3 |
| 21 | 8.5 | 10.0 | 3 |
| 22 | 8.4 | 9.5 | 4 |
| 23 | 9.0 | 10.0 | 2 |
| 24 | 8.9 | 9.5 | 3 |
| 25 | 8.8 | 10.0 | 3 |
| 26 | 8.9 | 10.0 | 3 |
| 27 | 8.8 | 9.4 | 3 |
| 28 | 9.0 | 10.0 | 3 |
| 29 | 9.1 | 10.0 | 3 |
| 30 | 8.8 | 10.0 | 3 |
| 31 | 8.6 | 9.8 | 3 |
| 32 | 8.8 | 9.8 | 3 |
| 33 | 8.0 | 10.0 | 3 |
| 34 | 8.8 | 9.4 | 3 |
| 35 | 8.9 | 9.9 | 3 |
| 36 | 8.0 | 9.3 | 4 |
| 37 | 9.5 | 10.0 | 3 |
| 38 | 6.0 | 9.9 | 3 |
| 39 | 8.8 | 10.0 | 3 |
| 40 | 8.5 | 9.5 | 4 |
| 41 | 8.6 | 9.4 | 5 |
| 42 | 8.4 | 9.3 | 4 |
| 43 | 8.6 | 10.0 | 3 |
| 44 | 9.4 | 10.0 | 2 |
| 45 | 9.0 | 9.3 | 2 |
| 46 | 8.9 | 10.0 | 2 |
| 47 | 8.6 | 9.9 | 4 |
| 48 | 9.4 | 10.0 | 3 |
| 49 | 9.0 | 9.9 | 3 |
| 50 | 7.9 | 10.0 | 3 |
| 51 | 8.8 | 10.0 | 3 |
| 52 | 8.3 | 9.9 | 3 |
| 53 | 8.0 | 8.9 | 3 |
| 54 | 8.8 | 9.1 | 3 |
| 55 | 9.3 | 9.9 | 3 |
| 56 | 8.8 | 9.9 | 3 |
| 57 | 8.4 | 10.0 | 4 |
| 58 | 7.5 | 9.7 | 3 |
| 59 | 7.8 | 9.2 | 4 |
| 60 | 9.3 | 9.8 | 3 |
| 61 | 8.4 | 10.0 | 3 |
| 62 | 8.9 | 9.9 | 3 |
| 63 | 8.6 | 9.9 | 3 |
| 64 | 8.3 | 9.9 | 3 |
| 65 | 8.9 | 10.0 | 3 |

What the first seven reviews kept finding, and what handles it now:

- Weak default copy in the preview (7 of 7): Content is step one
- Spacing over budget for the section (6 of 7): Linted
- Right-to-left tracking and CJK wrapping (6 of 7): Step five, other languages
- Reviewer pushing an unusual structure (4 of 7): Written into house taste
- Layout moving when state changes (4 of 7): Measured every capture
- Coverage not recorded (3 of 7): Validation fails without it

An agent connected over MCP can report a problem with `submit_component_feedback`. Feedback is screened, triaged weekly and turned into refine issues that run through the same pipeline.

## Where a person decides

- **Conventional structures, polished by craft.** When a reviewer asks for a more distinctive structure, we decline and improve type, spacing, states and copy instead.
- **Keep presence.** Deleting an element to raise a score reads as losing personality. Fix the badge; do not remove it.
- **Open with air.** Heroes and introductions keep their padding, even when a reviewer prices the space.
- **Hover only where a click lands.** Proved with a real hover and a real click in a browser, not by reading the CSS.

## What a score means

- One model’s judgement of real screenshots against a written standard, with its reasoning recorded.
- Its factual claims were checked against the source before anything changed.
- Every fixture and state the component has was captured at three widths and looked at.
- Accessibility and the export contract passed as gates, not traded for looks.

What it does not mean:

- User testing. Nobody watched a visitor use it.
- Perfection. A 10 means the reviewer could not name a change, and reviewers inflate: we have thrown out a 10 that re‑scored 9.3.
- That it fits your site as is. It is a starting point your agent adapts.
- That it has no bugs. If you find one, your agent can tell us through the MCP server.
