AI Model Reviewer

How this site was built

This site, its data pipeline and its checks were written by Claude Code (Claude Opus 5.5 at max effort), running unattended from written briefs.

The setup

Each session starts from a written brief (what the site must show, the data contract, the honesty rules and the checks), with Claude Code at --effort max in a terminal session that nobody answers. The agent cannot ask questions, so it makes every call itself; the project team only sends short notes, and each session's summary below says which.

Claude Code wrote ingest.py, which reads the results drop folder (one folder per run with meta.json, cost.json, the artifact, a poster and a transcript), checks each folder against the contract, links the media in (hardlinks, never copies), makes thumbnails and timelines, and merges each new run into data/*.json before these pages are rendered from it. Runs recovered after a machine loss keep their records; re-running the build with no new data changes nothing. Each session then loads every page type in headless browsers at phone and desktop sizes, collects console errors and failed requests, looks at the screenshots and fixes what is wrong.

The site itself is plain HTML and CSS with one small script, rendered ahead of time; there is no framework and nothing to build when you view it.

Dev sessions

Read from the dev-session logs. Headless sessions carry Claude Code's own cost total. Interactive sessions do not write one, so their cost is estimated from the token usage in the transcript and labelled that way.

session01: Session 01: build v1 of the site

Done

Date
2026-09-23
Duration
1h 24m
Turns
129 model calls; 6 human/operator messages
Cost
$20.41 estimated from token usage
Model
claude-opus-5-5, max effort
Harness
Claude Code 2.1.280 (interactive TUI in tmux, unattended)

Built v1 from the brief: the results contract, 14 verified seed results from the raw API material, the build script (validation, media, thumbnails, transcript excerpts, results, costs, speed, harness and build-log data), all eight page types, and headless-Chromium verification at desktop and phone sizes. Folded in the project lead's two mid-session additions (Remotion and HyperFrames categories; the harness-config-failure status). One turn died on a gateway timeout (HTTP 524) caused by our own Claude Code watchdog settings; the team fixed the settings and resumed the session.

Estimated from the token usage recorded in the transcript x Opus 5.5 list price (input $4, cache write $5 or $8 for 1-hour, cache read $0.20, output $20 per 1M). Interactive Claude Code does not write its own cost total to the transcript; any long-context surcharge is not included.

Ended before the machine move: the last transcript entry is at 21:07 UTC and the last commit is 0dccb62. A phone verification batch it had started finished at about 21:09 UTC with zero console errors; the desktop batch had passed at 21:04.

Transcript excerpt of session01

session02: Session 02: honesty pass on scores, new results, field notes, content sources

Done

Date
2026-09-23
Duration
38m 13s
Turns
115 model calls; 4 human/operator messages
Cost
$12.95 estimated from token usage
Model
claude-opus-5-5 [1m], max effort
Harness
Claude Code 2.1.280 (interactive TUI in tmux, unattended)

Estimated from the token usage recorded in the transcript x Opus 5.5 list price (input $4, cache write $5 or $8 for 1-hour, cache read $0.20, output $20 per 1M). Interactive Claude Code does not write its own cost total to the transcript; any long-context surcharge is not included.

Transcript excerpt of session02