About
About & data
How the runs were made and checked, what they cost, the raw data, and the older pages.
- MethodologyHow runs were done, labelled and scored; the rubrics and field notes.
- CostsSpend per model, harness and category; reported vs estimated.
- Model speedTime to first token and output speed.
- Data & downloadsresults.json and costs.json, the files behind every page.
- CompareEach prompt: every model in its home harness, the neutral harness and the raw API.
- LeaderboardsThe AI judge's provisional ranks within each scored prompt.
- GridEvery model in every harness, per prompt.
- Best pick for your jobThe pairings that did best on the recorded runs.
- Cost vs qualityCost against draft score, finish rate and wall time.
- ExperimentsThe earlier experiment view: per-model galleries per category.
- Subject categoriesRuns grouped by subject (animals, vehicles, ...).
- Harness research noteWhich harness each model family should run in.
- Legacy SVGs (older benchmark)1,113 SVGs from an older, separate benchmark: different prompts, never scored, not our runs.
- Devlog: how this site was builtEvery build session, logged.