AI Model Reviewer

Codex CLI

3 models ran in it: GPT-6 Astra, GPT-6 Sol, gpt-6.1-sol.

186 runs on 136 prompts · 99% finished

Results

The best run of each of its 136 tests first, then every other run. Filter them in the explorer

184 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
186
on 136 prompts
Finished
99%
184 of 186
Median cost, finished run
$1.80
n=166
Median wall time, finished run
14m 29s
n=184
Recorded spend
$601.04
$0 on unfinished runs; 20 without a recorded cost

From the numbers

Set against the other harnesses with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • 184 of 186 runs finished: 2nd of 14 harnesses with at least 3 (finish rate)
  • 14m 29s (n=184): 3rd of 12 harnesses with at least 3 (median wall time of a finished run)
  • Draft score 9.5/10 on Pelican riding a bicycle (control) (GPT-6 Astra): ranked 1st of 13
  • Draft score 9.5/10 on Isometric SVG town with sprite sheet (GPT-6 Astra): ranked 2nd of 11

Weaknesses

  • Other unfinished runs, not counted against the model: 2 TIMEOUT

Best and hardest run

Best run

Picked as its best draft-score rank on a prompt.

Hardest run

Picked as its lowest draft-score rank on a prompt.

Draft scores by prompt

Isometric SVG town with sprite sheetGPT-6 Astra9.5/10 draft2nd of 11
Lighthouse on a headland at nightGPT-6 Astra8/10 draft6th of 12
MacBook Pro 16" realistic 3D render (Playcode benchmark)GPT-6 Astra8.5/10 draft6th of 12
Pelican riding a bicycle (control)GPT-6 Astra9.5/10 draft1st of 13
Animated pelican riding a bicycle (SMIL/CSS)GPT-6 Astra9/10 draft6th of 12

By model

GPT-6 Sol9999100%$1.36 n=999m 44s n=99
GPT-6 Astra6767100%$5.99 n=6725m 22s n=67
gpt-6.1-sol201890%not recorded: no finished run with a recorded cost33m 31s n=18
All 136 tests