AI Model Reviewer

GPT-6 Sol

OpenAI. Run in Codex CLI, Raw API call. Run in Codex CLI (its home harness) from 2026-09-25.

238 runs on 234 prompts · 100% finished

Results

The best run of each of its 234 tests first, then every other run. Filter them in the explorer

238 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
238
on 234 prompts
Finished
100%
238 of 238
Median cost, finished run
$1.31
n=104
Median wall time, finished run
2m 25s
n=238
Recorded spend
$153.40
$0 on unfinished runs; 134 without a recorded cost

From the numbers

Set against the other models with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • 238 of 238 runs finished: 1st of 24 models with at least 3 (finish rate)
  • 2m 25s (n=238): 7th of 23 models with at least 3 (median wall time of a finished run)

Weaknesses

  • $1.31 (n=104): 11th of 14 models with at least 3 (median cost of a finished run)

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

  • GPT-6 Sol's output for Fruit Sales Bar Chart

    GPT-6 Sol, raw API call, Fruit Sales Bar Chart

    Raw API call, high effort, via OpenAI

    raw API baseline

    30 s

    not recorded: the run's token counts were never reported, so there is no estimate

    OK

    Unscored: no rubric for this prompt

Hardest run

Picked as its slowest finished run.

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By harness

Raw API call139139100%$0.44 n=51m 54s n=139
Codex CLI9999100%$1.36 n=999m 44s n=99
All 234 tests and experiment collections