AI Model Reviewer

gpt-6.1-sol

OpenAI. Run in Codex CLI, Raw API call.

31 runs on 31 prompts · 94% finished

Results

The best run of each of its 31 tests first, then every other run. Filter them in the explorer

29 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
31
on 31 prompts
Finished
94%
29 of 31
Median cost, finished run
not recorded
n=0
Median wall time, finished run
19m 46s
n=29
Recorded spend
$0
$0 on unfinished runs; 31 without a recorded cost

From the numbers

Set against the other models with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • Nothing stands out with enough runs to say so.

Weaknesses

  • 19m 46s (n=29): 21st of 23 models with at least 3 (median wall time of a finished run)
  • Other unfinished runs, not counted against the model: 2 TIMEOUT

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

Hardest run

Picked as its slowest finished run.

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By harness

Raw API call1111100%not recorded: no finished run with a recorded cost3m 35s n=11
Codex CLI201890%not recorded: no finished run with a recorded cost33m 31s n=18
All 31 tests and experiment collections