AI Model Reviewer

Raw API call

22 models ran in it: Claude Fable 5.1, Claude Haiku 5.5, Claude Opus 5.5, Claude Sonnet 5, Claude Sonnet 5.5, DeepSeek V4 Pro, DeepSeek V4.1 Flash, GLM-5.3, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol, Gemini 3.6 Flash, Gemini 3.7 Flash, Kimi K3, MiniMax M3, Qwen3.8 27B, Qwen3.8 Max, gemini-3.5-flash, gpt-6.1-sol, gpt-oss-120b, minimax-m3, qwen3p8-max.

2,097 runs on 375 prompts · 83% finished

Results

The best run of each of its 375 tests first, then every other run. Filter them in the explorer

1,744 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
2097
on 375 prompts
Finished
83%
1744 of 2097
Median cost, finished run
$0.59
n=371
Median wall time, finished run
1m 34s
n=1744
Recorded spend
$274.74
$38.77 on unfinished runs; 1696 without a recorded cost

From the numbers

Set against the other harnesses with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • 1m 34s (n=1744): 1st of 12 harnesses with at least 3 (median wall time of a finished run)
  • $0.59 (n=371): 3rd of 12 harnesses with at least 3 (median cost of a finished run)

Weaknesses

  • 348 runs ended as the model's own failure (DNF, class MODEL)
  • Other unfinished runs, not counted against the model: 1 INFRA, 2 TIMEOUT, 1 STOPPED, 1 ERROR

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

  • gpt-oss-120b's output for Settings Toggles

    gpt-oss-120b, raw API call, Settings Toggles

    Raw API call, high effort, via OpenAI

    raw API baseline

    2.3 s

    not recorded: the run's token counts were never reported, so there is no estimate

    OK

    Unscored: no rubric for this prompt

Hardest run

Picked as its costliest run that did not finish.

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By model

Show the table (20 rows)
GPT-6 Sol139139100%$0.44 n=51m 54s n=139
GPT-6 Luna134134100%not recorded: no finished run with a recorded cost1m 10s n=134
gpt-6.1-sol1111100%not recorded: no finished run with a recorded cost3m 35s n=11
Claude Haiku 5.522100%not recorded: no finished run with a recorded cost3m 08s n=2
Claude Sonnet 5999798%not recorded: no finished run with a recorded cost2m 17s n=97
GPT-6 Astra14714498%$0.61 n=1423m 46s n=144
DeepSeek V4 Pro17116798%$0.052 n=12m 23s n=167
Claude Fable 5.111110897%$0.72 n=1082m 53s n=108
DeepSeek V4.1 Flash16615795%$0.035 n=11m 02s n=157
Kimi K3716592%not recorded: no finished run with a recorded cost9m 10s n=65
gpt-oss-120b35432492%not recorded: no finished run with a recorded cost8.5 s n=324
Claude Opus 5.514111984%$0.40 n=1113m 28s n=119
Qwen3.8 Max715983%not recorded: no finished run with a recorded cost3m 56s n=59
MiniMax M3644875%not recorded: no finished run with a recorded cost2m 39s n=48
Claude Sonnet 5.54250%not recorded: no finished run with a recorded cost4m 30s n=2
gemini-3.5-flash2150%not recorded: no finished run with a recorded cost2m 15s n=1
Qwen3.8 27B34015044%not recorded: no finished run with a recorded cost26 s n=150
Gemini 3.6 Flash11327%$0 n=33m 12s n=3
GLM-5.3571425%not recorded: no finished run with a recorded cost11m 22s n=14
Gemini 3.7 Flash200%not recorded: no finished run with a recorded costnot recorded: no finished run with a recorded time
All 375 tests