AI Model Reviewer

Claude Sonnet 5.5

Anthropic. Run in Claude Code, OpenCode, Raw API call. Sonnet 5.5 sweep (2026-09-28): Claude Code and OpenCode, effort max, frozen MR prompts.

84 runs on 69 prompts · 64% finished

Results

The best run of each of its 69 tests first, then every other run. Filter them in the explorer

54 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
84
on 69 prompts
Finished
64%
54 of 84
Median cost, finished run
$6.65
n=10
Median wall time, finished run
26m 03s
n=54
Recorded spend
$177.92
$117.96 on unfinished runs; 54 without a recorded cost

From the numbers

Set against the other models with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • Nothing stands out with enough runs to say so.

Weaknesses

  • 54 of 84 runs finished: 19th of 24 models with at least 3 (finish rate)
  • 26m 03s (n=54): 22nd of 23 models with at least 3 (median wall time of a finished run)
  • $6.65 (n=10): 14th of 14 models with at least 3 (median cost of a finished run)
  • 2 runs ended as the model's own failure (DNF, class MODEL)
  • Other unfinished runs, not counted against the model: 1 HARNESS, 3 INFRA, 18 TIMEOUT, 6 ERROR

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

Hardest run

Picked as its costliest run that did not finish.

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By harness

Claude Code624877%$6.65 n=625m 20s n=48
Raw API call4250%not recorded: no finished run with a recorded cost4m 30s n=2
OpenCode18422%$6.56 n=437m 09s n=4
All 69 tests and experiment collections