AI Model Reviewer

Claude Opus 5.5

Anthropic. Run in Claude Code, Hermes Agent 0.21.5, Hermes Agent 0.21.5 + Oh-My-Hermes 2.0.5, Kilo Code CLI 7.7.9, Oh My Pi (omp) 18.3.1, OpenCode, Pi 0.87.1, Pi 0.87.1 + skills kit (ponytail, ECC, vetted skills), Raw API call.

355 runs on 229 prompts · 80% finished

Results

The best run of each of its 229 tests first, then every other run. Filter them in the explorer

283 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
355
on 229 prompts
Finished
80%
283 of 355
Median cost, finished run
$2.37
n=275
Median wall time, finished run
14m 39s
n=283
Recorded spend
$1,862.53
$262.79 on unfinished runs; 30 without a recorded cost

From the numbers

Set against the other models with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • Draft score 9.5/10 on Pelican riding a bicycle (control) (in Claude Code): ranked 1st of 13
  • Draft score 9.5/10 on MacBook Pro 16" realistic 3D render (Playcode benchmark) (in Claude Code): ranked 1st of 12
  • Draft score 10/10 on Animated pelican riding a bicycle (SMIL/CSS) (in Claude Code): ranked 1st of 12
  • Draft score 8.5/10 on Lighthouse on a headland at night (in OpenCode): ranked 2nd of 12
  • Draft score 9.5/10 on Animated pelican riding a bicycle (SMIL/CSS) (in OpenCode): ranked 2nd of 12

Weaknesses

  • 283 of 355 runs finished: 17th of 24 models with at least 3 (finish rate)
  • 14m 39s (n=283): 19th of 23 models with at least 3 (median wall time of a finished run)
  • $2.37 (n=275): 12th of 14 models with at least 3 (median cost of a finished run)
  • 27 runs ended as the model's own failure (DNF, class MODEL)
  • Other unfinished runs, not counted against the model: 3 HARNESS, 5 INFRA, 34 TIMEOUT, 3 STOPPED

Best and hardest run

Best run

Picked as its best draft-score rank on a prompt.

Hardest run

Picked as its costliest run that did not finish.

Draft scores by prompt

Lighthouse on a headland at nightOpenCode8.5/10 draft2nd of 12
MacBook Pro 16" realistic 3D render (Playcode benchmark)Claude Code9.5/10 draft1st of 12
Pelican riding a bicycle (control)Claude Code9.5/10 draft1st of 13
Pelican riding a bicycle (control)OpenCode9/10 draft6th of 13
Animated pelican riding a bicycle (SMIL/CSS)Claude Code10/10 draft1st of 12
Animated pelican riding a bicycle (SMIL/CSS)OpenCode9.5/10 draft2nd of 12

By harness

Claude Code12411190%$4.16 n=11128m 20s n=111
Pi 0.87.18788%$4.89 n=729m 04s n=7
Oh My Pi (omp) 18.3.17686%$7.36 n=633m 25s n=6
Raw API call14111984%$0.40 n=1113m 28s n=119
Kilo Code CLI 7.7.96583%$6.83 n=533m 50s n=5
OpenCode352571%$15.39 n=251h 04m n=25
Pi 0.87.1 + skills kit (ponytail, ECC, vetted skills)16956%$2.75 n=916m 06s n=9
Hermes Agent 0.21.510110%$23.83 n=12h 30m n=1
Hermes Agent 0.21.5 + Oh-My-Hermes 2.0.5800%not recorded: no finished run with a recorded costnot recorded: no finished run with a recorded time
All 229 tests and experiment collections