AI Model Reviewer

Claude Haiku 5.5

Anthropic. Run in Claude Code, Raw API call. One-shot API tests only so far (v5.4.19-api).

7 runs on 7 prompts · 100% finished

Results

The best run of each of its 7 tests first, then every other run. Filter them in the explorer

7 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
7
on 7 prompts
Finished
100%
7 of 7
Median cost, finished run
not recorded
n=0
Median wall time, finished run
7m 56s
n=7
Recorded spend
$0
$0 on unfinished runs; 7 without a recorded cost

From the numbers

Set against the other models with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • 7 of 7 runs finished: 1st of 24 models with at least 3 (finish rate)

Weaknesses

  • Nothing stands out with enough runs to say so.

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

Hardest run

Picked as its slowest finished run.

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By harness

Claude Code55100%not recorded: no finished run with a recorded cost8m 07s n=5
Raw API call22100%not recorded: no finished run with a recorded cost3m 08s n=2
All 7 tests and experiment collections