AI Model Reviewer

Claude Sonnet 5

Anthropic. Run in Raw API call.

99 runs on 99 prompts · 98% finished

Results

The best run of each of its 99 tests first, then every other run. Filter them in the explorer

97 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
99
on 99 prompts
Finished
98%
97 of 99
Median cost, finished run
not recorded
n=0
Median wall time, finished run
2m 17s
n=97
Recorded spend
$0
$0 on unfinished runs; 99 without a recorded cost

From the numbers

Set against the other models with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • 97 of 99 runs finished: 6th of 24 models with at least 3 (finish rate)
  • 2m 17s (n=97): 5th of 23 models with at least 3 (median wall time of a finished run)

Weaknesses

  • 2 runs ended as the model's own failure (DNF, class MODEL)

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

Hardest run

Picked as its slowest finished run.

  • Claude Sonnet 5's output for Grand Piano

    Claude Sonnet 5, raw API call, Grand Piano

    Raw API call, high effort, via Anthropic

    raw API baseline

    7m 22s

    not recorded: the run's token counts were never reported, so there is no estimate

    OK

    Unscored: no rubric for this prompt

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By harness

Raw API call999798%not recorded: no finished run with a recorded cost2m 17s n=97
All 99 tests and experiment collections