AI Model Reviewer

DeepSeek V4 Pro

DeepSeek. Run in Raw API call. Raw baseline only so far.

171 runs on 170 prompts · 98% finished

Results

The best run of each of its 170 tests first, then every other run. Filter them in the explorer

167 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
171
on 170 prompts
Finished
98%
167 of 171
Median cost, finished run
$0.052
n=1
Median wall time, finished run
2m 23s
n=167
Recorded spend
$0.052
$0 on unfinished runs; 170 without a recorded cost

From the numbers

Set against the other models with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • 167 of 171 runs finished: 7th of 24 models with at least 3 (finish rate)
  • 2m 23s (n=167): 6th of 23 models with at least 3 (median wall time of a finished run)

Weaknesses

  • 4 runs ended as the model's own failure (DNF, class MODEL)

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

Hardest run

Picked as its slowest finished run.

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By harness

Raw API call17116798%$0.052 n=12m 23s n=167
All 170 tests and experiment collections