AI Model Reviewer

dsh

1 model ran in it: DeepSeek V4.1 Flash.

22 runs on 21 prompts · 82% finished

Results

The best run of each of its 21 tests first, then every other run. Filter them in the explorer

18 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
22
on 21 prompts
Finished
82%
18 of 22
Median cost, finished run
$0.14
n=18
Median wall time, finished run
16m 40s
n=18
Recorded spend
$3.06
$0.23 on unfinished runs; 1 without a recorded cost

From the numbers

Set against the other harnesses with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • $0.14 (n=18): 2nd of 12 harnesses with at least 3 (median cost of a finished run)
  • Draft score 9.5/10 on Isometric SVG town with sprite sheet (DeepSeek V4.1 Flash): ranked 2nd of 11

Weaknesses

  • Draft score 7/10 on MacBook Pro 16" realistic 3D render (Playcode benchmark) (DeepSeek V4.1 Flash): ranked 10th of 12
  • Draft score 7/10 on Pelican riding a bicycle (control) (DeepSeek V4.1 Flash): ranked 12th of 13
  • Other unfinished runs, not counted against the model: 4 INFRA

Best and hardest run

Best run

Picked as its best draft-score rank on a prompt.

Hardest run

Picked as its costliest run that did not finish.

Draft scores by prompt

Isometric SVG town with sprite sheetDeepSeek V4.1 Flash9.5/10 draft2nd of 11
Lighthouse on a headland at nightDeepSeek V4.1 Flash8/10 draft6th of 12
MacBook Pro 16" realistic 3D render (Playcode benchmark)DeepSeek V4.1 Flash7/10 draft10th of 12
Pelican riding a bicycle (control)DeepSeek V4.1 Flash7/10 draft12th of 13
Animated pelican riding a bicycle (SMIL/CSS)DeepSeek V4.1 Flash9/10 draft6th of 12

By model

DeepSeek V4.1 Flash221882%$0.14 n=1816m 40s n=18
All 21 tests