AI Model Reviewer

Cline CLI

5 models ran in it: DeepSeek V4.1 Flash, Gemini 3.8 Flash, MiMo V2.6 Flash, Muse Spark 1.3, Space Bunny Alpha.

67 runs on 14 prompts · 52% finished

Results

The best run of each of its 14 tests first, then every other run. Filter them in the explorer

35 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
67
on 14 prompts
Finished
52%
35 of 67
Median cost, finished run
$0
n=23
Median wall time, finished run
5m 34s
n=35
Recorded spend
$0
$0 on unfinished runs; 42 without a recorded cost

From the numbers

Set against the other harnesses with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • 5m 34s (n=35): 2nd of 12 harnesses with at least 3 (median wall time of a finished run)
  • $0 (n=23): 1st of 12 harnesses with at least 3 (median cost of a finished run)

Weaknesses

  • 35 of 67 runs finished: 12th of 14 harnesses with at least 3 (finish rate)
  • 4 runs ended as the model's own failure (DNF, class MODEL)
  • Other unfinished runs, not counted against the model: 9 HARNESS, 12 INFRA, 7 TIMEOUT

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

Hardest run

Picked as its slowest finished run.

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By model

Muse Spark 1.31212100%$0 n=74m 35s n=12
Space Bunny Alpha1111100%$0 n=112m 26s n=11
Gemini 3.8 Flash15853%$0 n=39m 44s n=8
DeepSeek V4.1 Flash13215%not recorded: no finished run with a recorded cost14m 35s n=2
MiMo V2.6 Flash16212%$0 n=249m 26s n=2
All 14 tests