AI Model Reviewer

Qwen3.8 27B

Alibaba Qwen (via Cerebras). Run in Raw API call.

340 runs on 339 prompts · 44% finished

Results

The best run of each of its 339 tests first, then every other run. Filter them in the explorer

150 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
340
on 339 prompts
Finished
44%
150 of 340
Median cost, finished run
not recorded
n=0
Median wall time, finished run
26 s
n=150
Recorded spend
$0
$0 on unfinished runs; 340 without a recorded cost

From the numbers

Set against the other models with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • 26 s (n=150): 2nd of 23 models with at least 3 (median wall time of a finished run)

Weaknesses

  • 150 of 340 runs finished: 21st of 24 models with at least 3 (finish rate)
  • 190 runs ended as the model's own failure (DNF, class MODEL)

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

  • Qwen3.8 27B's output for Bookshop Logo

    Qwen3.8 27B, raw API call, Bookshop Logo

    Raw API call, default effort

    raw API baseline

    4.4 s

    not recorded: the run's token counts were never reported, so there is no estimate

    OK

    Unscored: no rubric for this prompt

Hardest run

Picked as its slowest finished run.

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By harness

Raw API call34015044%not recorded: no finished run with a recorded cost26 s n=150
All 339 tests and experiment collections