AI Model Reviewer

GLM-5.3 (fast router)

Z.ai. Run in OpenCode.

12 runs on 12 prompts · 17% finished

Results

The best run of each of its 12 tests first, then every other run. Filter them in the explorer

2 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
12
on 12 prompts
Finished
17%
2 of 12
Median cost, finished run
$3.11
n=2
Median wall time, finished run
13m 05s
n=2
Recorded spend
$8.65
$2.44 on unfinished runs

From the numbers

Set against the other models with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • Nothing stands out with enough runs to say so.

Weaknesses

  • 2 of 12 runs finished: 25th of 26 models with at least 3 (finish rate)
  • 10 runs ended as the model's own failure (DNF, class MODEL)

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

Hardest run

Picked as its costliest run that did not finish.

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By harness

OpenCode12217%$3.11 n=213m 05s n=2
All 12 tests and experiment collections