AI Model Reviewer

Kimi K3 (fast router)

Moonshot AI. Run in OpenCode.

12 runs on 12 prompts · 92% finished

Results

The best run of each of its 12 tests first, then every other run. Filter them in the explorer

11 finished runs

  • Impressionist world

    Kimi K3 (fast router)

    OpenCode 1.18.35

    skills on · OpenCode 1.18.35 + skills on (54: 33 video + 21 web/Three.js)

    Failed · Stopped by us

Numbers: finish rate, cost, wall time, draft scores
Runs
12
on 12 prompts
Finished
92%
11 of 12
Median cost, finished run
$2.51
n=11
Median wall time, finished run
16m 15s
n=11
Recorded spend
$29.33
$3.00 on unfinished runs

From the numbers

Set against the other models with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • Nothing stands out with enough runs to say so.

Weaknesses

  • 16m 15s (n=11): 20th of 25 models with at least 3 (median wall time of a finished run)
  • $2.51 (n=11): 15th of 17 models with at least 3 (median cost of a finished run)
  • Other unfinished runs, not counted against the model: 1 TIMEOUT

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

Hardest run

Picked as its costliest run that did not finish.

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By harness

OpenCode121192%$2.51 n=1116m 15s n=11
All 12 tests and experiment collections