AI Model Reviewer

Gemini 3.8 Flash

Google. Run in Cline CLI, OpenCode. No access to its home CLI or through OpenCode (those planned slots show as not tested); tested through the Cline free tier from 2026-09-24.

16 runs on 12 prompts · 50% finished

Results

The best run of each of its 12 tests first, then every other run. Filter them in the explorer

8 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
16
on 12 prompts
Finished
50%
8 of 16
Median cost, finished run
$0
n=3
Median wall time, finished run
9m 44s
n=8
Recorded spend
$0.073
$0.073 on unfinished runs; 10 without a recorded cost

From the numbers

Set against the other models with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • $0 (n=3): 1st of 14 models with at least 3 (median cost of a finished run)

Weaknesses

  • 8 of 16 runs finished: 20th of 24 models with at least 3 (finish rate)
  • 9m 44s (n=8): 17th of 23 models with at least 3 (median wall time of a finished run)
  • 4 runs ended as the model's own failure (DNF, class MODEL)
  • Other unfinished runs, not counted against the model: 3 INFRA, 1 TIMEOUT

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

Hardest run

Picked as its costliest run that did not finish.

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By harness

Cline CLI15853%$0 n=39m 44s n=8
OpenCode100%not recorded: no finished run with a recorded costnot recorded: no finished run with a recorded time
All 12 tests and experiment collections