AI Model Reviewer

Gemini 3.6 Flash

Google. Run in Raw API call.

11 runs on 11 prompts · 27% finished

Results

The best run of each of its 11 tests first, then every other run. Filter them in the explorer

3 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
11
on 11 prompts
Finished
27%
3 of 11
Median cost, finished run
$0
n=3
Median wall time, finished run
3m 12s
n=3
Recorded spend
$0
$0 on unfinished runs

From the numbers

Set against the other models with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • $0 (n=3): 1st of 14 models with at least 3 (median cost of a finished run)

Weaknesses

  • 3 of 11 runs finished: 23rd of 24 models with at least 3 (finish rate)
  • 7 runs ended as the model's own failure (DNF, class MODEL)
  • Other unfinished runs, not counted against the model: 1 ERROR

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

Hardest run

Picked as its slowest finished run.

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By harness

Raw API call11327%$0 n=33m 12s n=3
All 11 tests and experiment collections