AI Model Reviewer

gpt-oss-120b

OpenAI open-weight (via Cerebras). Run in Raw API call.

354 runs on 341 prompts · 92% finished

Results

The best run of each of its 341 tests first, then every other run. Filter them in the explorer

324 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
354
on 341 prompts
Finished
92%
324 of 354
Median cost, finished run
not recorded
n=0
Median wall time, finished run
8.5 s
n=324
Recorded spend
$0
$0 on unfinished runs; 354 without a recorded cost

From the numbers

Set against the other models with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • 8.5 s (n=324): 1st of 23 models with at least 3 (median wall time of a finished run)

Weaknesses

  • 30 runs ended as the model's own failure (DNF, class MODEL)

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

  • gpt-oss-120b's output for Settings Toggles

    gpt-oss-120b, raw API call, Settings Toggles

    Raw API call, high effort, via OpenAI

    raw API baseline

    2.3 s

    not recorded: the run's token counts were never reported, so there is no estimate

    OK

    Unscored: no rubric for this prompt

Hardest run

Picked as its slowest finished run.

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By harness

Raw API call35432492%not recorded: no finished run with a recorded cost8.5 s n=324
All 341 tests and experiment collections