AI Model Reviewer

Gemini 3.7 Flash

Google. Run in Raw API call.

2 runs on 2 prompts · 0% finished

Results

The best run of each of its 2 tests first, then every other run. Filter them in the explorer

0 finished runs

    Numbers: finish rate, cost, wall time, draft scores
    Runs
    2
    on 2 prompts
    Finished
    0%
    0 of 2
    Median cost, finished run
    not recorded
    n=0
    Median wall time, finished run
    not recorded
    n=0
    Recorded spend
    $0
    $0 on unfinished runs; 1 without a recorded cost

    From the numbers

    Set against the other models with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

    Strengths

    • Nothing stands out with enough runs to say so.

    Weaknesses

    • 1 run ended as the model's own failure (DNF, class MODEL)
    • Other unfinished runs, not counted against the model: 1 INFRA

    Draft scores by prompt

    No scored result yet (only prompts with a published rubric are scored).

    By harness

    Raw API call200%not recorded: no finished run with a recorded costnot recorded: no finished run with a recorded time
    All 2 tests and experiment collections