AI Model Reviewer

GPT-6 Luna

OpenAI. Run in Raw API call.

134 runs on 134 prompts · 100% finished

Results

The best run of each of its 134 tests first, then every other run. Filter them in the explorer

134 finished runs

  • GPT-6 Luna, Raw API (one-shot): 016 Heartbeat LineLive demo

    016 Heartbeat Line

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): 013 Soft Body BlobsLive demo

    013 Soft Body Blobs

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Koi Pond GardenLive demo

    Koi Pond Garden

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): 011 Pure CSS SunsetLive demo

    011 Pure CSS Sunset

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): 009 Forest Fire AutomatonLive demo

    009 Forest Fire Automaton

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): 008 Pocket ThereminLive demo

    008 Pocket Theremin

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): 002 Pendulum WaveLive demo

    002 Pendulum Wave

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): 003 Orbital Coffee LandingLive demo

    003 Orbital Coffee Landing

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Lavender Fields

    Lavender Fields

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Latte Art

    Latte Art

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Ladybug on a Leaf

    Ladybug on a Leaf

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Knight in Armour

    Knight in Armour

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Kingfisher Diving

    Kingfisher Diving

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Kids Building a Snowman

    Kids Building a Snowman

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Ring of Keys

    Ring of Keys

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Seigaiha Wave Pattern

    Seigaiha Wave Pattern

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Five-Storey Pagoda

    Five-Storey Pagoda

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Jack-o'-Lanterns

    Jack-o'-Lanterns

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Isometric Space Station

    Isometric Space Station

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Isometric Workspace

    Isometric Workspace

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Isometric Harbour

    Isometric Harbour

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Isometric Farm

    Isometric Farm

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Isometric Cube Illusion

    Isometric Cube Illusion

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

  • GPT-6 Luna, Raw API (one-shot): Isometric Coffee Shop

    Isometric Coffee Shop

    GPT-6 Luna

    Raw API (one-shot)

    raw API baseline

Numbers: finish rate, cost, wall time, draft scores
Runs
134
on 134 prompts
Finished
100%
134 of 134
Median cost, finished run
not recorded
n=0
Median wall time, finished run
1m 10s
n=134
Recorded spend
$0
$0 on unfinished runs; 134 without a recorded cost

From the numbers

Set against the other models with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • 134 of 134 runs finished: 1st of 24 models with at least 3 (finish rate)
  • 1m 10s (n=134): 3rd of 23 models with at least 3 (median wall time of a finished run)

Weaknesses

  • Nothing stands out with enough runs to say so.

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

  • GPT-6 Luna's output for File Upload Dropzone

    GPT-6 Luna, raw API call, File Upload Dropzone

    Raw API call, high effort, via OpenAI

    raw API baseline

    24 s

    not recorded: the run's token counts were never reported, so there is no estimate

    OK

    Unscored: no rubric for this prompt

Hardest run

Picked as its slowest finished run.

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By harness

Raw API call134134100%not recorded: no finished run with a recorded cost1m 10s n=134
All 134 tests and experiment collections