Best pick for your job
Choose the kind of work and what matters most. The picks are the model and harness pairings that did best on the runs recorded here, each with the evidence and the number of runs behind it. Picks are calculated from recorded runs, including costs labeled as estimates where applicable, and a pick with few runs says so.
SVG: cheapest that finishes
At least half of its runs finished; ranked by the median cost of a finished run.
31 pairings qualify; the top 3 are shown.
-
1
Muse Spark 1.3 in Cline CLI
- Finished 12 of 12 (100%)
- Median cost, finished run $0 n=7
- Median wall time, finished run 4m 35s n=12
-
2
Space Bunny Alpha in OpenCode
- Finished 12 of 12 (100%)
- Median cost, finished run $0 n=11
- Median wall time, finished run 7m 19s n=12
-
3
Space Bunny Alpha in Cline CLI
- Finished 9 of 9 (100%)
- Median cost, finished run $0 n=9
- Median wall time, finished run 2m 03s n=9
SVG: best draft score
Ranked by how often it placed in the top third of a scored prompt, then by its best rank there. Ranks come from single prompts; scores are never averaged across prompts.
13 pairings qualify; the top 3 are shown.
-
1
Claude Fable 5.1 in OpenCode
- Finished 12 of 12 (100%)
- Median cost, finished run $6.58 n=12
- Median wall time, finished run 22m 19s n=12
- Ranks on scored prompts 1 of 13, 1 of 12, 1 of 11, 2 of 12, 2 of 12
-
2
Claude Opus 5.5 in Claude Code
- Finished 15 of 15 (100%)
- Median cost, finished run $7.14 n=15
- Median wall time, finished run 39m 14s n=15
- Ranks on scored prompts 1 of 13, 1 of 12, 1 of 12
-
3
Grok 4.7 in Grok Build
- Finished 4 of 4 (100%)
- Median cost, finished run $4.50 n=4
- Median wall time, finished run 38m 25s n=4
- Ranks on scored prompts 1 of 13, 2 of 12, 2 of 11, 6 of 12
SVG: fastest
At least half of its runs finished; ranked by the median wall time of a finished run.
37 pairings qualify; the top 3 are shown.
-
1
gpt-oss-120b in Raw API call
- Finished 225 of 240 (94%)
- Median cost, finished run not recorded
- Median wall time, finished run 8.0 s n=225
-
2
DeepSeek V4.1 Flash in Raw API call
- Finished 132 of 135 (98%)
- Median cost, finished run $0.035 n=1
- Median wall time, finished run 58 s n=132
-
3
GPT-6 Luna in Raw API call
- Finished 113 of 113 (100%)
- Median cost, finished run not recorded
- Median wall time, finished run 1m 02s n=113
SVG: free only
Only its runs that cost nothing although they used tokens (free routes); ranked by finish rate, then the median wall time of a finished run.
8 pairings qualify; the top 3 are shown.
-
1
Space Bunny Alpha in Cline CLI
- Finished 9 of 9 (100%)
- Median cost, finished run $0 n=9
- Median wall time, finished run 2m 03s n=9
-
2
Muse Spark 1.3 in Cline CLI
- Finished 7 of 7 (100%)
- Median cost, finished run $0 n=7
- Median wall time, finished run 5m 20s n=7
-
3
Muse Spark 1.3 in OpenCode
- Finished 11 of 11 (100%)
- Median cost, finished run $0 n=11
- Median wall time, finished run 5m 38s n=11
Three.js: cheapest that finishes
At least half of its runs finished; ranked by the median cost of a finished run.
17 pairings qualify; the top 3 are shown.
-
1
DeepSeek V4.1 Flash in OpenCode
- Finished 9 of 11 (82%)
- Median cost, finished run $0.15 n=9
- Median wall time, finished run 28m 46s n=9
-
2
DeepSeek V4.1 Flash in dsh
- Finished 5 of 7 (71%)
- Median cost, finished run $0.19 n=5
- Median wall time, finished run 45m 32s n=5
-
3
GPT-6 Sol in Raw API call
- Finished 3 of 3 (100%)
- Median cost, finished run $0.56 n=1
- Median wall time, finished run 2m 54s n=3
Three.js: best draft score
Ranked by how often it placed in the top third of a scored prompt, then by its best rank there. Ranks come from single prompts; scores are never averaged across prompts.
No scored prompt in this category yet (only prompts with a published rubric are scored).
Three.js: fastest
At least half of its runs finished; ranked by the median wall time of a finished run.
29 pairings qualify; the top 3 are shown.
-
1No render
gpt-oss-120b in Raw API call
- Finished 4 of 4 (100%)
- Median cost, finished run not recorded
- Median wall time, finished run 14 s n=4
-
2
GPT-6 Sol in Raw API call
- Finished 3 of 3 (100%)
- Median cost, finished run $0.56 n=1
- Median wall time, finished run 2m 54s n=3
-
3
minimax-m3 in Raw API call
- Finished 7 of 8 (88%)
- Median cost, finished run not recorded
- Median wall time, finished run 3m 37s n=7
Three.js: free only
Only its runs that cost nothing although they used tokens (free routes); ranked by finish rate, then the median wall time of a finished run.
2 pairings qualify; the top 2 are shown.
-
1No render
Gemini 3.6 Flash in Raw API call thin evidence: 2 runs
- Finished 2 of 2 (100%)
- Median cost, finished run $0 n=2
- Median wall time, finished run 2m 41s n=2
-
2
Space Bunny Alpha in Cline CLI thin evidence: 1 run
- Finished 1 of 1 (100%)
- Median cost, finished run $0 n=1
- Median wall time, finished run 4m 55s n=1
HTML animation: cheapest that finishes
At least half of its runs finished; ranked by the median cost of a finished run.
10 pairings qualify; the top 3 are shown.
-
1
Claude Opus 5.5 in Raw API call
- Finished 21 of 25 (84%)
- Median cost, finished run $0.37 n=21
- Median wall time, finished run 2m 42s n=21
-
2
GPT-6 Sol in Raw API call
- Finished 21 of 21 (100%)
- Median cost, finished run $0.59 n=1
- Median wall time, finished run 2m 33s n=21
-
3
GPT-6 Sol in Codex CLI
- Finished 5 of 5 (100%)
- Median cost, finished run $0.64 n=5
- Median wall time, finished run 8m 09s n=5
HTML animation: best draft score
Ranked by how often it placed in the top third of a scored prompt, then by its best rank there. Ranks come from single prompts; scores are never averaged across prompts.
No scored prompt in this category yet (only prompts with a published rubric are scored).
HTML animation: fastest
At least half of its runs finished; ranked by the median wall time of a finished run.
19 pairings qualify; the top 3 are shown.
-
1No render
gpt-oss-120b in Raw API call
- Finished 95 of 110 (86%)
- Median cost, finished run not recorded
- Median wall time, finished run 9.8 s n=95
-
2No render
DeepSeek V4.1 Flash in Raw API call
- Finished 22 of 26 (85%)
- Median cost, finished run not recorded
- Median wall time, finished run 1m 41s n=22
-
3No render
GPT-6 Luna in Raw API call
- Finished 19 of 19 (100%)
- Median cost, finished run not recorded
- Median wall time, finished run 1m 53s n=19
HTML animation: free only
Only its runs that cost nothing although they used tokens (free routes); ranked by finish rate, then the median wall time of a finished run.
1 pairing qualify; the top 1 are shown.
-
1
Space Bunny Alpha in Cline CLI thin evidence: 1 run
- Finished 1 of 1 (100%)
- Median cost, finished run $0 n=1
- Median wall time, finished run 9m 22s n=1
Blender: cheapest that finishes
At least half of its runs finished; ranked by the median cost of a finished run.
8 pairings qualify; the top 3 are shown.
-
1
DeepSeek V4.1 Flash in OpenCode
- Finished 3 of 4 (75%)
- Median cost, finished run $0.15 n=3
- Median wall time, finished run 1h 20m n=3
-
2
Claude Opus 5.5 in Claude Code
- Finished 3 of 4 (75%)
- Median cost, finished run $8.85 n=3
- Median wall time, finished run 1h 22m n=3
-
3
GPT-6 Astra in Codex CLI
- Finished 10 of 10 (100%)
- Median cost, finished run $10.85 n=10
- Median wall time, finished run 37m 21s n=10
Blender: best draft score
Ranked by how often it placed in the top third of a scored prompt, then by its best rank there. Ranks come from single prompts; scores are never averaged across prompts.
No scored prompt in this category yet (only prompts with a published rubric are scored).
Blender: fastest
At least half of its runs finished; ranked by the median wall time of a finished run.
10 pairings qualify; the top 3 are shown.
-
1
GPT-6 Astra in Codex CLI
- Finished 10 of 10 (100%)
- Median cost, finished run $10.85 n=10
- Median wall time, finished run 37m 21s n=10
-
2
DeepSeek V4.1 Flash in OpenCode
- Finished 3 of 4 (75%)
- Median cost, finished run $0.15 n=3
- Median wall time, finished run 1h 20m n=3
-
3
Claude Opus 5.5 in Claude Code
- Finished 3 of 4 (75%)
- Median cost, finished run $8.85 n=3
- Median wall time, finished run 1h 22m n=3
Blender: free only
Only its runs that cost nothing although they used tokens (free routes); ranked by finish rate, then the median wall time of a finished run.
No free run in this category finished.
Video: cheapest that finishes
At least half of its runs finished; ranked by the median cost of a finished run.
15 pairings qualify; the top 3 are shown.
-
1
DeepSeek V4.1 Flash in OpenCode
- Finished 2 of 4 (50%)
- Median cost, finished run $0.25 n=2
- Median wall time, finished run 2h 12m n=2
-
2
GPT-6 Sol in Codex CLI
- Finished 19 of 19 (100%)
- Median cost, finished run $2.10 n=19
- Median wall time, finished run 14m 28s n=19
-
3
Claude Opus 5.5 in Claude Code
- Finished 16 of 18 (89%)
- Median cost, finished run $5.00 n=16
- Median wall time, finished run 29m 45s n=16
Video: best draft score
Ranked by how often it placed in the top third of a scored prompt, then by its best rank there. Ranks come from single prompts; scores are never averaged across prompts.
No scored prompt in this category yet (only prompts with a published rubric are scored).
Video: fastest
At least half of its runs finished; ranked by the median wall time of a finished run.
16 pairings qualify; the top 3 are shown.
-
1
GPT-6 Sol in Codex CLI
- Finished 19 of 19 (100%)
- Median cost, finished run $2.10 n=19
- Median wall time, finished run 14m 28s n=19
-
2
Claude Opus 5.5 in Claude Code
- Finished 16 of 18 (89%)
- Median cost, finished run $5.00 n=16
- Median wall time, finished run 29m 45s n=16
-
3
Claude Opus 5.5 in Pi 0.87.1 + skills kit (ponytail, ECC, vetted skills)
- Finished 2 of 3 (67%)
- Median cost, finished run $12.27 n=2
- Median wall time, finished run 1h 19m n=2
Video: free only
Only its runs that cost nothing although they used tokens (free routes); ranked by finish rate, then the median wall time of a finished run.
No free run in this category finished.
Every attempt in the category counts, failures included; runs lost to infrastructure are not results. The same runs are on the Leaderboards, the Cost vs quality charts and the model pages.