Model speed
A quick streaming benchmark: time to first token and output speed for every model we can reach. Sort any column; the default order is output speed.
- Measured
- 2026-09-23T19:07:34Z
- Prompt
- “Write a 400-word technical explanation of how a DualSense controller's adaptive triggers work.”
Streaming, one user message, max output 800 tokens (2000 for Opus 5.5/Fable 5.1, which cannot turn thinking off), lowest reasoning setting each model accepted. Measured from our test machine; includes one network hop. TTFT = request start to first visible text token (or first streamed reasoning token where reasoning is streamed). TPS = output tokens from usage / (end - first token); where reasoning is hidden, hidden reasoning tokens are excluded from the count; where reasoning is streamed (Grok, Kimi code, Cerebras gpt-oss) they are included. One run each; two where total < 5 s (median shown). Up to 4 parallel. Single samples: expect ±20-30% run to run.
33 models measured
Show the table (33 rows)
| Setting and notes | |||||||
|---|---|---|---|---|---|---|---|
| gpt-oss-120b | cerebras | 0.54 | 1325.8 | 1.03 | 800 | 8 | {"reasoning_effort": "low"}Cerebras; reasoning low; hit 800 cap |
| qwen-3.8-27b | cerebras | 0.31 | 989.2 | 0.82 | 584 | 0 | {"disable_reasoning": true}Cerebras; reasoning disabled |
| gemini-3.1-flash-lite | google-ai | 16.69 | 245.3 | 18.73 | 501 | 0 | {"thinkingConfig": {"thinkingLevel": "minimal"}}thinking minimal; image input; 16.7 s queue before first token |
| gemini-flash-lite-latestserved as gemini-3.5-flash-lite | google-ai | 0.98 | 218.8 | 3.17 | 494 | 0 | {"thinkingConfig": {"thinkingLevel": "minimal"}}alias; thinking minimal; image input |
| gemini-3.5-flash-lite | google-ai | 1.31 | 199.7 | 3.77 | 466 | 0 | {"thinkingConfig": {"thinkingLevel": "minimal"}}thinking minimal; image input |
| kimi-k2.7-code-highspeed | moonshot | 1.32 | 194.9 | 5.42 | 800 | 461 | {"reasoning_effort": "low"}thinking cannot be disabled; effort low; 461 of 800 tokens were reasoning (answer cut at cap), TPS counts streamed reasoning |
| gpt-5.4-miniserved as gpt-5.4-mini-2026-03-17 | openai | 0.73 | 162.0 | 3.94 | 517 | 0 | {"reasoning_effort": "none"}effort none; image input |
| deepseek-flash | deepseek | 0.87 | 153.9 | 4.08 | 483 | 0 | {"thinking": {"type": "disabled"}}thinking disabled (V4.1 Flash per root; takes images per earlier check) |
| accounts/fireworks/routers/glm-5p3-fastserved as accounts/fireworks/models/glm-5p3 | fireworks | 1.61 | 153.2 | 5.24 | 556 | 0 | {"reasoning_effort": "low"}Fireworks fast router; thinking-only, effort low |
| gpt-5.4-nanoserved as gpt-5.4-nano-2026-03-17 | openai | 0.75 | 132.6 | 4.42 | 488 | 0 | {"reasoning_effort": "none"}effort none; image input |
| gpt-4.1-miniserved as gpt-4.1-mini-2025-04-14 | openai | 0.67 | 120.5 | 4.46 | 479 | 0 | default settingsnon-reasoning; image input |
| gpt-6-luna | openai | 0.64 | 106.9 | 5.41 | 510 | 0 | {"reasoning_effort": "none"}reasoning model run at effort none; image input |
| claude-opus-5-5 | anthropic | 4.35 | 96.2 | 13.22 | 1,162 | 309 | {"thinking": {"type": "adaptive"}, "output_config": {"effort": "low"}}reasoning model; thinking cannot be disabled (adaptive only), ran effort low; image input; 309 hidden thinking tokens before text |
| accounts/fireworks/models/minimax-m3 | fireworks | 0.91 | 94.0 | 5.95 | 474 | 0 | {"reasoning_effort": "none"}Fireworks; effort none |
| accounts/fireworks/models/deepseek-v4p1-flash | fireworks | 0.62 | 82.6 | 6.96 | 524 | 0 | {"reasoning_effort": "none"}Fireworks; effort none |
| claude-haiku-4-5-20251001 | anthropic | 0.67 | 80.8 | 7.98 | 590 | 0 | {"thinking": {"type": "disabled"}}thinking disabled; image input |
| claude-sonnet-5 | anthropic | 0.79 | 80.7 | 10.70 | 800 | 0 | {"thinking": {"type": "disabled"}}thinking disabled; image input; hit 800-token cap |
| claude-opus-5 | anthropic | 1.22 | 71.5 | 12.40 | 800 | 0 | {"thinking": {"type": "disabled"}}thinking disabled; image input; hit 800-token cap |
| claude-fable-5-1 | anthropic | 2.13 | 61.0 | 17.35 | 929 | 0 | {"thinking": {"type": "adaptive"}, "output_config": {"effort": "low"}}reasoning model; thinking cannot be disabled, ran effort low (0 thinking tokens used); image input |
| accounts/fireworks/routers/kimi-k3-fastserved as accounts/fireworks/models/kimi-k3 | fireworks | 1.05 | 60.4 | 10.35 | 561 | 0 | {"reasoning_effort": "none"}Fireworks fast router; effort none |
| gpt-6-sol | openai | 0.87 | 60.3 | 9.07 | 495 | 0 | {"reasoning_effort": "none"}reasoning model run at effort none; image input |
| gpt-5.5served as gpt-5.5-2026-04-23 | openai | 0.91 | 59.8 | 9.61 | 520 | 0 | {"reasoning_effort": "none"}reasoning model run at effort none; image input |
| accounts/fireworks/models/glm-5p3 | fireworks | 1.13 | 57.2 | 10.60 | 542 | 0 | {"reasoning_effort": "low"}Fireworks; thinking-only model, ran effort low |
| grok-build-0.1 | xai | 0.50 | 55.8 | 36.48 | 2,008 | 1,538 | default settingsalways reasons (no effort param); 1,538 reasoning tokens streamed and counted, 36 s total for 400 words |
| kimi-k2.6 | moonshot | 1.00 | 55.7 | 10.38 | 523 | 1 | {"thinking": {"type": "disabled"}}thinking disabled |
| grok-4.20-0309-non-reasoning | xai | 0.66 | 54.2 | 14.20 | 734 | 0 | default settingsnon-reasoning |
| deepseek-v4-pro | deepseek | 1.13 | 52.7 | 10.00 | 467 | 0 | {"thinking": {"type": "disabled"}}thinking disabled |
| grok-4.7 | xai | 0.82 | 52.4 | 13.36 | 657 | 45 | {"reasoning_effort": "low"}reasoning model, cannot go below low; reasoning streamed and counted |
| grok-4.6 | xai | 0.65 | 39.6 | 22.13 | 851 | 326 | {"reasoning_effort": "low"}reasoning model, cannot go below low; 326 reasoning tokens streamed and counted |
| kimi-k3 | moonshot | 6.09 | 37.0 | 18.53 | 461 | 0 | {"thinking": {"type": "disabled"}}thinking disabled; slow first token (6 s) |
| gpt-6-astra | openai | 4.41 | 27.0 | 22.16 | 519 | 39 | {"reasoning_effort": "low"}reasoning model; lowest accepted effort is low (no none/minimal); image input |
| accounts/fireworks/models/qwen3p8-maxserved as Qwen 3.8 Max | fireworks | 1.27 | 18.4 | 30.78 | 543 | 0 | {"reasoning_effort": "none"}Fireworks; effort none; slow |
| gemma-4-31b-it | google-ai | 47.86 | 7.7 | 111.22 | 485 | 0 | {"thinkingConfig": {"thinkingLevel": "minimal"}}open model on Google; 48 s to first token, 7.7 TPS (heavily loaded) |
Could not measure (7)
| What happened | ||
|---|---|---|
| gemini-3.8-flash | google-ai | HTTP 400: Thinking level MINIMAL is not supported for this model. Please retry with other thinking level. HTTP 503: This model is currently experiencing high demand. Spikes in demand are usually temporary. Please try again later.MINIMAL rejected; repeated 503 "high demand"; the one run that got through was cut off after 62 tokens — no valid TPS |
| gemini-3.1-pro-preview | google-ai | HTTP 429: quota exceeded at test time |
| gemini-pro-latest | google-ai | HTTP 429: quota exceeded at test time |
| glm-5-3-flash-260828 | byteplus | HTTP 404: not available to us at test time |
| deepseek-v4-1-flash-260910 | byteplus | HTTP 404: not available to us at test time |
| gemini-3.7-flash | google-ai | HTTP 400: Thinking level MINIMAL is not supported for this model. Please retry with other thinking level. HTTP 503: This model is currently experiencing high demand. Spikes in demand are usually temporary. Please try again later.MINIMAL rejected; 503 high demand on low/budget-0; only default-thinking run got through: 625 thinking tokens ate the 800 cap, 122 words out — no valid TPS |
| (all) | groq | HTTP 400: not available to us at test timenot available to us at test time |
Not benchmarked
- groq: not available to us at test time
- byteplus: not available to us at test time
- google-ai pro models: quota exceeded at test time
- google-ai OpenAI-compat endpoint: not usable for streaming at test time; the native streaming API was used instead