AI Model Reviewer

Model speed

A quick streaming benchmark: time to first token and output speed for every model we can reach. Sort any column; the default order is output speed.

Measured
2026-09-23T19:07:34Z
Prompt
“Write a 400-word technical explanation of how a DualSense controller's adaptive triggers work.”

Streaming, one user message, max output 800 tokens (2000 for Opus 5.5/Fable 5.1, which cannot turn thinking off), lowest reasoning setting each model accepted. Measured from our test machine; includes one network hop. TTFT = request start to first visible text token (or first streamed reasoning token where reasoning is streamed). TPS = output tokens from usage / (end - first token); where reasoning is hidden, hidden reasoning tokens are excluded from the count; where reasoning is streamed (Grok, Kimi code, Cerebras gpt-oss) they are included. One run each; two where total < 5 s (median shown). Up to 4 parallel. Single samples: expect ±20-30% run to run.

33 models measured

Show the table (33 rows)
Setting and notes
gpt-oss-120bcerebras0.541325.81.038008{"reasoning_effort": "low"}Cerebras; reasoning low; hit 800 cap
qwen-3.8-27bcerebras0.31989.20.825840{"disable_reasoning": true}Cerebras; reasoning disabled
gemini-3.1-flash-litegoogle-ai16.69245.318.735010{"thinkingConfig": {"thinkingLevel": "minimal"}}thinking minimal; image input; 16.7 s queue before first token
gemini-flash-lite-latestserved as gemini-3.5-flash-litegoogle-ai0.98218.83.174940{"thinkingConfig": {"thinkingLevel": "minimal"}}alias; thinking minimal; image input
gemini-3.5-flash-litegoogle-ai1.31199.73.774660{"thinkingConfig": {"thinkingLevel": "minimal"}}thinking minimal; image input
kimi-k2.7-code-highspeedmoonshot1.32194.95.42800461{"reasoning_effort": "low"}thinking cannot be disabled; effort low; 461 of 800 tokens were reasoning (answer cut at cap), TPS counts streamed reasoning
gpt-5.4-miniserved as gpt-5.4-mini-2026-03-17openai0.73162.03.945170{"reasoning_effort": "none"}effort none; image input
deepseek-flashdeepseek0.87153.94.084830{"thinking": {"type": "disabled"}}thinking disabled (V4.1 Flash per root; takes images per earlier check)
accounts/fireworks/routers/glm-5p3-fastserved as accounts/fireworks/models/glm-5p3fireworks1.61153.25.245560{"reasoning_effort": "low"}Fireworks fast router; thinking-only, effort low
gpt-5.4-nanoserved as gpt-5.4-nano-2026-03-17openai0.75132.64.424880{"reasoning_effort": "none"}effort none; image input
gpt-4.1-miniserved as gpt-4.1-mini-2025-04-14openai0.67120.54.464790default settingsnon-reasoning; image input
gpt-6-lunaopenai0.64106.95.415100{"reasoning_effort": "none"}reasoning model run at effort none; image input
claude-opus-5-5anthropic4.3596.213.221,162309{"thinking": {"type": "adaptive"}, "output_config": {"effort": "low"}}reasoning model; thinking cannot be disabled (adaptive only), ran effort low; image input; 309 hidden thinking tokens before text
accounts/fireworks/models/minimax-m3fireworks0.9194.05.954740{"reasoning_effort": "none"}Fireworks; effort none
accounts/fireworks/models/deepseek-v4p1-flashfireworks0.6282.66.965240{"reasoning_effort": "none"}Fireworks; effort none
claude-haiku-4-5-20251001anthropic0.6780.87.985900{"thinking": {"type": "disabled"}}thinking disabled; image input
claude-sonnet-5anthropic0.7980.710.708000{"thinking": {"type": "disabled"}}thinking disabled; image input; hit 800-token cap
claude-opus-5anthropic1.2271.512.408000{"thinking": {"type": "disabled"}}thinking disabled; image input; hit 800-token cap
claude-fable-5-1anthropic2.1361.017.359290{"thinking": {"type": "adaptive"}, "output_config": {"effort": "low"}}reasoning model; thinking cannot be disabled, ran effort low (0 thinking tokens used); image input
accounts/fireworks/routers/kimi-k3-fastserved as accounts/fireworks/models/kimi-k3fireworks1.0560.410.355610{"reasoning_effort": "none"}Fireworks fast router; effort none
gpt-6-solopenai0.8760.39.074950{"reasoning_effort": "none"}reasoning model run at effort none; image input
gpt-5.5served as gpt-5.5-2026-04-23openai0.9159.89.615200{"reasoning_effort": "none"}reasoning model run at effort none; image input
accounts/fireworks/models/glm-5p3fireworks1.1357.210.605420{"reasoning_effort": "low"}Fireworks; thinking-only model, ran effort low
grok-build-0.1xai0.5055.836.482,0081,538default settingsalways reasons (no effort param); 1,538 reasoning tokens streamed and counted, 36 s total for 400 words
kimi-k2.6moonshot1.0055.710.385231{"thinking": {"type": "disabled"}}thinking disabled
grok-4.20-0309-non-reasoningxai0.6654.214.207340default settingsnon-reasoning
deepseek-v4-prodeepseek1.1352.710.004670{"thinking": {"type": "disabled"}}thinking disabled
grok-4.7xai0.8252.413.3665745{"reasoning_effort": "low"}reasoning model, cannot go below low; reasoning streamed and counted
grok-4.6xai0.6539.622.13851326{"reasoning_effort": "low"}reasoning model, cannot go below low; 326 reasoning tokens streamed and counted
kimi-k3moonshot6.0937.018.534610{"thinking": {"type": "disabled"}}thinking disabled; slow first token (6 s)
gpt-6-astraopenai4.4127.022.1651939{"reasoning_effort": "low"}reasoning model; lowest accepted effort is low (no none/minimal); image input
accounts/fireworks/models/qwen3p8-maxserved as Qwen 3.8 Maxfireworks1.2718.430.785430{"reasoning_effort": "none"}Fireworks; effort none; slow
gemma-4-31b-itgoogle-ai47.867.7111.224850{"thinkingConfig": {"thinkingLevel": "minimal"}}open model on Google; 48 s to first token, 7.7 TPS (heavily loaded)

Could not measure (7)

What happened
gemini-3.8-flashgoogle-aiHTTP 400: Thinking level MINIMAL is not supported for this model. Please retry with other thinking level.
HTTP 503: This model is currently experiencing high demand. Spikes in demand are usually temporary. Please try again later.MINIMAL rejected; repeated 503 "high demand"; the one run that got through was cut off after 62 tokens — no valid TPS
gemini-3.1-pro-previewgoogle-aiHTTP 429: quota exceeded at test time
gemini-pro-latestgoogle-aiHTTP 429: quota exceeded at test time
glm-5-3-flash-260828byteplusHTTP 404: not available to us at test time
deepseek-v4-1-flash-260910byteplusHTTP 404: not available to us at test time
gemini-3.7-flashgoogle-aiHTTP 400: Thinking level MINIMAL is not supported for this model. Please retry with other thinking level.
HTTP 503: This model is currently experiencing high demand. Spikes in demand are usually temporary. Please try again later.MINIMAL rejected; 503 high demand on low/budget-0; only default-thinking run got through: 625 thinking tokens ate the 800 cap, 122 words out — no valid TPS
(all)groqHTTP 400: not available to us at test timenot available to us at test time

Not benchmarked

  • groq: not available to us at test time
  • byteplus: not available to us at test time
  • google-ai pro models: quota exceeded at test time
  • google-ai OpenAI-compat endpoint: not usable for streaming at test time; the native streaming API was used instead