DeepSeek V4.1 Flash model for reasoning and agentic coding
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
A next-generation productivity model with significantly enhanced Agent and complex task execution capabilities.
Open-weight experimental preview of the Qwen4 architecture: hybrid-attention MoE (125B total, 6B active) with vision encoder for coding, agent tasks, and image and video understanding
Native multimodal GLM model for efficient coding and long-horizon agent tasks
Mixture-of-experts coding-reasoning model for agentic software tasks, tool use, and image understanding
Flagship GLM model for long-horizon coding, agents, and complex project delivery
Dense 27B vision-language model for coding, agent tasks, and image and video understanding
Open-weight sparse MoE (2.4T total, 95B active), the open-weight twin of Qwen3.8 Max for coding, research, complex reasoning, and agentic workflows
DeepSeek V4 Pro snapshot with million-token context and support for thinking and non-thinking modes
Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads
Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads
Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding
Multimodal MoE reasoning model (276B total, 12B active) for text, image, and audio
Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Tencent Hy reasoning model for coding, instruction following, and agent tasks
Open flagship GLM for long-horizon coding agents and million-token context work
Lower-latency Kimi Code variant for interactive edits and coding-agent loops
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
Cohere coding model for practical software engineering and agentic edits
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
MiniMax multimodal model for long-context coding, perception, and agent planning
Newer StepFun flash model for faster agents, coding, and multimodal prompts
Balanced Mistral model for enterprise assistants, multilingual work, and tools
Balanced Mistral model for enterprise assistants, multilingual work, and tools
Open Nemotron omni model combining reasoning with text, vision, and audio
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Open MoE flagship with million-token context for coding and long agent runs
Initial DeepSeek V4 Flash snapshot for economical reasoning, coding, and million-token agent workloads
DeepSeek V4 Pro initial snapshot with million-token context and support for thinking and non-thinking modes
Open MiMo model for multimodal coding agents and long-context automation
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
Qwen vision-language model for visual reasoning, documents, and agent tasks
Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
Tencent Hy reasoning model for coding, instruction following, and agent tasks
Open multimodal Qwen MoE for local agents that need vision, audio, and code
Strong GLM coding model for agentic engineering, terminals, and repository generation
StepFun flash model for efficient multimodal reasoning, coding, and tool use
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Open Gemma instruction model for efficient chat and self-hosted deployments
Nemotron model for efficient reasoning, coding, and specialized AI agents
Low-latency M2.7 variant for interactive coding plans and agent loops
Open MiniMax flagship for coding agents, office automation, and complex environments
Efficient Mistral model for fast chat, extraction, and production assistants
Fast Mistral production model for chat, extraction, and cost-sensitive agents
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash | 1M | $0.15 | $0.6 | 2026-09-10 | |||
| Hy4 previewtencent/hy4-preview | 1.024M | $0.67 | $2 | 2026-08-28 | |||
| Qwen3.8 Flash Nextalibaba/qwen3.8-flash-next | 262.144K | $0.12 | $0.4 | 2026-08-27 | |||
| GLM-5.3-Flashzhipuai/glm-5.3-flash | 1M | $0.075 | $0.25 | 2026-08-26 | |||
| Ornith 1.5 35B A3Bdeepreinforce/ornith-1.5-35b-a3b | 262.144K | $0.1 | $0.4 | 2026-08-18 | |||
| GLM-5.3zhipuai/glm-5.3 | 1M | $1.4 | $4.4 | 2026-08-14 | |||
| Qwen3.8 27Balibaba/qwen3.8-27b | 262.144K | $0.1 | $0.4 | 2026-08-14 | |||
| Qwen3.8 2.4T A95Balibaba/qwen3.8-2.4t-a95b | 262.144K | $2 | $6 | 2026-08-12 | |||
| DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 | 1M | $0.442 | $0.884 | 2026-08-12 | |||
| Nemotron 3.5 Lightning 30B A3Bnvidia/nemotron-3.5-lightning-30b-a3b | 262.144K | — | — | 2026-08-11 | |||
| Nemotron 3.5 Lightning 30B A3Bnvidia/nemotron-3.5-lightning | 262.144K | $0.05 | $0.2 | 2026-08-11 | |||
| DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | 1M | $0.035 | $0.07 | 2026-07-31 | |||
| Inkling Smallthinkingmachines/inkling-small | 1.04858M | $0.45 | $1.2 | 2026-07-30 | |||
| Kimi K3moonshotai/kimi-k3 | 1.04858M | $3 | $15 | 2026-07-16 | |||
| Inklingthinkingmachines/inkling | 1.04858M | $1.87 | $4.68 | 2026-07-15 | |||
| Hy3tencent/hy3 | 256K | $0.066 | $0.26 | 2026-07-06 | |||
| GLM-5.2zhipuai/glm-5.2 | 1M | $1.4 | $4.4 | 2026-06-13 | |||
| Kimi K2.7 Code Highspeedmoonshotai/kimi-k2.7-code-highspeed | 262.144K | $1.9 | $8 | 2026-06-12 | |||
| Kimi K2.7 Codemoonshotai/kimi-k2.7-code | 262.144K | $0.95 | $4 | 2026-06-12 | |||
| North Mini Codecohere/north-mini-code-1-0 | 256K | — | — | 2026-06-09 | |||
| Nemotron 3 Ultra 550B A55Bnvidia/nemotron-3-ultra-550b-a55b | 1M | $0.5 | $2.5 | 2026-06-04 | |||
| MiniMax-M3minimax/MiniMax-M3 | 1.04858M | $0.3 | $1.2 | 2026-06-01 | |||
| Step 3.7 Flashstepfun/step-3.7-flash | 256K | $0.185 | $1.11 | 2026-05-29 | |||
| Mistral Medium 3.5mistral/mistral-medium-2604 | 262.144K | $1.5 | $7.5 | 2026-04-29 | |||
| Mistral Medium (latest)mistral/mistral-medium-latest | 262.144K | $1.5 | $7.5 | 2026-04-29 | |||
| Nemotron 3 Nano Omni 30B A3B Reasoningnvidia/nemotron-3-nano-omni-30b-a3b-reasoning | 256K | $0.2 | $0.8 | 2026-04-28 | |||
| DeepSeek V4 Flashdeepseek/deepseek-v4-flash | 1M | $0.15 | $0.6 | 2026-04-24 | |||
| DeepSeek V4 Prodeepseek/deepseek-v4-pro | 1M | $0.435 | $0.87 | 2026-04-24 | |||
| DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash-0423 | 1M | $0.139 | $0.278 | 2026-04-23 | |||
| DeepSeek V4 Pro 0423deepseek/deepseek-v4-pro-0423 | 1M | $1.32 | $3.96 | 2026-04-23 | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 1.04858M | $0.435 | $0.87 | 2026-04-22 | |||
| Qwen3.6 27Balibaba/qwen3.6-27b | 262.144K | $0.6 | $3.6 | 2026-04-22 | |||
| Kimi K2.6moonshotai/kimi-k2.6 | 262.144K | $0.95 | $4 | 2026-04-21 | |||
| Hy3 previewtencent/hy3-preview | 256K | $0.066 | $0.26 | 2026-04-20 | |||
| Qwen3.6 35B-A3Balibaba/qwen3.6-35b-a3b | 262.144K | $0.248 | $1.485 | 2026-04-17 | |||
| GLM-5.1zhipuai/glm-5.1 | 200K | $1.4 | $4.4 | 2026-04-07 | |||
| Step 3.5 Flash 2603stepfun/step-3.5-flash-2603 | 256K | $0.1 | $0.3 | 2026-04-02 | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 262.144K | $0.042 | $0.22 | 2026-04-02 | |||
| Nemotron Cascade 2 30B A3Bnvidia/nemotron-cascade-2-30b-a3b | 256K | — | — | 2026-03-24 | |||
| MiniMax-M2.7-highspeedminimax/MiniMax-M2.7-highspeed | 204.8K | $0.6 | $2.4 | 2026-03-18 | |||
| MiniMax-M2.7minimax/MiniMax-M2.7 | 204.8K | $0.3 | $1.2 | 2026-03-18 | |||
| Mistral Small (latest)mistral/mistral-small-latest | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Mistral Small 4mistral/mistral-small-2603 | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Nemotron 3 Super 120B A12Bnvidia/nemotron-3-super-120b-a12b | 262.144K | $0.2 | $0.8 | 2026-03-11 | |||
| Qwen3.5 9Balibaba/qwen3.5-9b | 262.144K | $0.04 | $0.15 | 2026-02-23 | |||
| Qwen3.5 35B-A3Balibaba/qwen3.5-35b-a3b | 262.144K | $0.25 | $2 | 2026-02-23 | |||
| Qwen3.5 27Balibaba/qwen3.5-27b | 262.144K | $0.3 | $2.4 | 2026-02-23 | |||
| Qwen3.5 122B-A10Balibaba/qwen3.5-122b-a10b | 262.144K | $0.4 | $3.2 | 2026-02-23 |