DeepSeek V4.1 Flash model for reasoning and agentic coding
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
A next-generation productivity model with significantly enhanced Agent and complex task execution capabilities.
Open-weight experimental preview of the Qwen4 architecture: hybrid-attention MoE (125B total, 6B active) with vision encoder for coding, agent tasks, and image and video understanding
Dense 27B vision-language model for coding, agent tasks, and image and video understanding
DeepSeek V4 Pro snapshot with million-token context and support for thinking and non-thinking modes
Open-weight sparse MoE (2.4T total, 95B active), the open-weight twin of Qwen3.8 Max for coding, research, complex reasoning, and agentic workflows
Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads
Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding
Multimodal MoE reasoning model (276B total, 12B active) for text, image, and audio
Agentic coding model from Poolside in the XS size class for local deployment
Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Tencent Hy reasoning model for coding, instruction following, and agent tasks
Lower-latency Kimi Code variant for interactive edits and coding-agent loops
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
MiMo pro model for strong multimodal reasoning and agent execution
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
Newer StepFun flash model for faster agents, coding, and multimodal prompts
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Open MoE flagship with million-token context for coding and long agent runs
Open MiMo model for multimodal coding agents and long-context automation
Qwen vision-language model for visual reasoning, documents, and agent tasks
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Open Gemma instruction model for efficient chat and self-hosted deployments
Reasoning-optimized 398B MoE agent model with extended thinking for long-horizon and multi-turn tool use
Efficient Mistral model for fast chat, extraction, and production assistants
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
General GLM flagship for coding, analysis, and tool-heavy engineering workflows
Open-weight Qwen coding model for agents, repository edits, and multi-turn tool use
StepFun flash lane for quick multimodal reasoning and coding assistance
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
MiMo flash model for fast multimodal assistance and agent workflows
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
Compact multimodal coding model for repository exploration, file editing, and software agents
Mistral's coding-agent model for repository work, terminal tasks, and software fixes
Compact multimodal Mistral model for local assistants, edge agents, and efficient tool use
Mistral's largest general model for enterprise agents, coding, and multilingual reasoning
Thinking Kimi model for slower research passes, planning, and hard technical questions
Mistral code model for completions, refactors, and developer IDE workflows
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash | 1M | $0.15 | $0.6 | 2026-09-10 | |||
| Hy4 previewtencent/hy4-preview | 1.024M | $0.67 | $2 | 2026-08-28 | |||
| Qwen3.8 Flash Nextalibaba/qwen3.8-flash-next | 262.144K | $0.12 | $0.4 | 2026-08-27 | |||
| Qwen3.8 27Balibaba/qwen3.8-27b | 262.144K | $0.1 | $0.4 | 2026-08-14 | |||
| DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 | 1M | $0.442 | $0.884 | 2026-08-12 | |||
| Qwen3.8 2.4T A95Balibaba/qwen3.8-2.4t-a95b | 262.144K | $2 | $6 | 2026-08-12 | |||
| Nemotron 3.5 Lightning 30B A3Bnvidia/nemotron-3.5-lightning | 262.144K | $0.05 | $0.2 | 2026-08-11 | |||
| DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | 1M | $0.035 | $0.07 | 2026-07-31 | |||
| Inkling Smallthinkingmachines/inkling-small | 1.04858M | $0.45 | $1.2 | 2026-07-30 | |||
| Laguna S 2.1poolside/laguna-s-2.1 | 1.04858M | $0.09 | $0.18 | 2026-07-21 | |||
| Kimi K3moonshotai/kimi-k3 | 1.04858M | $3 | $15 | 2026-07-16 | |||
| Inklingthinkingmachines/inkling | 1.04858M | $1.87 | $4.68 | 2026-07-15 | |||
| Hy3tencent/hy3 | 256K | $0.066 | $0.26 | 2026-07-06 | |||
| Kimi K2.7 Code Highspeedmoonshotai/kimi-k2.7-code-highspeed | 262.144K | $1.9 | $8 | 2026-06-12 | |||
| Kimi K2.7 Codemoonshotai/kimi-k2.7-code | 262.144K | $0.95 | $4 | 2026-06-12 | |||
| MiMo-V2.5-Pro-UltraSpeedxiaomi/mimo-v2.5-pro-ultraspeed | 1.04858M | $1.305 | $2.61 | 2026-06-08 | |||
| Nemotron 3 Ultra 550B A55Bnvidia/nemotron-3-ultra-550b-a55b | 1M | $0.5 | $2.5 | 2026-06-04 | |||
| Step 3.7 Flashstepfun/step-3.7-flash | 256K | $0.185 | $1.11 | 2026-05-29 | |||
| DeepSeek V4 Flashdeepseek/deepseek-v4-flash | 1M | $0.15 | $0.6 | 2026-04-24 | |||
| DeepSeek V4 Prodeepseek/deepseek-v4-pro | 1M | $0.435 | $0.87 | 2026-04-24 | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| Qwen3.6 27Balibaba/qwen3.6-27b | 262.144K | $0.6 | $3.6 | 2026-04-22 | |||
| MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 1.04858M | $0.435 | $0.87 | 2026-04-22 | |||
| Kimi K2.6moonshotai/kimi-k2.6 | 262.144K | $0.95 | $4 | 2026-04-21 | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 262.144K | $0.042 | $0.22 | 2026-04-02 | |||
| Trinity Large Thinkingarcee-ai/trinity-large-thinking | 524.288K | $0.25 | $0.9 | 2026-04-01 | |||
| Mistral Small (latest)mistral/mistral-small-latest | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Nemotron 3 Super 120B A12Bnvidia/nemotron-3-super-120b-a12b | 262.144K | $0.2 | $0.8 | 2026-03-11 | |||
| GLM-5zhipuai/glm-5 | 204.8K | $1 | $3.2 | 2026-02-12 | |||
| Qwen3 Coder Nextalibaba/qwen3-coder-next | 262.144K | $0.108 | $0.675 | 2026-02-03 | |||
| Step 3.5 Flashstepfun/step-3.5-flash | 256K | $0.1 | $0.3 | 2026-01-29 | |||
| Kimi K2.5moonshotai/kimi-k2.5 | 262.144K | $0.3 | $1.9 | 2026-01 | |||
| MiMo-V2-Flashxiaomi/mimo-v2-flash | 262.144K | $0.14 | $0.28 | 2025-12-16 | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 262.144K | $0.05 | $0.2 | 2025-12-15 | |||
| Devstral Small 2mistral/devstral-small-2 | 262.144K | $0.1 | $0.3 | 2025-12-09 | |||
| Devstral 2mistral/devstral-2512 | 262.144K | $0.4 | $2 | 2025-12-09 | |||
| Ministral 14Bmistral/ministral-14b | 262.144K | $0.2 | $0.2 | 2025-12-02 | |||
| Mistral Large 3mistral/mistral-large-2512 | 262.144K | $0.5 | $1.5 | 2025-12-02 | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 262.144K | $0.4 | $2.5 | 2025-11-06 | |||
| Codestral (latest)mistral/codestral-latest | 256K | $0.3 | $0.9 | 2024-05-29 |