DeepSeek V4.1 Flash model for reasoning and agentic coding
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Dense 27B vision-language model for coding, agent tasks, and image and video understanding
DeepSeek V4 Pro snapshot with million-token context and support for thinking and non-thinking modes
Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads
Muse Glimmer is a 30-billion-parameter open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark for always-on local agents, tool use, coding, and image understanding.
Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding
Multimodal MoE reasoning model (276B total, 12B active) for text, image, and audio
Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Tencent Hy reasoning model for coding, instruction following, and agent tasks
Agentic coding model from Poolside in the XS size class for local deployment
Open flagship GLM for long-horizon coding agents and million-token context work
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
Cohere coding model for practical software engineering and agentic edits
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
MiniMax multimodal model for long-context coding, perception, and agent planning
Newer StepFun flash model for faster agents, coding, and multimodal prompts
Balanced Mistral model for enterprise assistants, multilingual work, and tools
Balanced Mistral model for enterprise assistants, multilingual work, and tools
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Open MoE flagship with million-token context for coding and long agent runs
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
Open MiMo model for multimodal coding agents and long-context automation
Qwen vision-language model for visual reasoning, documents, and agent tasks
Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
Tencent Hy reasoning model for coding, instruction following, and agent tasks
Open multimodal Qwen MoE for local agents that need vision, audio, and code
Strong GLM coding model for agentic engineering, terminals, and repository generation
StepFun flash model for efficient multimodal reasoning, coding, and tool use
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Open Gemma instruction model for efficient chat and self-hosted deployments
Reasoning-optimized 398B MoE agent model with extended thinking for long-horizon and multi-turn tool use
Open MiniMax flagship for coding agents, office automation, and complex environments
Fast Mistral production model for chat, extraction, and cost-sensitive agents
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Large open Qwen multimodal MoE for visual agents and long technical tasks
Prior MiniMax coding model for agent workflows, office edits, and automation
General GLM flagship for coding, analysis, and tool-heavy engineering workflows
StepFun flash lane for quick multimodal reasoning and coding assistance
Budget GLM lane for fast coding help, routing, and everyday automation
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
Earlier MiniMax agent model for practical coding and productivity tasks
Mature GLM model for dependable coding, reasoning, and structured agent tasks
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
Mistral's coding-agent model for repository work, terminal tasks, and software fixes
DeepSeek chat model for instruction following, coding, and analysis
DeepSeek reasoning model for multi-step analysis, math, coding, and tools
Hybrid-reasoning DeepSeek model with thinking and non-thinking modes, sparse attention, and tool-use
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash | 1M | $0.15 | $0.6 | 2026-09-10 | |||
| Qwen3.8 27Balibaba/qwen3.8-27b | 262.144K | $0.1 | $0.4 | 2026-08-14 | |||
| DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 | 1M | $0.442 | $0.884 | 2026-08-12 | |||
| Nemotron 3.5 Lightning 30B A3Bnvidia/nemotron-3.5-lightning | 262.144K | $0.05 | $0.2 | 2026-08-11 | |||
| Muse Glimmer 30Bmeta/muse-glimmer-30b | 131.072K | $0.2 | $0.8 | 2026-08-10 | |||
| DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | 1M | $0.05 | $0.16 | 2026-07-31 | |||
| Inkling Smallthinkingmachines/inkling-small | 1.04858M | $0.45 | $1.2 | 2026-07-30 | |||
| Kimi K3moonshotai/kimi-k3 | 1.04858M | $3 | $15 | 2026-07-16 | |||
| Inklingthinkingmachines/inkling | 1.04858M | $1.87 | $4.68 | 2026-07-15 | |||
| Hy3tencent/hy3 | 256K | $0.066 | $0.26 | 2026-07-06 | |||
| Laguna XS 2.1poolside/laguna-xs-2.1 | 262.144K | $0.06 | $0.12 | 2026-07-02 | |||
| GLM-5.2zhipuai/glm-5.2 | 1M | $1.4 | $4.4 | 2026-06-13 | |||
| Kimi K2.7 Codemoonshotai/kimi-k2.7-code | 262.144K | $0.95 | $4 | 2026-06-12 | |||
| North Mini Codecohere/north-mini-code-1-0 | 256K | — | — | 2026-06-09 | |||
| Nemotron 3 Ultra 550B A55Bnvidia/nemotron-3-ultra-550b-a55b | 1M | $0.5 | $2.5 | 2026-06-04 | |||
| MiniMax-M3minimax/MiniMax-M3 | 1.04858M | $0.3 | $1.2 | 2026-06-01 | |||
| Step 3.7 Flashstepfun/step-3.7-flash | 256K | $0.185 | $1.11 | 2026-05-29 | |||
| Mistral Medium 3.5mistral/mistral-medium-2604 | 262.144K | $1.5 | $7.5 | 2026-04-29 | |||
| Mistral Medium (latest)mistral/mistral-medium-latest | 262.144K | $1.5 | $7.5 | 2026-04-29 | |||
| DeepSeek V4 Flashdeepseek/deepseek-v4-flash | 1M | $0.15 | $0.6 | 2026-04-24 | |||
| DeepSeek V4 Prodeepseek/deepseek-v4-pro | 1M | $0.435 | $0.87 | 2026-04-24 | |||
| MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 1.04858M | $0.435 | $0.87 | 2026-04-22 | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| Qwen3.6 27Balibaba/qwen3.6-27b | 262.144K | $0.6 | $3.6 | 2026-04-22 | |||
| Kimi K2.6moonshotai/kimi-k2.6 | 262.144K | $0.95 | $4 | 2026-04-21 | |||
| Hy3 previewtencent/hy3-preview | 256K | $0.066 | $0.26 | 2026-04-20 | |||
| Qwen3.6 35B-A3Balibaba/qwen3.6-35b-a3b | 262.144K | $0.248 | $1.485 | 2026-04-17 | |||
| GLM-5.1zhipuai/glm-5.1 | 200K | $1.4 | $4.4 | 2026-04-07 | |||
| Step 3.5 Flash 2603stepfun/step-3.5-flash-2603 | 256K | $0.1 | $0.3 | 2026-04-02 | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 262.144K | $0.042 | $0.22 | 2026-04-02 | |||
| Trinity Large Thinkingarcee-ai/trinity-large-thinking | 524.288K | $0.25 | $0.9 | 2026-04-01 | |||
| MiniMax-M2.7minimax/MiniMax-M2.7 | 204.8K | $0.3 | $1.2 | 2026-03-18 | |||
| Mistral Small 4mistral/mistral-small-2603 | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Nemotron 3 Super 120B A12Bnvidia/nemotron-3-super-120b-a12b | 262.144K | $0.2 | $0.8 | 2026-03-11 | |||
| Qwen3.5 122B-A10Balibaba/qwen3.5-122b-a10b | 262.144K | $0.4 | $3.2 | 2026-02-23 | |||
| Qwen3.5 27Balibaba/qwen3.5-27b | 262.144K | $0.3 | $2.4 | 2026-02-23 | |||
| Qwen3.5 397B-A17Balibaba/qwen3.5-397b-a17b | 262.144K | $0.6 | $3.6 | 2026-02-15 | |||
| MiniMax-M2.5minimax/MiniMax-M2.5 | 204.8K | $0.3 | $1.2 | 2026-02-12 | |||
| GLM-5zhipuai/glm-5 | 204.8K | $1 | $3.2 | 2026-02-12 | |||
| Step 3.5 Flashstepfun/step-3.5-flash | 256K | $0.1 | $0.3 | 2026-01-29 | |||
| GLM-4.7-Flashzhipuai/glm-4.7-flash | 200K | $0.06 | $0.4 | 2026-01-19 | |||
| Kimi K2.5moonshotai/kimi-k2.5 | 262.144K | $0.3 | $1.9 | 2026-01 | |||
| MiniMax-M2.1minimax/MiniMax-M2.1 | 204.8K | $0.3 | $1.2 | 2025-12-23 | |||
| GLM-4.7zhipuai/glm-4.7 | 204.8K | $0.6 | $2.2 | 2025-12-22 | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 262.144K | $0.05 | $0.2 | 2025-12-15 | |||
| Devstral 2mistral/devstral-2512 | 262.144K | $0.4 | $2 | 2025-12-09 | |||
| DeepSeek Chatdeepseek/deepseek-chat | 1M | $0.147 | $0.295 | 2025-12-01 | |||
| DeepSeek Reasonerdeepseek/deepseek-reasoner | 1M | $0.147 | $0.295 | 2025-12-01 | |||
| DeepSeek V3.2deepseek/deepseek-v3.2 | 128K | $0.18 | $0.35 | 2025-12-01 |