Nemotron model for efficient reasoning, coding, and specialized AI agents
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Low-latency M2.7 variant for interactive coding plans and agent loops
Open MiniMax flagship for coding agents, office automation, and complex environments
Efficient Mistral model for fast chat, extraction, and production assistants
Fast Mistral production model for chat, extraction, and cost-sensitive agents
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen instruction model for multilingual chat, reasoning, and tool use
Large open Qwen multimodal MoE for visual agents and long technical tasks
High-speed MiniMax model for low-latency coding and agent workflows
General GLM flagship for coding, analysis, and tool-heavy engineering workflows
Prior MiniMax coding model for agent workflows, office edits, and automation
Open-weight Qwen coding model for agents, repository edits, and multi-turn tool use
StepFun flash lane for quick multimodal reasoning and coding assistance
Lightly post-trained 398B MoE chat model for creative work, long-context prompts, and tool-using agents
Efficient GLM model for fast reasoning, coding, and agent workflows
Budget GLM lane for fast coding help, routing, and everyday automation
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
Earlier MiniMax agent model for practical coding and productivity tasks
Mature GLM model for dependable coding, reasoning, and structured agent tasks
MiMo flash model for fast multimodal assistance and agent workflows
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
Compact multimodal coding model for repository exploration, file editing, and software agents
Mistral's coding-agent model for repository work, terminal tasks, and software fixes
Open vision-language model for efficient local deployment, instruction following, and tool use
Mistral's largest general model for enterprise agents, coding, and multilingual reasoning
Compact multimodal Mistral model for local assistants, edge agents, and efficient tool use
Compact open vision-language model for edge deployment, instruction following, and tool use
Compact open vision-language model for edge deployment, instruction following, and tool use
Mistral coding agent model for repository tasks and software engineering workflows
DeepSeek reasoning model for multi-step analysis, math, coding, and tools
DeepSeek chat model for instruction following, coding, and analysis
Kimi reasoning model for long-horizon research, planning, and tool use
Thinking Kimi model for slower research passes, planning, and hard technical questions
Efficient open MiniMax model built for coding agents and tool-heavy workflows
Late GLM-4 workhorse for coding agents, reasoning, and structured tasks
Cohere reasoning model for multilingual enterprise agents, tools, and complex workflows
Updated large open Qwen3 MoE instruct model for multilingual chat, coding, and tool use
Open multimodal Llama for strong reasoning with efficient everyday serving
Open Llama with long-context vision for efficient multimodal agents
Smaller Qwen coder for efficient local agents and repo-level fixes
Open Qwen coding heavyweight for repository reasoning and agentic engineering
Cohere command model for multilingual enterprise agents, tools, and chat
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Mistral code model for completions, refactors, and developer IDE workflows
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Nemotron Cascade 2 30B A3Bnvidia/nemotron-cascade-2-30b-a3b | 256K | — | — | 2026-03-24 | |||
| MiniMax-M2.7-highspeedminimax/MiniMax-M2.7-highspeed | 204.8K | $0.6 | $2.4 | 2026-03-18 | |||
| MiniMax-M2.7minimax/MiniMax-M2.7 | 204.8K | $0.3 | $1.2 | 2026-03-18 | |||
| Mistral Small (latest)mistral/mistral-small-latest | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Mistral Small 4mistral/mistral-small-2603 | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Nemotron 3 Super 120B A12Bnvidia/nemotron-3-super-120b-a12b | 262.144K | $0.2 | $0.8 | 2026-03-11 | |||
| Qwen3.5 35B-A3Balibaba/qwen3.5-35b-a3b | 262.144K | $0.25 | $2 | 2026-02-23 | |||
| Qwen3.5 122B-A10Balibaba/qwen3.5-122b-a10b | 262.144K | $0.4 | $3.2 | 2026-02-23 | |||
| Qwen3.5 27Balibaba/qwen3.5-27b | 262.144K | $0.3 | $2.4 | 2026-02-23 | |||
| Qwen3.5 9Balibaba/qwen3.5-9b | 262.144K | $0.04 | $0.15 | 2026-02-23 | |||
| Qwen3.5 397B-A17Balibaba/qwen3.5-397b-a17b | 262.144K | $0.6 | $3.6 | 2026-02-15 | |||
| MiniMax-M2.5-highspeedminimax/MiniMax-M2.5-highspeed | 204.8K | $0.6 | $2.4 | 2026-02-13 | |||
| GLM-5zhipuai/glm-5 | 204.8K | $1 | $3.2 | 2026-02-12 | |||
| MiniMax-M2.5minimax/MiniMax-M2.5 | 204.8K | $0.3 | $1.2 | 2026-02-12 | |||
| Qwen3 Coder Nextalibaba/qwen3-coder-next | 262.144K | $0.108 | $0.675 | 2026-02-03 | |||
| Step 3.5 Flashstepfun/step-3.5-flash | 256K | $0.1 | $0.3 | 2026-01-29 | |||
| Trinity Large Previewarcee-ai/trinity-large-preview | 524.288K | — | — | 2026-01-27 | |||
| GLM-4.7-FlashXzhipuai/glm-4.7-flashx | 200K | $0.07 | $0.4 | 2026-01-19 | |||
| GLM-4.7-Flashzhipuai/glm-4.7-flash | 200K | $0.06 | $0.4 | 2026-01-19 | |||
| Kimi K2.5moonshotai/kimi-k2.5 | 262.144K | $0.3 | $1.9 | 2026-01 | |||
| MiniMax-M2.1minimax/MiniMax-M2.1 | 204.8K | $0.3 | $1.2 | 2025-12-23 | |||
| GLM-4.7zhipuai/glm-4.7 | 204.8K | $0.6 | $2.2 | 2025-12-22 | |||
| MiMo-V2-Flashxiaomi/mimo-v2-flash | 262.144K | $0.14 | $0.28 | 2025-12-16 | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 262.144K | $0.05 | $0.2 | 2025-12-15 | |||
| Devstral Small 2mistral/devstral-small-2 | 262.144K | $0.1 | $0.3 | 2025-12-09 | |||
| Devstral 2mistral/devstral-2512 | 262.144K | $0.4 | $2 | 2025-12-09 | |||
| Ministral 3 14Bmistral/ministral-3-14b-instruct-2512 | 262.144K | $0.1 | $0.4 | 2025-12-02 | |||
| Mistral Large 3mistral/mistral-large-2512 | 262.144K | $0.5 | $1.5 | 2025-12-02 | |||
| Ministral 14Bmistral/ministral-14b | 262.144K | $0.2 | $0.2 | 2025-12-02 | |||
| Ministral 3 3Bmistral/ministral-3-3b-instruct-2512 | 262.144K | $0.1 | $0.1 | 2025-12-02 | |||
| Ministral 3 8Bmistral/ministral-3-8b-instruct-2512 | 262.144K | $0.15 | $0.15 | 2025-12-02 | |||
| Devstral 2 (latest)mistral/devstral-medium-latest | 262.144K | $0.4 | $2 | 2025-12-02 | |||
| DeepSeek Reasonerdeepseek/deepseek-reasoner | 1M | $0.147 | $0.295 | 2025-12-01 | |||
| DeepSeek Chatdeepseek/deepseek-chat | 1M | $0.147 | $0.295 | 2025-12-01 | |||
| Kimi K2 Thinking Turbomoonshotai/kimi-k2-thinking-turbo | 262.144K | $1.15 | $8 | 2025-11-06 | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 262.144K | $0.4 | $2.5 | 2025-11-06 | |||
| MiniMax-M2minimax/MiniMax-M2 | 204.8K | $0.3 | $1.2 | 2025-10-27 | |||
| GLM-4.6zhipuai/glm-4.6 | 204.8K | $0.6 | $2.2 | 2025-09-30 | |||
| Command A Reasoningcohere/command-a-reasoning-08-2025 | 256K | $2.5 | $10 | 2025-08-21 | |||
| Qwen3 235B-A22B Instruct 2507alibaba/qwen3-235b-a22b-instruct-2507 | 262.144K | $0.069 | $0.455 | 2025-07-21 | |||
| Llama 4 Maverick 17B Instructmeta/llama-4-maverick-17b-instruct | 1M | $0.14 | $0.59 | 2025-04-05 | |||
| Llama 4 Scout 17B Instructmeta/llama-4-scout-17b-instruct | 10M | $0.17 | $0.66 | 2025-04-05 | |||
| Qwen3-Coder 30B-A3B Instructalibaba/qwen3-coder-30b-a3b-instruct | 262.144K | $0.45 | $2.25 | 2025-04 | |||
| Qwen3-Coder 480B-A35B Instructalibaba/qwen3-coder-480b-a35b-instruct | 262.144K | $1.5 | $7.5 | 2025-04 | |||
| Command Acohere/command-a-03-2025 | 256K | $2.5 | $10 | 2025-03-13 | |||
| Mistral Large (latest)mistral/mistral-large-latest | 262.144K | $0.5 | $1.5 | 2024-11-01 | |||
| Codestral (latest)mistral/codestral-latest | 256K | $0.3 | $0.9 | 2024-05-29 |