Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Multimodal MoE reasoning model (276B total, 12B active) for text, image, and audio
Agentic coding model from Poolside in the XS size class for local deployment
Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work
Open flagship GLM for long-horizon coding agents and million-token context work
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
MiniMax multimodal model for long-context coding, perception, and agent planning
Balanced Mistral model for enterprise assistants, multilingual work, and tools
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Open MoE flagship with million-token context for coding and long agent runs
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
Qwen vision-language model for visual reasoning, documents, and agent tasks
Open MiMo model for multimodal coding agents and long-context automation
Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
Open multimodal Qwen MoE for local agents that need vision, audio, and code
Strong GLM coding model for agentic engineering, terminals, and repository generation
Open Gemma instruction model for efficient chat and self-hosted deployments
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Open Gemma instruction model for efficient chat and self-hosted deployments
Open MiniMax flagship for coding agents, office automation, and complex environments
Fast Mistral production model for chat, extraction, and cost-sensitive agents
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
Qwen instruction model for multilingual chat, reasoning, and tool use
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
Mistral's coding-agent model for repository work, terminal tasks, and software fixes
Compact multimodal coding model for repository exploration, file editing, and software agents
Compact multimodal Mistral model for local assistants, edge agents, and efficient tool use
Hybrid-reasoning DeepSeek model with thinking and non-thinking modes
Open GPT reasoning model for self-hosted agents and controllable deployments
Open GPT reasoning model for self-hosted agents and controllable deployments
Updated large open Qwen3 MoE instruct model for multilingual chat, coding, and tool use
Open Mistral reasoning model for transparent step-by-step problem solving
Dense open Qwen model for self-hosted chat, reasoning, and coding
Open DeepSeek MoE chat model for coding, math, and general reasoning
Popular open Llama workhorse for multilingual chat, coding, and self-hosting
Tiny open Qwen code model for lightweight completion and on-device coding
Mistral's largest general model for enterprise agents, coding, and multilingual reasoning
Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads
Efficient open Mistral edge model for on-device chat and function calling
Small open Llama base model for lightweight text generation and self-hosting
Compact open Llama base model for lightweight and on-device use
Mistral vision-language model for image understanding and multimodal chat
Efficient Mistral-NVIDIA open model for multilingual chat and local deployment
Open Mistral code model for fill-in-the-middle and 80+ programming languages
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Nemotron 3.5 Lightning 30B A3Bnvidia/nemotron-3.5-lightning | 262.144K | $0.05 | $0.2 | 2026-08-11 | |||
| Inkling Smallthinkingmachines/inkling-small | 1.04858M | $0.45 | $1.2 | 2026-07-30 | |||
| Laguna S 2.1poolside/laguna-s-2.1 | 1.04858M | $0.09 | $0.18 | 2026-07-21 | |||
| Kimi K3moonshotai/kimi-k3 | 1.04858M | $3 | $15 | 2026-07-16 | |||
| GLM-5.2zhipuai/glm-5.2 | 1M | $1.4 | $4.4 | 2026-06-13 | |||
| Kimi K2.7 Codemoonshotai/kimi-k2.7-code | 262.144K | $0.95 | $4 | 2026-06-12 | |||
| Nemotron 3 Ultra 550B A55Bnvidia/nemotron-3-ultra-550b-a55b | 1M | $0.5 | $2.5 | 2026-06-04 | |||
| MiniMax-M3minimax/MiniMax-M3 | 1.04858M | $0.3 | $1.2 | 2026-06-01 | |||
| Mistral Medium 3.5mistral/mistral-medium-2604 | 262.144K | $1.5 | $7.5 | 2026-04-29 | |||
| DeepSeek V4 Flashdeepseek/deepseek-v4-flash | 1M | $0.15 | $0.6 | 2026-04-24 | |||
| DeepSeek V4 Prodeepseek/deepseek-v4-pro | 1M | $0.435 | $0.87 | 2026-04-24 | |||
| MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 1.04858M | $0.435 | $0.87 | 2026-04-22 | |||
| Qwen3.6 27Balibaba/qwen3.6-27b | 262.144K | $0.6 | $3.6 | 2026-04-22 | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| Kimi K2.6moonshotai/kimi-k2.6 | 262.144K | $0.95 | $4 | 2026-04-21 | |||
| Qwen3.6 35B-A3Balibaba/qwen3.6-35b-a3b | 262.144K | $0.248 | $1.485 | 2026-04-17 | |||
| GLM-5.1zhipuai/glm-5.1 | 200K | $1.4 | $4.4 | 2026-04-07 | |||
| Gemma 4 E2B ITgoogle/gemma-4-E2B-it | 131.072K | $0.04 | $0.08 | 2026-04-02 | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| Gemma 4 E4B ITgoogle/gemma-4-E4B-it | 131.072K | $0.02 | $0.1 | 2026-04-02 | |||
| MiniMax-M2.7minimax/MiniMax-M2.7 | 204.8K | $0.3 | $1.2 | 2026-03-18 | |||
| Mistral Small 4mistral/mistral-small-2603 | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Nemotron 3 Super 120B A12Bnvidia/nemotron-3-super-120b-a12b | 262.144K | $0.2 | $0.8 | 2026-03-11 | |||
| Qwen3.5 9Balibaba/qwen3.5-9b | 262.144K | $0.04 | $0.15 | 2026-02-23 | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 262.144K | $0.05 | $0.2 | 2025-12-15 | |||
| Devstral 2mistral/devstral-2512 | 262.144K | $0.4 | $2 | 2025-12-09 | |||
| Devstral Small 2mistral/devstral-small-2 | 262.144K | $0.1 | $0.3 | 2025-12-09 | |||
| Ministral 14Bmistral/ministral-14b | 262.144K | $0.2 | $0.2 | 2025-12-02 | |||
| DeepSeek-V3.1deepseek/deepseek-v3.1 | 131.072K | $0.19 | $0.71 | 2025-08-21 | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| GPT OSS 120Bopenai/gpt-oss-120b | 131.072K | $0.03 | $0.17 | 2025-08-05 | |||
| Qwen3 235B-A22B Instruct 2507alibaba/qwen3-235b-a22b-instruct-2507 | 262.144K | $0.069 | $0.455 | 2025-07-21 | |||
| Magistral Smallmistral/magistral-small-2506 | 131.072K | $0.5 | $1.5 | 2025-06-10 | |||
| Qwen3 32Balibaba/qwen3-32b | 131.072K | $0.7 | $2.8 | 2025-04 | |||
| DeepSeek-V3deepseek/deepseek-v3 | 131.072K | $0.27 | $1.12 | 2024-12-26 | |||
| Llama-3.3-70B-Instructmeta/llama-3.3-70b-instruct | 128K | $0.1 | $0.32 | 2024-12-06 | |||
| Qwen2.5-Coder-0.5Balibaba/qwen2.5-coder-0.5b | 32.768K | $0.1 | $0.1 | 2024-11-12 | |||
| Mistral Large 3mistral/mistral-large-2512 | 262.144K | $0.5 | $1.5 | 2024-11-01 | |||
| Ministral 3Bmistral/ministral-3b | 128K | $0.04 | $0.04 | 2024-10-16 | |||
| Ministral 8B Instructmistral/ministral-8b-instruct-2410 | 131.072K | $0.15 | $0.15 | 2024-10-16 | |||
| Llama-3.2-3Bmeta/llama-3.2-3b | 131.072K | $0.1 | $0.1 | 2024-09-25 | |||
| Llama-3.2-1Bmeta/llama-3.2-1b | 131.072K | $0.1 | $0.1 | 2024-09-25 | |||
| Pixtral 12Bmistral/pixtral-12b | 128K | $0.15 | $0.15 | 2024-09-01 | |||
| Mistral Nemomistral/mistral-nemo | 128K | $0.15 | $0.15 | 2024-07-01 | |||
| Codestral-22B-v0.1mistral/codestral-22b-v0.1 | 32.768K | $0.3 | $0.9 | 2024-05-29 |