Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work
DeepSeek V4.1 Flash model for reasoning and agentic coding
DeepSeek V4 Pro snapshot with million-token context and support for thinking and non-thinking modes
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding
Hybrid-reasoning DeepSeek model with thinking and non-thinking modes, sparse attention, and tool-use
Open MoE flagship with million-token context for coding and long agent runs
Newer StepFun flash model for faster agents, coding, and multimodal prompts
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
Multimodal MoE reasoning model (276B total, 12B active) for text, image, and audio
Open Gemma instruction model for efficient chat and self-hosted deployments
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
Open MiMo model for multimodal coding agents and long-context automation
Thinking Kimi model for slower research passes, planning, and hard technical questions
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads
Open GPT reasoning model for self-hosted agents and controllable deployments
Classic open reasoning model for transparent math, coding, and deliberate problem solving
Reasoning-optimized 398B MoE agent model with extended thinking for long-horizon and multi-turn tool use
Open GPT reasoning model for self-hosted agents and controllable deployments
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| Kimi K2.6moonshotai/kimi-k2.6 | 45.1 | 262.144K | $0.95 | $4 | 2026-04-21 | |||
| Kimi K3moonshotai/kimi-k3 | 43.8 | 1.04858M | $3 | $15 | 2026-07-16 | |||
| DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash | 39.5 | 1M | $0.15 | $0.6 | 2026-09-10 | |||
| DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 | 36.3 | 1M | $0.442 | $0.884 | 2026-08-12 | |||
| Kimi K2.5moonshotai/kimi-k2.5 | 36.0 | 262.144K | $0.3 | $1.9 | 2026-01 | |||
| DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | 34.5 | 1M | $0.035 | $0.07 | 2026-07-31 | |||
| DeepSeek V3.2deepseek/deepseek-v3.2 | 32.6 | 128K | $0.18 | $0.35 | 2025-12-01 | |||
| DeepSeek V4 Prodeepseek/deepseek-v4-pro | 30.9 | 1M | $0.435 | $0.87 | 2026-04-24 | |||
| Step 3.7 Flashstepfun/step-3.7-flash | 30.9 | 256K | $0.185 | $1.11 | 2026-05-29 | |||
| MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 26.4 | 1.04858M | $0.435 | $0.87 | 2026-04-22 | |||
| Kimi K2.7 Codemoonshotai/kimi-k2.7-code | 26.3 | 262.144K | $0.95 | $4 | 2026-06-12 | |||
| Inkling Smallthinkingmachines/inkling-small | 26.1 | 1.04858M | $0.45 | $1.2 | 2026-07-30 | |||
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 26.1 | 262.144K | $0.042 | $0.22 | 2026-04-02 | |||
| Inklingthinkingmachines/inkling | 25.5 | 1.04858M | $1.87 | $4.68 | 2026-07-15 | |||
| DeepSeek V4 Flashdeepseek/deepseek-v4-flash | 24.8 | 1M | $0.15 | $0.6 | 2026-04-24 | |||
| Nemotron 3 Ultra 550B A55Bnvidia/nemotron-3-ultra-550b-a55b | 23.4 | 1M | $0.5 | $2.5 | 2026-06-04 | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 22.3 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 17.2 | 262.144K | $0.4 | $2.5 | 2025-11-06 | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 15.4 | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| Nemotron 3 Super 120B A12Bnvidia/nemotron-3-super-120b-a12b | 13.6 | 262.144K | $0.2 | $0.8 | 2026-03-11 | |||
| Nemotron 3.5 Lightning 30B A3Bnvidia/nemotron-3.5-lightning | 13.6 | 262.144K | $0.05 | $0.2 | 2026-08-11 | |||
| GPT OSS 120Bopenai/gpt-oss-120b | 12.3 | 131.072K | $0.03 | $0.17 | 2025-08-05 | |||
| DeepSeek-R1deepseek/deepseek-r1 | 11.4 | 128K | $0.7 | $2.5 | 2025-01-20 | |||
| Trinity Large Thinkingarcee-ai/trinity-large-thinking | 10.9 | 524.288K | $0.25 | $0.9 | 2026-04-01 | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 9.0 | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 8.9 | 262.144K | $0.05 | $0.2 | 2025-12-15 |