Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work
DeepSeek V4.1 Flash model for reasoning and agentic coding
DeepSeek V4 Pro snapshot with million-token context and support for thinking and non-thinking modes
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding
Hybrid-reasoning DeepSeek model with thinking and non-thinking modes, sparse attention, and tool-use
Open MoE flagship with million-token context for coding and long agent runs
Newer StepFun flash model for faster agents, coding, and multimodal prompts
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
Open Gemma instruction model for efficient chat and self-hosted deployments
Multimodal MoE reasoning model (276B total, 12B active) for text, image, and audio
Tencent Hy reasoning model for coding, instruction following, and agent tasks
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
Open MiMo model for multimodal coding agents and long-context automation
Thinking Kimi model for slower research passes, planning, and hard technical questions
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads
Open GPT reasoning model for self-hosted agents and controllable deployments
Classic open reasoning model for transparent math, coding, and deliberate problem solving
Reasoning-optimized 398B MoE agent model with extended thinking for long-horizon and multi-turn tool use
Open GPT reasoning model for self-hosted agents and controllable deployments
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
Largest open Gemma 3 instruction model for multilingual text generation and visual understanding
Open multimodal Gemma instruction model for multilingual text generation and image understanding
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| Kimi K2.6moonshotai/kimi-k2.6 | 45.1 | 262.144K | $0.95 | $4 | 2026-04-21 | |||
| Kimi K3moonshotai/kimi-k3 | 43.8 | 1.04858M | $3 | $15 | 2026-07-16 | |||
| DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash | 39.5 | 1M | $0.15 | $0.6 | 2026-09-10 | |||
| DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 | 36.3 | 1M | $0.442 | $0.884 | 2026-08-12 | |||
| Kimi K2.5moonshotai/kimi-k2.5 | 36.0 | 262.144K | $0.3 | $1.9 | 2026-01 | |||
| DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | 34.5 | 1M | $0.05 | $0.16 | 2026-07-31 | |||
| DeepSeek V3.2deepseek/deepseek-v3.2 | 32.6 | 128K | $0.18 | $0.35 | 2025-12-01 | |||
| DeepSeek V4 Prodeepseek/deepseek-v4-pro | 30.9 | 1M | $0.435 | $0.87 | 2026-04-24 | |||
| Step 3.7 Flashstepfun/step-3.7-flash | 30.9 | 256K | $0.185 | $1.11 | 2026-05-29 | |||
| MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 26.4 | 1.04858M | $0.435 | $0.87 | 2026-04-22 | |||
| Kimi K2.7 Codemoonshotai/kimi-k2.7-code | 26.3 | 262.144K | $0.95 | $4 | 2026-06-12 | |||
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 26.1 | 262.144K | $0.042 | $0.22 | 2026-04-02 | |||
| Inkling Smallthinkingmachines/inkling-small | 26.1 | 1.04858M | $0.45 | $1.2 | 2026-07-30 | |||
| Hy3 previewtencent/hy3-preview | 25.8 | 256K | $0.066 | $0.26 | 2026-04-20 | |||
| Inklingthinkingmachines/inkling | 25.5 | 1.04858M | $1.87 | $4.68 | 2026-07-15 | |||
| DeepSeek V4 Flashdeepseek/deepseek-v4-flash | 24.8 | 1M | $0.15 | $0.6 | 2026-04-24 | |||
| Nemotron 3 Ultra 550B A55Bnvidia/nemotron-3-ultra-550b-a55b | 23.4 | 1M | $0.5 | $2.5 | 2026-06-04 | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 22.3 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 17.2 | 262.144K | $0.4 | $2.5 | 2025-11-06 | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 15.4 | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| Nemotron 3 Super 120B A12Bnvidia/nemotron-3-super-120b-a12b | 13.6 | 262.144K | $0.2 | $0.8 | 2026-03-11 | |||
| Nemotron 3.5 Lightning 30B A3Bnvidia/nemotron-3.5-lightning | 13.6 | 262.144K | $0.05 | $0.2 | 2026-08-11 | |||
| GPT OSS 120Bopenai/gpt-oss-120b | 12.3 | 131.072K | $0.03 | $0.17 | 2025-08-05 | |||
| DeepSeek-R1deepseek/deepseek-r1 | 11.4 | 128K | $0.7 | $2.5 | 2025-01-20 | |||
| Trinity Large Thinkingarcee-ai/trinity-large-thinking | 10.9 | 524.288K | $0.25 | $0.9 | 2026-04-01 | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 9.0 | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 8.9 | 262.144K | $0.05 | $0.2 | 2025-12-15 | |||
| Gemma 3 27B ITgoogle/gemma-3-27b-it | 4.9 | 131.072K | $0.08 | $0.16 | 2025-03-12 | |||
| Gemma 3 12B ITgoogle/gemma-3-12b-it | 3.8 | 131.072K | $0.05 | $0.1 | 2025-03-12 |