DeepSeek V4.1 Flash model for reasoning and agentic coding
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Dense 27B vision-language model for coding, agent tasks, and image and video understanding
Muse Glimmer is a 30-billion-parameter open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark for always-on local agents, tool use, coding, and image understanding.
Multimodal MoE reasoning model (276B total, 12B active) for text, image, and audio
Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
Lower-latency Kimi Code variant for interactive edits and coding-agent loops
Balanced Mistral model for enterprise assistants, multilingual work, and tools
Open Nemotron omni model combining reasoning with text, vision, and audio
Qwen vision-language model for visual reasoning, documents, and agent tasks
Open MiMo model for multimodal coding agents and long-context automation
Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Open Gemma instruction model for efficient chat and self-hosted deployments
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen instruction model for multilingual chat, reasoning, and tool use
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
Qwen vision-language instruct model for visual reasoning, documents, and agent tasks
Qwen vision-language thinking model for visual reasoning, documents, and agent tasks
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash | 1M | $0.15 | $0.6 | 2026-09-10 | |||
| Qwen3.8 27Balibaba/qwen3.8-27b | 262.144K | $0.1 | $0.4 | 2026-08-14 | |||
| Muse Glimmer 30Bmeta/muse-glimmer-30b | 131.072K | $0.2 | $0.8 | 2026-08-10 | |||
| Inkling Smallthinkingmachines/inkling-small | 1.04858M | $0.45 | $1.2 | 2026-07-30 | |||
| Kimi K3moonshotai/kimi-k3 | 1.04858M | $3 | $15 | 2026-07-16 | |||
| Inklingthinkingmachines/inkling | 1.04858M | $1.87 | $4.68 | 2026-07-15 | |||
| Kimi K2.7 Codemoonshotai/kimi-k2.7-code | 262.144K | $0.95 | $4 | 2026-06-12 | |||
| Kimi K2.7 Code Highspeedmoonshotai/kimi-k2.7-code-highspeed | 262.144K | $1.9 | $8 | 2026-06-12 | |||
| Mistral Medium 3.5mistral/mistral-medium-2604 | 262.144K | $1.5 | $7.5 | 2026-04-29 | |||
| Nemotron 3 Nano Omni 30B A3B Reasoningnvidia/nemotron-3-nano-omni-30b-a3b-reasoning | 256K | $0.2 | $0.8 | 2026-04-28 | |||
| Qwen3.6 27Balibaba/qwen3.6-27b | 262.144K | $0.6 | $3.6 | 2026-04-22 | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| Kimi K2.6moonshotai/kimi-k2.6 | 262.144K | $0.95 | $4 | 2026-04-21 | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 262.144K | $0.042 | $0.22 | 2026-04-02 | |||
| Qwen3.5 27Balibaba/qwen3.5-27b | 262.144K | $0.3 | $2.4 | 2026-02-23 | |||
| Qwen3.5 9Balibaba/qwen3.5-9b | 262.144K | $0.04 | $0.15 | 2026-02-23 | |||
| Kimi K2.5moonshotai/kimi-k2.5 | 262.144K | $0.3 | $1.9 | 2026-01 | |||
| Qwen3 VL 235B A22B Instructalibaba/qwen3-vl-235b-a22b-instruct | 131.072K | $0.2 | $0.88 | 2025-09-23 | |||
| Qwen3 VL 235B A22B Thinkingalibaba/qwen3-vl-235b-a22b-thinking | 131.072K | $0.4 | $4 | 2025-09-23 |