DeepSeek V4.1 Flash model for reasoning and agentic coding
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Native multimodal GLM model for efficient coding and long-horizon agent tasks
Dense 27B vision-language model for coding, agent tasks, and image and video understanding
Flagship GLM model for long-horizon coding, agents, and complex project delivery
Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work
Tencent Hy reasoning model for coding, instruction following, and agent tasks
Open flagship GLM for long-horizon coding agents and million-token context work
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
Lower-latency Kimi Code variant for interactive edits and coding-agent loops
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
MiniMax multimodal model for long-context coding, perception, and agent planning
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Open MoE flagship with million-token context for coding and long agent runs
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
Open MiMo model for multimodal coding agents and long-context automation
Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
Open multimodal Qwen MoE for local agents that need vision, audio, and code
Strong GLM coding model for agentic engineering, terminals, and repository generation
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Open Gemma instruction model for efficient chat and self-hosted deployments
Open MiniMax flagship for coding agents, office automation, and complex environments
Low-latency M2.7 variant for interactive coding plans and agent loops
Qwen instruction model for multilingual chat, reasoning, and tool use
Large open Qwen multimodal MoE for visual agents and long technical tasks
High-speed MiniMax model for low-latency coding and agent workflows
General GLM flagship for coding, analysis, and tool-heavy engineering workflows
Prior MiniMax coding model for agent workflows, office edits, and automation
Open-weight Qwen coding model for agents, repository edits, and multi-turn tool use
Budget GLM lane for fast coding help, routing, and everyday automation
Efficient GLM model for fast reasoning, coding, and agent workflows
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
Earlier MiniMax agent model for practical coding and productivity tasks
Mature GLM model for dependable coding, reasoning, and structured agent tasks
Mistral's coding-agent model for repository work, terminal tasks, and software fixes
GLM vision model for visual reasoning, documents, and multimodal agents
Hybrid-reasoning DeepSeek model with thinking and non-thinking modes, sparse attention, and tool-use
Thinking Kimi model for slower research passes, planning, and hard technical questions
Efficient open MiniMax model built for coding agents and tool-heavy workflows
Late GLM-4 workhorse for coding agents, reasoning, and structured tasks
Qwen vision-language thinking model for visual reasoning, documents, and agent tasks
Qwen vision-language instruct model for visual reasoning, documents, and agent tasks
Qwen instruction model for multilingual chat, reasoning, and tool use
Efficient Qwen thinking model for local reasoning, math, and coding agents
GLM vision model for visual reasoning, documents, and multimodal agents
Open GPT reasoning model for self-hosted agents and controllable deployments
Open GPT reasoning model for self-hosted agents and controllable deployments
Hybrid-reasoning GLM release that made the 4.5 line broadly useful
Lighter GLM-4.5 variant for fast coding assistance and cheaper agents
Mistral coding agent model for repository tasks and software engineering workflows
Efficient Mistral model for fast chat, extraction, and production assistants
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash | 1M | $0.15 | $0.6 | 2026-09-10 | |||
| GLM-5.3-Flashzhipuai/glm-5.3-flash | 1M | $0.075 | $0.25 | 2026-08-26 | |||
| Qwen3.8 27Balibaba/qwen3.8-27b | 262.144K | $0.1 | $0.4 | 2026-08-14 | |||
| GLM-5.3zhipuai/glm-5.3 | 1M | $1.4 | $4.4 | 2026-08-14 | |||
| Kimi K3moonshotai/kimi-k3 | 1.04858M | $3 | $15 | 2026-07-16 | |||
| Hy3tencent/hy3 | 256K | $0.066 | $0.26 | 2026-07-06 | |||
| GLM-5.2zhipuai/glm-5.2 | 1M | $1.4 | $4.4 | 2026-06-13 | |||
| Kimi K2.7 Codemoonshotai/kimi-k2.7-code | 262.144K | $0.95 | $4 | 2026-06-12 | |||
| Kimi K2.7 Code Highspeedmoonshotai/kimi-k2.7-code-highspeed | 262.144K | $1.9 | $8 | 2026-06-12 | |||
| Nemotron 3 Ultra 550B A55Bnvidia/nemotron-3-ultra-550b-a55b | 1M | $0.5 | $2.5 | 2026-06-04 | |||
| MiniMax-M3minimax/MiniMax-M3 | 1.04858M | $0.3 | $1.2 | 2026-06-01 | |||
| DeepSeek V4 Flashdeepseek/deepseek-v4-flash | 1M | $0.15 | $0.6 | 2026-04-24 | |||
| DeepSeek V4 Prodeepseek/deepseek-v4-pro | 1M | $0.435 | $0.87 | 2026-04-24 | |||
| MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 1.04858M | $0.435 | $0.87 | 2026-04-22 | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| Kimi K2.6moonshotai/kimi-k2.6 | 262.144K | $0.95 | $4 | 2026-04-21 | |||
| Qwen3.6 35B-A3Balibaba/qwen3.6-35b-a3b | 262.144K | $0.248 | $1.485 | 2026-04-17 | |||
| GLM-5.1zhipuai/glm-5.1 | 200K | $1.4 | $4.4 | 2026-04-07 | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 262.144K | $0.042 | $0.22 | 2026-04-02 | |||
| MiniMax-M2.7minimax/MiniMax-M2.7 | 204.8K | $0.3 | $1.2 | 2026-03-18 | |||
| MiniMax-M2.7-highspeedminimax/MiniMax-M2.7-highspeed | 204.8K | $0.6 | $2.4 | 2026-03-18 | |||
| Qwen3.5 9Balibaba/qwen3.5-9b | 262.144K | $0.04 | $0.15 | 2026-02-23 | |||
| Qwen3.5 397B-A17Balibaba/qwen3.5-397b-a17b | 262.144K | $0.6 | $3.6 | 2026-02-15 | |||
| MiniMax-M2.5-highspeedminimax/MiniMax-M2.5-highspeed | 204.8K | $0.6 | $2.4 | 2026-02-13 | |||
| GLM-5zhipuai/glm-5 | 204.8K | $1 | $3.2 | 2026-02-12 | |||
| MiniMax-M2.5minimax/MiniMax-M2.5 | 204.8K | $0.3 | $1.2 | 2026-02-12 | |||
| Qwen3 Coder Nextalibaba/qwen3-coder-next | 262.144K | $0.108 | $0.675 | 2026-02-03 | |||
| GLM-4.7-Flashzhipuai/glm-4.7-flash | 200K | $0.06 | $0.4 | 2026-01-19 | |||
| GLM-4.7-FlashXzhipuai/glm-4.7-flashx | 200K | $0.07 | $0.4 | 2026-01-19 | |||
| Kimi K2.5moonshotai/kimi-k2.5 | 262.144K | $0.3 | $1.9 | 2026-01 | |||
| MiniMax-M2.1minimax/MiniMax-M2.1 | 204.8K | $0.3 | $1.2 | 2025-12-23 | |||
| GLM-4.7zhipuai/glm-4.7 | 204.8K | $0.6 | $2.2 | 2025-12-22 | |||
| Devstral 2mistral/devstral-2512 | 262.144K | $0.4 | $2 | 2025-12-09 | |||
| GLM-4.6Vzhipuai/glm-4.6v | 128K | $0.3 | $0.9 | 2025-12-08 | |||
| DeepSeek V3.2deepseek/deepseek-v3.2 | 128K | $0.18 | $0.35 | 2025-12-01 | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 262.144K | $0.4 | $2.5 | 2025-11-06 | |||
| MiniMax-M2minimax/MiniMax-M2 | 204.8K | $0.3 | $1.2 | 2025-10-27 | |||
| GLM-4.6zhipuai/glm-4.6 | 204.8K | $0.6 | $2.2 | 2025-09-30 | |||
| Qwen3 VL 235B A22B Thinkingalibaba/qwen3-vl-235b-a22b-thinking | 131.072K | $0.4 | $4 | 2025-09-23 | |||
| Qwen3 VL 235B A22B Instructalibaba/qwen3-vl-235b-a22b-instruct | 131.072K | $0.2 | $0.88 | 2025-09-23 | |||
| Qwen3-Next 80B-A3B Instructalibaba/qwen3-next-80b-a3b-instruct | 131.072K | $0.5 | $2 | 2025-09 | |||
| Qwen3-Next 80B-A3B (Thinking)alibaba/qwen3-next-80b-a3b-thinking | 131.072K | $0.5 | $6 | 2025-09 | |||
| GLM-4.5Vzhipuai/glm-4.5v | 64K | $0.6 | $1.8 | 2025-08-11 | |||
| GPT OSS 120Bopenai/gpt-oss-120b | 131.072K | $0.03 | $0.17 | 2025-08-05 | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| GLM-4.5zhipuai/glm-4.5 | 131.072K | $0.6 | $2.2 | 2025-07-28 | |||
| GLM-4.5-Airzhipuai/glm-4.5-air | 131.072K | $0.2 | $1.1 | 2025-07-28 | |||
| Devstral Smallmistral/devstral-small-2507 | 128K | $0.1 | $0.3 | 2025-07-10 | |||
| Mistral Small 3.2mistral/mistral-small-2506 | 128K | $0.1 | $0.3 | 2025-06-20 |