Open Gemma instruction model for efficient chat and self-hosted deployments
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Earlier Qwen multimodal workhorse for million-token agent and document tasks
Low-latency M2.7 variant for interactive coding plans and agent loops
Open MiniMax flagship for coding agents, office automation, and complex environments
Strong small GPT for coding subagents, quick tool use, and high-volume work
Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation
Grok model for agentic tool use, reasoning, coding, and live assistance
Reasoning Grok for document-heavy analysis and long-horizon tool use
More exact GPT-5.4 tier for demanding professional reasoning and agent tasks
Agent-ready GPT for coding and computer-use workflows at a lower cost
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
Qwen instruction model for multilingual chat, reasoning, and tool use
Reasoning-first Gemini preview for agentic coding and complex problem solving
Claude workhorse for coding agents, careful analysis, and production cost control
Large open Qwen multimodal MoE for visual agents and long technical tasks
High-speed MiniMax model for low-latency coding and agent workflows
Prior MiniMax coding model for agent workflows, office edits, and automation
General GLM flagship for coding, analysis, and tool-heavy engineering workflows
Coding-optimized GPT model for repository edits, reviews, and agentic software work
High-end Claude for difficult coding, planning, and slower expert reasoning
Open-weight Qwen coding model for agents, repository edits, and multi-turn tool use
Budget GLM lane for fast coding help, routing, and everyday automation
Efficient GLM model for fast reasoning, coding, and agent workflows
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
Earlier MiniMax agent model for practical coding and productivity tasks
Mature GLM model for dependable coding, reasoning, and structured agent tasks
New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
Higher-accuracy GPT-5.2 variant for tougher reasoning and review workflows
Code-specialist GPT for repository edits, reviews, and long-running software agents
Reliable GPT generation for broad coding, writing, and tool-assisted product work
Mistral's coding-agent model for repository work, terminal tasks, and software fixes
GLM vision model for visual reasoning, documents, and multimodal agents
Hybrid-reasoning DeepSeek model with thinking and non-thinking modes, sparse attention, and tool-use
Coding-optimized GPT model for repository edits, reviews, and agentic software work
Sharper GPT-5 generation for coding, product work, and tool-assisted tasks
Codex GPT for repository edits, code review, and practical software agents
Thinking Kimi model for slower research passes, planning, and hard technical questions
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Efficient open MiniMax model built for coding agents and tool-heavy workflows
Fast Claude lane for lightweight agents, office tasks, and responsive chat
Fast Claude model for responsive assistance, classification, and lightweight agents
Higher-accuracy GPT-5 tier for tough analysis, coding reviews, and planning
Late GLM-4 workhorse for coding agents, reasoning, and structured tasks
Balanced Claude model for coding, analysis, agent workflows, and cost control
Balanced Claude model for coding, analysis, agent workflows, and cost control
Qwen vision-language thinking model for visual reasoning, documents, and agent tasks
Flagship Qwen3 model for coding agents, complex reasoning, and tool use
Qwen vision-language instruct model for visual reasoning, documents, and agent tasks
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 262.144K | $0.042 | $0.22 | 2026-04-02 | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| Qwen3.6 Plusalibaba/qwen3.6-plus | 1M | $0.5 | $3 | 2026-04-02 | |||
| MiniMax-M2.7-highspeedminimax/MiniMax-M2.7-highspeed | 204.8K | $0.6 | $2.4 | 2026-03-18 | |||
| MiniMax-M2.7minimax/MiniMax-M2.7 | 204.8K | $0.3 | $1.2 | 2026-03-18 | |||
| GPT-5.4 miniopenai/gpt-5.4-mini | 400K | $0.75 | $4.5 | 2026-03-17 | |||
| GPT-5.4 nanoopenai/gpt-5.4-nano | 400K | $0.2 | $1.25 | 2026-03-17 | |||
| Grok 4.20 (Non-Reasoning)xai/grok-4.20-0309-non-reasoning | 1M | $1.25 | $2.5 | 2026-03-09 | |||
| Grok 4.20 (Reasoning)xai/grok-4.20-0309-reasoning | 1M | $1.25 | $2.5 | 2026-03-09 | |||
| GPT-5.4 Proopenai/gpt-5.4-pro | 1.05M | $30 | $180 | 2026-03-05 | |||
| GPT-5.4openai/gpt-5.4 | 1.05M | $2.5 | $15 | 2026-03-05 | |||
| GPT-5.3 Chat (latest)openai/gpt-5.3-chat-latest | 128K | $1.75 | $14 | 2026-03-03 | |||
| Qwen3.5 9Balibaba/qwen3.5-9b | 262.144K | $0.04 | $0.15 | 2026-02-23 | |||
| Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | 1.04858M | $2 | $12 | 2026-02-19 | |||
| Claude Sonnet 4.6anthropic/claude-sonnet-4-6 | 1M | $3 | $15 | 2026-02-17 | |||
| Qwen3.5 397B-A17Balibaba/qwen3.5-397b-a17b | 262.144K | $0.6 | $3.6 | 2026-02-15 | |||
| MiniMax-M2.5-highspeedminimax/MiniMax-M2.5-highspeed | 204.8K | $0.6 | $2.4 | 2026-02-13 | |||
| MiniMax-M2.5minimax/MiniMax-M2.5 | 204.8K | $0.3 | $1.2 | 2026-02-12 | |||
| GLM-5zhipuai/glm-5 | 204.8K | $1 | $3.2 | 2026-02-12 | |||
| GPT-5.3 Codexopenai/gpt-5.3-codex | 400K | $1.75 | $14 | 2026-02-05 | |||
| Claude Opus 4.6anthropic/claude-opus-4-6 | 1M | $5 | $25 | 2026-02-05 | |||
| Qwen3 Coder Nextalibaba/qwen3-coder-next | 262.144K | $0.108 | $0.675 | 2026-02-03 | |||
| GLM-4.7-Flashzhipuai/glm-4.7-flash | 200K | $0.06 | $0.4 | 2026-01-19 | |||
| GLM-4.7-FlashXzhipuai/glm-4.7-flashx | 200K | $0.07 | $0.4 | 2026-01-19 | |||
| Kimi K2.5moonshotai/kimi-k2.5 | 262.144K | $0.3 | $1.9 | 2026-01 | |||
| MiniMax-M2.1minimax/MiniMax-M2.1 | 204.8K | $0.3 | $1.2 | 2025-12-23 | |||
| GLM-4.7zhipuai/glm-4.7 | 204.8K | $0.6 | $2.2 | 2025-12-22 | |||
| Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | 1.04858M | $0.5 | $3 | 2025-12-17 | |||
| GPT-5.2 Chatopenai/gpt-5.2-chat-latest | 128K | $1.75 | $14 | 2025-12-11 | |||
| GPT-5.2 Proopenai/gpt-5.2-pro | 400K | $21 | $168 | 2025-12-11 | |||
| GPT-5.2 Codexopenai/gpt-5.2-codex | 400K | $0.14 | $1.14 | 2025-12-11 | |||
| GPT-5.2openai/gpt-5.2 | 400K | $1.75 | $14 | 2025-12-11 | |||
| Devstral 2mistral/devstral-2512 | 262.144K | $0.4 | $2 | 2025-12-09 | |||
| GLM-4.6Vzhipuai/glm-4.6v | 128K | $0.3 | $0.9 | 2025-12-08 | |||
| DeepSeek V3.2deepseek/deepseek-v3.2 | 128K | $0.18 | $0.35 | 2025-12-01 | |||
| GPT-5.1 Codex miniopenai/gpt-5.1-codex-mini | 400K | $0.22 | $1.8 | 2025-11-13 | |||
| GPT-5.1openai/gpt-5.1 | 400K | $1.25 | $10 | 2025-11-13 | |||
| GPT-5.1 Codexopenai/gpt-5.1-codex | 400K | $1.07 | $8.5 | 2025-11-13 | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 262.144K | $0.4 | $2.5 | 2025-11-06 | |||
| Claude Opus 4.5anthropic/claude-opus-4-5-20251101 | 200K | $5 | $25 | 2025-11-01 | |||
| MiniMax-M2minimax/MiniMax-M2 | 204.8K | $0.3 | $1.2 | 2025-10-27 | |||
| Claude Haiku 4.5 (latest)anthropic/claude-haiku-4-5 | 200K | $1 | $5 | 2025-10-15 | |||
| Claude Haiku 4.5anthropic/claude-haiku-4-5-20251001 | 200K | $1 | $5 | 2025-10-15 | |||
| GPT-5 Proopenai/gpt-5-pro | 400K | $15 | $120 | 2025-10-06 | |||
| GLM-4.6zhipuai/glm-4.6 | 204.8K | $0.6 | $2.2 | 2025-09-30 | |||
| Claude Sonnet 4.5anthropic/claude-sonnet-4-5-20250929 | 200K | $3 | $15 | 2025-09-29 | |||
| Claude Sonnet 4.5 (latest)anthropic/claude-sonnet-4-5 | 200K | $3 | $15 | 2025-09-29 | |||
| Qwen3 VL 235B A22B Thinkingalibaba/qwen3-vl-235b-a22b-thinking | 131.072K | $0.4 | $4 | 2025-09-23 | |||
| Qwen3 Maxalibaba/qwen3-max | 262.144K | $1.2 | $6 | 2025-09-23 | |||
| Qwen3 VL 235B A22B Instructalibaba/qwen3-vl-235b-a22b-instruct | 131.072K | $0.2 | $0.88 | 2025-09-23 |