Nemotron middle tier for collaborative agents and high-volume reasoning workloads
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
More exact GPT-5.4 tier for demanding professional reasoning and agent tasks
Agent-ready GPT for coding and computer-use workflows at a lower cost
Qwen vision-language model for visual reasoning, documents, and agent tasks
Reasoning-first Gemini preview for agentic coding and complex problem solving
Claude workhorse for coding agents, careful analysis, and production cost control
Coding-optimized GPT model for repository edits, reviews, and agentic software work
StepFun flash lane for quick multimodal reasoning and coding assistance
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
Code-specialist GPT for repository edits, reviews, and long-running software agents
Reliable GPT generation for broad coding, writing, and tool-assisted product work
DeepSeek reasoning model for multi-step analysis, math, coding, and tools
Hybrid-reasoning DeepSeek model with thinking and non-thinking modes, sparse attention, and tool-use
Sharper GPT-5 generation for coding, product work, and tool-assisted tasks
Coding-optimized GPT model for repository edits, reviews, and agentic software work
Codex GPT for repository edits, code review, and practical software agents
Thinking Kimi model for slower research passes, planning, and hard technical questions
Flagship Qwen3 model for coding agents, complex reasoning, and tool use
Coding-optimized GPT model for repository edits, reviews, and agentic software work
Small GPT-5 for responsive agents, coding help, and everyday automation
Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows
Tiny GPT-5 lane for routing, extraction, classification, and bulk jobs
Open GPT reasoning model for self-hosted agents and controllable deployments
Open GPT reasoning model for self-hosted agents and controllable deployments
Google's proven reasoning model for coding, math, and multimodal analysis
Fast Gemini workhorse for multimodal apps where latency and price matter
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
Deliberate o-series reasoner for hard math, coding, and multi-step analysis
Fast o-series model for compact reasoning, coding, and tool use
Affordable GPT-4.1 lane for fast coding help and structured extraction
Long-lived GPT workhorse for coding, instruction following, and production apps
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
Smaller o-series reasoner for economical coding, math, and planning tasks
O-series reasoning model for hard analysis, math, coding, and planning
GPT model for general reasoning, writing, coding, and tool-assisted tasks
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Small omni GPT for cheap multimodal assistance and production-scale traffic
Omni-era GPT for multimodal chat, practical coding, and general assistants
Compact GPT model for low-latency assistance and high-volume workloads
Compact GPT model for low-latency assistance and high-volume workloads
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Nemotron 3 Super 120B A12Bnvidia/nemotron-3-super-120b-a12b | 262.144K | $0.2 | $0.8 | 2026-03-11 | |||
| GPT-5.4 Proopenai/gpt-5.4-pro | 1.05M | $30 | $180 | 2026-03-05 | |||
| GPT-5.4openai/gpt-5.4 | 1.05M | $2.5 | $15 | 2026-03-05 | |||
| Qwen3.5 27Balibaba/qwen3.5-27b | 262.144K | $0.3 | $2.4 | 2026-02-23 | |||
| Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | 1.04858M | $2 | $12 | 2026-02-19 | |||
| Claude Sonnet 4.6anthropic/claude-sonnet-4-6 | 1M | $3 | $15 | 2026-02-17 | |||
| GPT-5.3 Codexopenai/gpt-5.3-codex | 400K | $1.75 | $14 | 2026-02-05 | |||
| Step 3.5 Flashstepfun/step-3.5-flash | 256K | $0.1 | $0.3 | 2026-01-29 | |||
| Kimi K2.5moonshotai/kimi-k2.5 | 262.144K | $0.3 | $1.9 | 2026-01 | |||
| Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | 1.04858M | $0.5 | $3 | 2025-12-17 | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 262.144K | $0.05 | $0.2 | 2025-12-15 | |||
| GPT-5.2 Codexopenai/gpt-5.2-codex | 400K | $0.14 | $1.14 | 2025-12-11 | |||
| GPT-5.2openai/gpt-5.2 | 400K | $1.75 | $14 | 2025-12-11 | |||
| DeepSeek Reasonerdeepseek/deepseek-reasoner | 1M | $0.147 | $0.295 | 2025-12-01 | |||
| DeepSeek V3.2deepseek/deepseek-v3.2 | 128K | $0.18 | $0.35 | 2025-12-01 | |||
| GPT-5.1openai/gpt-5.1 | 400K | $1.25 | $10 | 2025-11-13 | |||
| GPT-5.1 Codex miniopenai/gpt-5.1-codex-mini | 400K | $0.22 | $1.8 | 2025-11-13 | |||
| GPT-5.1 Codexopenai/gpt-5.1-codex | 400K | $1.07 | $8.5 | 2025-11-13 | |||
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 262.144K | $0.4 | $2.5 | 2025-11-06 | |||
| Qwen3 Maxalibaba/qwen3-max | 262.144K | $1.2 | $6 | 2025-09-23 | |||
| GPT-5-Codexopenai/gpt-5-codex | 400K | $1.1 | $9 | 2025-09-15 | |||
| GPT-5 Miniopenai/gpt-5-mini | 400K | $0.25 | $2 | 2025-08-07 | |||
| GPT-5openai/gpt-5 | 400K | $1.25 | $10 | 2025-08-07 | |||
| GPT-5 Nanoopenai/gpt-5-nano | 400K | $0.05 | $0.4 | 2025-08-07 | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| GPT OSS 120Bopenai/gpt-oss-120b | 131.072K | $0.03 | $0.17 | 2025-08-05 | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1.04858M | $0.3 | $2.5 | 2025-06-17 | |||
| Mistral Medium 3mistral/mistral-medium-2505 | 131.072K | $0.4 | $2 | 2025-05-07 | |||
| o3openai/o3 | 200K | $2 | $8 | 2025-04-16 | |||
| o4-miniopenai/o4-mini | 200K | $1.1 | $4.4 | 2025-04-16 | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| GPT-4.1openai/gpt-4.1 | 1.04758M | $2 | $8 | 2025-04-14 | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| o3-miniopenai/o3-mini | 200K | $1.1 | $4.4 | 2024-12-20 | |||
| o1openai/o1 | 200K | $15 | $60 | 2024-12-05 | |||
| GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | 128K | $2.5 | $10 | 2024-11-20 | |||
| GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | 128K | $2.5 | $10 | 2024-08-06 | |||
| GPT-4o miniopenai/gpt-4o-mini | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| GPT-4oopenai/gpt-4o | 128K | $2.5 | $10 | 2024-05-13 | |||
| GPT-4 Turboopenai/gpt-4-turbo | 128K | $10 | $30 | 2023-11-06 | |||
| GPT-3.5-turboopenai/gpt-3.5-turbo | 16.385K | $0.5 | $1.5 | 2023-03-01 |