Tiny GPT-5 lane for routing, extraction, classification, and bulk jobs
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Open GPT reasoning model for self-hosted agents and controllable deployments
Open GPT reasoning model for self-hosted agents and controllable deployments
Qwen coding model for software agents, repository edits, and code reasoning
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Hybrid-reasoning GLM release that made the 4.5 line broadly useful
Nemotron model for efficient reasoning, coding, and specialized AI agents
Hosted Qwen coder for software agents, repo edits, and long-context code
Updated large open Qwen3 MoE instruct model for multilingual chat, coding, and tool use
Open audio-language model for speech transcription, audio understanding, and voice-driven tool use
Open audio-language model for speech transcription, audio understanding, and voice-driven tool use
Instruct model with native audio input for speech understanding and tool use
Fast Gemini workhorse for multimodal apps where latency and price matter
Google's proven reasoning model for coding, math, and multimodal analysis
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents
High-effort o3 tier for difficult technical reasoning and careful answers
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Balanced Claude model for coding, analysis, agent workflows, and cost control
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
Sparse MoE Qwen model with 3B active parameters for efficient chat and reasoning
Fast o-series model for compact reasoning, coding, and tool use
Deliberate o-series reasoner for hard math, coding, and multi-step analysis
Affordable GPT-4.1 lane for fast coding help and structured extraction
Long-lived GPT workhorse for coding, instruction following, and production apps
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
Mistral vision-language model for image understanding and multimodal chat
Flagship Nemotron model for high-throughput reasoning and complex agents
Open multimodal Llama for strong reasoning with efficient everyday serving
Large open Qwen MoE for multilingual reasoning, coding, and tool use
Open Qwen coding heavyweight for repository reasoning and agentic engineering
Dense open Qwen model for self-hosted chat, reasoning, and coding
Smaller Qwen coder for efficient local agents and repo-level fixes
March 2025 checkpoint of DeepSeek-V3 with improved reasoning and coding
O-series reasoning model for hard analysis, math, coding, and planning
Mistral reasoning model for transparent analysis, math, and complex decisions
Efficient multimodal model for instruction following, coding, reasoning, and function calling
Cohere command model for multilingual enterprise agents, tools, and chat
Open multimodal Gemma instruction model for multilingual text generation and image understanding
Open multimodal Gemma instruction model for efficient text generation and image understanding
Largest open Gemma 3 instruction model for multilingual text generation and visual understanding
Balanced Claude model for coding, analysis, agent workflows, and cost control
Classic open reasoning model for transparent math, coding, and deliberate problem solving
Open DeepSeek MoE chat model for coding, math, and general reasoning
Smaller o-series reasoner for economical coding, math, and planning tasks
Earlier Gemini Flash workhorse for responsive multimodal apps and tool use
Low-latency Gemini model for high-volume multimodal and agent workloads
Popular open Llama workhorse for multilingual chat, coding, and self-hosting
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| GPT-5 Nanoopenai/gpt-5-nano | 400K | $0.05 | $0.4 | 2025-08-07 | |||
| GPT-5openai/gpt-5 | 400K | $1.25 | $10 | 2025-08-07 | |||
| Claude Opus 4.1 (latest)anthropic/claude-opus-4-1 | 200K | $15 | $75 | 2025-08-05 | |||
| Claude Opus 4.1anthropic/claude-opus-4-1-20250805 | 200K | $15 | $75 | 2025-08-05 | |||
| GPT OSS 120Bopenai/gpt-oss-120b | 131.072K | $0.03 | $0.17 | 2025-08-05 | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| Qwen3 Coder Flashalibaba/qwen3-coder-flash | 1M | $0.3 | $1.5 | 2025-07-28 | |||
| Qwen Flashalibaba/qwen-flash | 1M | $0.05 | $0.4 | 2025-07-28 | |||
| GLM-4.5zhipuai/glm-4.5 | 131.072K | $0.6 | $2.2 | 2025-07-28 | |||
| Llama 3.3 Nemotron Super 49B v1.5nvidia/llama-3.3-nemotron-super-49b-v1.5 | 131.072K | $0.4 | $0.4 | 2025-07-25 | |||
| Qwen3 Coder Plusalibaba/qwen3-coder-plus | 1.04858M | $1 | $5 | 2025-07-23 | |||
| Qwen3 235B-A22B Instruct 2507alibaba/qwen3-235b-a22b-instruct-2507 | 262.144K | $0.069 | $0.455 | 2025-07-21 | |||
| Voxtral Small 24B 2507mistral/voxtral-small-24b-2507 | 32.768K | $0.1 | $0.3 | 2025-07-15 | |||
| Voxtral Mini 3B 2507mistral/voxtral-mini-3b-2507 | 32.768K | $0.04 | $0.04 | 2025-07-15 | |||
| Voxtral Small (latest)mistral/voxtral-small-latest | 32K | $0.1 | $0.3 | 2025-07-15 | |||
| Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1.04858M | $0.3 | $2.5 | 2025-06-17 | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| Gemini 2.5 Flash-Litegoogle/gemini-2.5-flash-lite | 1.04858M | $0.1 | $0.4 | 2025-06-17 | |||
| o3-proopenai/o3-pro | 200K | $20 | $80 | 2025-06-10 | |||
| Claude Opus 4anthropic/claude-opus-4-20250514 | 200K | $15 | $75 | 2025-05-22 | |||
| Claude Sonnet 4 (latest)anthropic/claude-sonnet-4-0 | 200K | $2.898 | $14.493 | 2025-05-22 | |||
| Mistral Medium 3mistral/mistral-medium-2505 | 131.072K | $0.4 | $2 | 2025-05-07 | |||
| Qwen3 30B A3Balibaba/qwen3-30b-a3b | 131.072K | $0.08 | $0.29 | 2025-04-28 | |||
| o4-miniopenai/o4-mini | 200K | $1.1 | $4.4 | 2025-04-16 | |||
| o3openai/o3 | 200K | $2 | $8 | 2025-04-16 | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| GPT-4.1openai/gpt-4.1 | 1.04758M | $2 | $8 | 2025-04-14 | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| Pixtral Large (25.02)mistral/pixtral-large-2502 | 128K | $1.993 | $5.978 | 2025-04-08 | |||
| Llama 3.1 Nemotron Ultra 253Bnvidia/llama-3.1-nemotron-ultra-253b | 128K | — | — | 2025-04-07 | |||
| Llama 4 Maverick 17B Instructmeta/llama-4-maverick-17b-instruct | 1M | $0.14 | $0.59 | 2025-04-05 | |||
| Qwen3 235B-A22Balibaba/qwen3-235b-a22b | 131.072K | $0.7 | $2.8 | 2025-04 | |||
| Qwen3-Coder 480B-A35B Instructalibaba/qwen3-coder-480b-a35b-instruct | 262.144K | $1.5 | $7.5 | 2025-04 | |||
| Qwen3 32Balibaba/qwen3-32b | 131.072K | $0.7 | $2.8 | 2025-04 | |||
| Qwen3-Coder 30B-A3B Instructalibaba/qwen3-coder-30b-a3b-instruct | 262.144K | $0.45 | $2.25 | 2025-04 | |||
| DeepSeek V3 0324deepseek/deepseek-v3-0324 | 163.84K | $0.2 | $0.8 | 2025-03-24 | |||
| o1-proopenai/o1-pro | 200K | $150 | $600 | 2025-03-19 | |||
| Magistral Medium (latest)mistral/magistral-medium-latest | 128K | $2 | $5 | 2025-03-17 | |||
| Mistral Small 3.1 24Bmistral/mistral-small-3-1-24b-instruct-2503 | 128K | $0.106 | $0.318 | 2025-03-17 | |||
| Command Acohere/command-a-03-2025 | 256K | $2.5 | $10 | 2025-03-13 | |||
| Gemma 3 12B ITgoogle/gemma-3-12b-it | 131.072K | $0.05 | $0.1 | 2025-03-12 | |||
| Gemma 3 4B ITgoogle/gemma-3-4b-it | 131.072K | $0.04 | $0.08 | 2025-03-12 | |||
| Gemma 3 27B ITgoogle/gemma-3-27b-it | 131.072K | $0.08 | $0.16 | 2025-03-12 | |||
| Claude Sonnet 3.7anthropic/claude-3-7-sonnet-20250219 | 200K | $3 | $15 | 2025-02-19 | |||
| DeepSeek-R1deepseek/deepseek-r1 | 128K | $0.7 | $2.5 | 2025-01-20 | |||
| DeepSeek-V3deepseek/deepseek-v3 | 131.072K | $0.27 | $1.12 | 2024-12-26 | |||
| o3-miniopenai/o3-mini | 200K | $1.1 | $4.4 | 2024-12-20 | |||
| Gemini 2.0 Flashgoogle/gemini-2.0-flash | 1.04858M | $0.1 | $0.42 | 2024-12-11 | |||
| Gemini 2.0 Flash-Litegoogle/gemini-2.0-flash-lite | 1.04858M | $0.052 | $0.21 | 2024-12-11 | |||
| Llama-3.3-70B-Instructmeta/llama-3.3-70b-instruct | 128K | $0.1 | $0.32 | 2024-12-06 |