Flagship Claude model for deep reasoning, coding, and long-horizon agents
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Qwen coding model for software agents, repository edits, and code reasoning
Hosted Qwen coder for software agents, repo edits, and long-context code
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents
Fast Gemini workhorse for multimodal apps where latency and price matter
Google's proven reasoning model for coding, math, and multimodal analysis
Fast o-series model for compact reasoning, coding, and tool use
Deliberate o-series reasoner for hard math, coding, and multi-step analysis
Affordable GPT-4.1 lane for fast coding help and structured extraction
Long-lived GPT workhorse for coding, instruction following, and production apps
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
Open Llama with long-context vision for efficient multimodal agents
Open multimodal Llama for strong reasoning with efficient everyday serving
Open Qwen coding heavyweight for repository reasoning and agentic engineering
Smaller Qwen coder for efficient local agents and repo-level fixes
Smaller o-series reasoner for economical coding, math, and planning tasks
O-series reasoning model for hard analysis, math, coding, and planning
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Qwen instruction model for multilingual chat, reasoning, and tool use
Deeper Sonar search model with broader retrieval and stronger synthesis
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Claude Opus 4.1anthropic/claude-opus-4-1-20250805 | 200K | $15 | $75 | 2025-08-05 | |||
| Qwen Flashalibaba/qwen-flash | 1M | $0.05 | $0.4 | 2025-07-28 | |||
| Qwen3 Coder Flashalibaba/qwen3-coder-flash | 1M | $0.3 | $1.5 | 2025-07-28 | |||
| Qwen3 Coder Plusalibaba/qwen3-coder-plus | 1.04858M | $1 | $5 | 2025-07-23 | |||
| Gemini 2.5 Flash-Litegoogle/gemini-2.5-flash-lite | 1.04858M | $0.1 | $0.4 | 2025-06-17 | |||
| Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1.04858M | $0.3 | $2.5 | 2025-06-17 | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| o4-miniopenai/o4-mini | 200K | $1.1 | $4.4 | 2025-04-16 | |||
| o3openai/o3 | 200K | $2 | $8 | 2025-04-16 | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| GPT-4.1openai/gpt-4.1 | 1.04758M | $2 | $8 | 2025-04-14 | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| Llama 4 Scout 17B Instructmeta/llama-4-scout-17b-instruct | 10M | $0.17 | $0.66 | 2025-04-05 | |||
| Llama 4 Maverick 17B Instructmeta/llama-4-maverick-17b-instruct | 1M | $0.14 | $0.59 | 2025-04-05 | |||
| Qwen3-Coder 480B-A35B Instructalibaba/qwen3-coder-480b-a35b-instruct | 262.144K | $1.5 | $7.5 | 2025-04 | |||
| Qwen3-Coder 30B-A3B Instructalibaba/qwen3-coder-30b-a3b-instruct | 262.144K | $0.45 | $2.25 | 2025-04 | |||
| o3-miniopenai/o3-mini | 200K | $1.1 | $4.4 | 2024-12-20 | |||
| o1openai/o1 | 200K | $15 | $60 | 2024-12-05 | |||
| Qwen Turboalibaba/qwen-turbo | 1M | $0.05 | $0.2 | 2024-11-01 | |||
| Mistral Large (latest)mistral/mistral-large-latest | 262.144K | $0.5 | $1.5 | 2024-11-01 | |||
| Qwen Plusalibaba/qwen-plus | 1M | $0.4 | $1.2 | 2024-01-25 | |||
| Sonar Properplexity/sonar-pro | 200K | $3 | $15 | 2024-01-01 |