Chat-tuned GPT model for conversational assistance, writing, and tool workflows
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Tiny GPT-5 lane for routing, extraction, classification, and bulk jobs
Small GPT-5 for responsive agents, coding help, and everyday automation
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Open GPT reasoning model for self-hosted agents and controllable deployments
Open GPT reasoning model for self-hosted agents and controllable deployments
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Hybrid-reasoning GLM release that made the 4.5 line broadly useful
Lighter GLM-4.5 variant for fast coding assistance and cheaper agents
Efficient Mistral model for fast chat, extraction, and production assistants
Fast Gemini workhorse for multimodal apps where latency and price matter
Google's proven reasoning model for coding, math, and multimodal analysis
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents
Fast o-series model for compact reasoning, coding, and tool use
Deliberate o-series reasoner for hard math, coding, and multi-step analysis
Flagship Nemotron model for high-throughput reasoning and complex agents
Dense open Qwen model for self-hosted chat, reasoning, and coding
Qwen reasoning model for deliberate problem solving, math, and coding
Smaller o-series reasoner for economical coding, math, and planning tasks
Popular open Llama workhorse for multilingual chat, coding, and self-hosting
O-series reasoning model for hard analysis, math, coding, and planning
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Qwen instruction model for multilingual chat, reasoning, and tool use
Web-grounded Sonar for multi-step research questions that need cited reasoning
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| GPT-5 Chat (latest)openai/gpt-5-chat-latest | 400K | $1.25 | $10 | 2025-08-07 | |||
| GPT-5 Nanoopenai/gpt-5-nano | 400K | $0.05 | $0.4 | 2025-08-07 | |||
| GPT-5 Miniopenai/gpt-5-mini | 400K | $0.25 | $2 | 2025-08-07 | |||
| Claude Opus 4.1anthropic/claude-opus-4-1-20250805 | 200K | $15 | $75 | 2025-08-05 | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| GPT OSS 120Bopenai/gpt-oss-120b | 131.072K | $0.03 | $0.17 | 2025-08-05 | |||
| Qwen Flashalibaba/qwen-flash | 1M | $0.05 | $0.4 | 2025-07-28 | |||
| GLM-4.5zhipuai/glm-4.5 | 131.072K | $0.6 | $2.2 | 2025-07-28 | |||
| GLM-4.5-Airzhipuai/glm-4.5-air | 131.072K | $0.2 | $1.1 | 2025-07-28 | |||
| Mistral Small 3.2mistral/mistral-small-2506 | 128K | $0.1 | $0.3 | 2025-06-20 | |||
| Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1.04858M | $0.3 | $2.5 | 2025-06-17 | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| Gemini 2.5 Flash-Litegoogle/gemini-2.5-flash-lite | 1.04858M | $0.1 | $0.4 | 2025-06-17 | |||
| o4-miniopenai/o4-mini | 200K | $1.1 | $4.4 | 2025-04-16 | |||
| o3openai/o3 | 200K | $2 | $8 | 2025-04-16 | |||
| Llama 3.1 Nemotron Ultra 253Bnvidia/llama-3.1-nemotron-ultra-253b | 128K | — | — | 2025-04-07 | |||
| Qwen3 32Balibaba/qwen3-32b | 131.072K | $0.7 | $2.8 | 2025-04 | |||
| QwQ Plusalibaba/qwq-plus | 131.072K | $0.8 | $2.4 | 2025-03-05 | |||
| o3-miniopenai/o3-mini | 200K | $1.1 | $4.4 | 2024-12-20 | |||
| Llama-3.3-70B-Instructmeta/llama-3.3-70b-instruct | 128K | $0.1 | $0.32 | 2024-12-06 | |||
| o1openai/o1 | 200K | $15 | $60 | 2024-12-05 | |||
| Qwen Turboalibaba/qwen-turbo | 1M | $0.05 | $0.2 | 2024-11-01 | |||
| Qwen Plusalibaba/qwen-plus | 1M | $0.4 | $1.2 | 2024-01-25 | |||
| Sonar Reasoning Properplexity/sonar-reasoning-pro | 128K | $2 | $8 | 2024-01-01 |