Chat-tuned GPT model for conversational assistance, writing, and tool workflows
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Small GPT-5 for responsive agents, coding help, and everyday automation
Tiny GPT-5 lane for routing, extraction, classification, and bulk jobs
Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows
Open GPT reasoning model for self-hosted agents and controllable deployments
Open GPT reasoning model for self-hosted agents and controllable deployments
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Cohere vision model for multilingual document analysis, OCR, and image understanding
Lighter GLM-4.5 variant for fast coding assistance and cheaper agents
Efficient GLM model for fast reasoning, coding, and agent workflows
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Qwen coding model for software agents, repository edits, and code reasoning
Hybrid-reasoning GLM release that made the 4.5 line broadly useful
Nemotron model for efficient reasoning, coding, and specialized AI agents
Hosted Qwen coder for software agents, repo edits, and long-context code
Updated large open Qwen3 MoE instruct model for multilingual chat, coding, and tool use
Open audio-language model for speech transcription, audio understanding, and voice-driven tool use
Open audio-language model for speech transcription, audio understanding, and voice-driven tool use
Instruct model with native audio input for speech understanding and tool use
Mistral coding agent model for repository tasks and software engineering workflows
Mistral coding agent model for repository tasks and software engineering workflows
Efficient Mistral model for fast chat, extraction, and production assistants
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents
Fast Gemini workhorse for multimodal apps where latency and price matter
Google's proven reasoning model for coding, math, and multimodal analysis
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
Open Mistral reasoning model for transparent step-by-step problem solving
High-effort o3 tier for difficult technical reasoning and careful answers
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Balanced Claude model for coding, analysis, agent workflows, and cost control
Balanced Claude model for coding, analysis, agent workflows, and cost control
Flagship model for demanding analysis, coding, and production agent workflows
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
Multimodal model for complex analysis, long-context understanding, tool use, and model distillation
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
Sparse MoE Qwen model with 3B active parameters for efficient chat and reasoning
Fast o-series model for compact reasoning, coding, and tool use
Deliberate o-series reasoner for hard math, coding, and multi-step analysis
Nemotron model for efficient reasoning, coding, and specialized AI agents
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
Affordable GPT-4.1 lane for fast coding help and structured extraction
Long-lived GPT workhorse for coding, instruction following, and production apps
Mistral vision-language model for image understanding and multimodal chat
Nemotron model for efficient reasoning, coding, and specialized AI agents
Flagship Nemotron model for high-throughput reasoning and complex agents
Open multimodal Llama for strong reasoning with efficient everyday serving
Open Llama with long-context vision for efficient multimodal agents
Dense open Qwen model for self-hosted chat, reasoning, and coding
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| GPT-5 Chat (latest)openai/gpt-5-chat-latest | 400K | $1.25 | $10 | 2025-08-07 | |||
| GPT-5 Miniopenai/gpt-5-mini | 400K | $0.25 | $2 | 2025-08-07 | |||
| GPT-5 Nanoopenai/gpt-5-nano | 400K | $0.05 | $0.4 | 2025-08-07 | |||
| GPT-5openai/gpt-5 | 400K | $1.25 | $10 | 2025-08-07 | |||
| GPT OSS 120Bopenai/gpt-oss-120b | 131.072K | $0.03 | $0.17 | 2025-08-05 | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| Claude Opus 4.1anthropic/claude-opus-4-1-20250805 | 200K | $15 | $75 | 2025-08-05 | |||
| Claude Opus 4.1 (latest)anthropic/claude-opus-4-1 | 200K | $15 | $75 | 2025-08-05 | |||
| Command A Visioncohere/command-a-vision-07-2025 | 128K | $2.5 | $10 | 2025-07-31 | |||
| GLM-4.5-Airzhipuai/glm-4.5-air | 131.072K | $0.2 | $1.1 | 2025-07-28 | |||
| GLM-4.5-Flashzhipuai/glm-4.5-flash | 131.072K | — | — | 2025-07-28 | |||
| Qwen Flashalibaba/qwen-flash | 1M | $0.05 | $0.4 | 2025-07-28 | |||
| Qwen3 Coder Flashalibaba/qwen3-coder-flash | 1M | $0.3 | $1.5 | 2025-07-28 | |||
| GLM-4.5zhipuai/glm-4.5 | 131.072K | $0.6 | $2.2 | 2025-07-28 | |||
| Llama 3.3 Nemotron Super 49B v1.5nvidia/llama-3.3-nemotron-super-49b-v1.5 | 131.072K | $0.4 | $0.4 | 2025-07-25 | |||
| Qwen3 Coder Plusalibaba/qwen3-coder-plus | 1.04858M | $1 | $5 | 2025-07-23 | |||
| Qwen3 235B-A22B Instruct 2507alibaba/qwen3-235b-a22b-instruct-2507 | 262.144K | $0.069 | $0.455 | 2025-07-21 | |||
| Voxtral Mini 3B 2507mistral/voxtral-mini-3b-2507 | 32.768K | $0.04 | $0.04 | 2025-07-15 | |||
| Voxtral Small 24B 2507mistral/voxtral-small-24b-2507 | 32.768K | $0.1 | $0.3 | 2025-07-15 | |||
| Voxtral Small (latest)mistral/voxtral-small-latest | 32K | $0.1 | $0.3 | 2025-07-15 | |||
| Devstral Smallmistral/devstral-small-2507 | 128K | $0.1 | $0.3 | 2025-07-10 | |||
| Devstral Mediummistral/devstral-medium-2507 | 128K | $0.4 | $2 | 2025-07-10 | |||
| Mistral Small 3.2mistral/mistral-small-2506 | 128K | $0.1 | $0.3 | 2025-06-20 | |||
| Gemini 2.5 Flash-Litegoogle/gemini-2.5-flash-lite | 1.04858M | $0.1 | $0.4 | 2025-06-17 | |||
| Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1.04858M | $0.3 | $2.5 | 2025-06-17 | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| Mistral Nemotronnvidia/mistral-nemotron | 128K | — | — | 2025-06-11 | |||
| Magistral Smallmistral/magistral-small-2506 | 131.072K | $0.5 | $1.5 | 2025-06-10 | |||
| o3-proopenai/o3-pro | 200K | $20 | $80 | 2025-06-10 | |||
| Claude Opus 4anthropic/claude-opus-4-20250514 | 200K | $15 | $75 | 2025-05-22 | |||
| Claude Opus 4 (latest)anthropic/claude-opus-4-0 | 200K | — | — | 2025-05-22 | |||
| Claude Sonnet 4 (latest)anthropic/claude-sonnet-4-0 | 200K | $2.898 | $14.493 | 2025-05-22 | |||
| Claude Sonnet 4anthropic/claude-sonnet-4-20250514 | 200K | $3 | $15 | 2025-05-22 | |||
| Solar Pro 2upstage/solar-pro2 | 65.536K | $0.15 | $0.6 | 2025-05-20 | |||
| Mistral Medium 3mistral/mistral-medium-2505 | 131.072K | $0.4 | $2 | 2025-05-07 | |||
| Nova Premieramazon/nova-premier | 1M | — | — | 2025-04-30 | |||
| Palmyra X5writer/palmyra-x5 | 1M | $0.6 | $6 | 2025-04-28 | |||
| Qwen3 30B A3Balibaba/qwen3-30b-a3b | 131.072K | $0.08 | $0.29 | 2025-04-28 | |||
| o4-miniopenai/o4-mini | 200K | $1.1 | $4.4 | 2025-04-16 | |||
| o3openai/o3 | 200K | $2 | $8 | 2025-04-16 | |||
| Llama 3.1 Nemotron 70B Instructnvidia/llama-3.1-nemotron-70b-instruct | 128K | — | — | 2025-04-15 | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| GPT-4.1openai/gpt-4.1 | 1.04758M | $2 | $8 | 2025-04-14 | |||
| Pixtral Large (25.02)mistral/pixtral-large-2502 | 128K | $1.993 | $5.978 | 2025-04-08 | |||
| Llama 3.3 Nemotron Super 49B v1nvidia/llama-3.3-nemotron-super-49b-v1 | 131.072K | — | — | 2025-04-07 | |||
| Llama 3.1 Nemotron Ultra 253Bnvidia/llama-3.1-nemotron-ultra-253b | 128K | — | — | 2025-04-07 | |||
| Llama 4 Maverick 17B Instructmeta/llama-4-maverick-17b-instruct | 1M | $0.14 | $0.59 | 2025-04-05 | |||
| Llama 4 Scout 17B Instructmeta/llama-4-scout-17b-instruct | 10M | $0.17 | $0.66 | 2025-04-05 | |||
| Qwen3 32Balibaba/qwen3-32b | 131.072K | $0.7 | $2.8 | 2025-04 |