High-effort o3 tier for difficult technical reasoning and careful answers
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Balanced Claude model for coding, analysis, agent workflows, and cost control
Balanced Claude model for coding, analysis, agent workflows, and cost control
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
Flagship model for demanding analysis, coding, and production agent workflows
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
Multimodal model for complex analysis, long-context understanding, tool use, and model distillation
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
OpenAI image model for production generation, edits, and brand-safe visual workflows
Deliberate o-series reasoner for hard math, coding, and multi-step analysis
Fast o-series model for compact reasoning, coding, and tool use
Long-lived GPT workhorse for coding, instruction following, and production apps
Affordable GPT-4.1 lane for fast coding help and structured extraction
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
Mistral vision-language model for image understanding and multimodal chat
O-series reasoning model for hard analysis, math, coding, and planning
Mistral reasoning model for transparent analysis, math, and complex decisions
Qwen reasoning model for deliberate problem solving, math, and coding
Balanced Claude model for coding, analysis, agent workflows, and cost control
MiniMax text-to-image generation model with reference-image support
Sonar search model for autonomous research and citation-backed long-form reports
Qwen omni model for text, vision, audio, and multimodal agent tasks
Smaller o-series reasoner for economical coding, math, and planning tasks
Earlier Gemini Flash workhorse for responsive multimodal apps and tool use
Low-latency Gemini model for high-volume multimodal and agent workloads
O-series reasoning model for hard analysis, math, coding, and planning
Efficient model for low-latency assistance, extraction, and routine automation
Flagship model for demanding analysis, coding, and production agent workflows
Efficient model for low-latency assistance, extraction, and routine automation
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Balanced Claude model for coding, analysis, agent workflows, and cost control
Fast Claude model for responsive assistance, classification, and lightweight agents
Enterprise language model for workflow automation, coding, data analysis, and tool use
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Small omni GPT for cheap multimodal assistance and production-scale traffic
Research model for long-horizon investigation, synthesis, and analytical reports
Research model for long-horizon investigation, synthesis, and analytical reports
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Omni-era GPT for multimodal chat, practical coding, and general assistants
Qwen vision-language model for visual reasoning, documents, and agent tasks
Flagship Qwen model for complex reasoning, coding, and agentic workflows
Legacy model retained for compatibility with older integrations
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen vision-language model for visual reasoning, documents, and agent tasks
Deeper Sonar search model with broader retrieval and stronger synthesis
Web-grounded Sonar for multi-step research questions that need cited reasoning
Fast web-grounded Sonar for current answers, citations, and lightweight retrieval
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| o3-proopenai/o3-pro | 200K | $20 | $80 | 2025-06-10 | |||
| Claude Opus 4anthropic/claude-opus-4-20250514 | 200K | $15 | $75 | 2025-05-22 | |||
| Claude Opus 4 (latest)anthropic/claude-opus-4-0 | 200K | — | — | 2025-05-22 | |||
| Claude Sonnet 4anthropic/claude-sonnet-4-20250514 | 200K | $3 | $15 | 2025-05-22 | |||
| Claude Sonnet 4 (latest)anthropic/claude-sonnet-4-0 | 200K | $2.898 | $14.493 | 2025-05-22 | |||
| Gemini Embedding 001google/gemini-embedding-001 | 2.048K | $0.15 | — | 2025-05-20 | |||
| Solar Pro 2upstage/solar-pro2 | 65.536K | $0.15 | $0.6 | 2025-05-20 | |||
| Mistral Medium 3mistral/mistral-medium-2505 | 131.072K | $0.4 | $2 | 2025-05-07 | |||
| Nova Premieramazon/nova-premier | 1M | — | — | 2025-04-30 | |||
| Palmyra X5writer/palmyra-x5 | 1M | $0.6 | $6 | 2025-04-28 | |||
| GPT-Image-1openai/gpt-image-1 | Not documented | $5 | $40 | 2025-04-24 | |||
| o3openai/o3 | 200K | $2 | $8 | 2025-04-16 | |||
| o4-miniopenai/o4-mini | 200K | $1.1 | $4.4 | 2025-04-16 | |||
| GPT-4.1openai/gpt-4.1 | 1.04758M | $2 | $8 | 2025-04-14 | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| Pixtral Large (25.02)mistral/pixtral-large-2502 | 128K | $1.993 | $5.978 | 2025-04-08 | |||
| o1-proopenai/o1-pro | 200K | $150 | $600 | 2025-03-19 | |||
| Magistral Medium (latest)mistral/magistral-medium-latest | 128K | $2 | $5 | 2025-03-17 | |||
| QwQ Plusalibaba/qwq-plus | 131.072K | $0.8 | $2.4 | 2025-03-05 | |||
| Claude Sonnet 3.7anthropic/claude-3-7-sonnet-20250219 | 200K | $3 | $15 | 2025-02-19 | |||
| MiniMax image-01minimax/image-01 | Not documented | — | — | 2025-02-15 | |||
| Sonar Deep Researchperplexity/sonar-deep-research | 128K | $2 | $8 | 2025-02-01 | |||
| Qwen-Omni Turboalibaba/qwen-omni-turbo | 32.768K | $0.07 | $0.27 | 2025-01-19 | |||
| o3-miniopenai/o3-mini | 200K | $1.1 | $4.4 | 2024-12-20 | |||
| Gemini 2.0 Flashgoogle/gemini-2.0-flash | 1.04858M | $0.1 | $0.42 | 2024-12-11 | |||
| Gemini 2.0 Flash-Litegoogle/gemini-2.0-flash-lite | 1.04858M | $0.052 | $0.21 | 2024-12-11 | |||
| o1openai/o1 | 200K | $15 | $60 | 2024-12-05 | |||
| Nova Microamazon/nova-micro | 128K | $0.035 | $0.14 | 2024-12-03 | |||
| Nova Proamazon/nova-pro | 300K | $0.8 | $3.2 | 2024-12-03 | |||
| Nova Liteamazon/nova-lite | 300K | $0.06 | $0.24 | 2024-12-03 | |||
| GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | 128K | $2.5 | $10 | 2024-11-20 | |||
| Qwen Turboalibaba/qwen-turbo | 1M | $0.05 | $0.2 | 2024-11-01 | |||
| Claude Sonnet 3.5 v2anthropic/claude-3-5-sonnet-20241022 | 200K | — | — | 2024-10-22 | |||
| Claude Haiku 3.5anthropic/claude-3-5-haiku-20241022 | 200K | $0.8 | $4 | 2024-10-22 | |||
| Palmyra X4writer/palmyra-x4 | 128K | $2.5 | $10 | 2024-10-09 | |||
| GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | 128K | $2.5 | $10 | 2024-08-06 | |||
| GPT-4o miniopenai/gpt-4o-mini | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| o3-deep-researchopenai/o3-deep-research | 200K | $9 | $36 | 2024-06-26 | |||
| o4-mini-deep-researchopenai/o4-mini-deep-research | 200K | $1.8 | $7.2 | 2024-06-26 | |||
| GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | 128K | $5 | $15 | 2024-05-13 | |||
| GPT-4oopenai/gpt-4o | 128K | $2.5 | $10 | 2024-05-13 | |||
| Qwen-VL Maxalibaba/qwen-vl-max | 131.072K | $0.8 | $3.2 | 2024-04-08 | |||
| Qwen Maxalibaba/qwen-max | 32.768K | $1.6 | $6.4 | 2024-04-03 | |||
| Claude Haiku 3anthropic/claude-3-haiku-20240307 | 200K | $0.25 | $1.25 | 2024-03-13 | |||
| Qwen Plusalibaba/qwen-plus | 1M | $0.4 | $1.2 | 2024-01-25 | |||
| Qwen-VL Plusalibaba/qwen-vl-plus | 131.072K | $0.21 | $0.63 | 2024-01-25 | |||
| Sonar Properplexity/sonar-pro | 200K | $3 | $15 | 2024-01-01 | |||
| Sonar Reasoning Properplexity/sonar-reasoning-pro | 128K | $2 | $8 | 2024-01-01 | |||
| Sonarperplexity/sonar | 128K | $1 | $1 | 2024-01-01 |