Efficient Qwen model for fast chat, extraction, and high-volume workloads
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Hosted Qwen coder for software agents, repo edits, and long-context code
Updated large open Qwen3 MoE instruct model for multilingual chat, coding, and tool use
Google's proven reasoning model for coding, math, and multimodal analysis
Fast Gemini workhorse for multimodal apps where latency and price matter
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents
High-effort o3 tier for difficult technical reasoning and careful answers
Balanced Claude model for coding, analysis, agent workflows, and cost control
Balanced Claude model for coding, analysis, agent workflows, and cost control
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Multimodal model for complex analysis, long-context understanding, tool use, and model distillation
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
Deliberate o-series reasoner for hard math, coding, and multi-step analysis
Fast o-series model for compact reasoning, coding, and tool use
Long-lived GPT workhorse for coding, instruction following, and production apps
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
Affordable GPT-4.1 lane for fast coding help and structured extraction
Open Llama with long-context vision for efficient multimodal agents
Open multimodal Llama for strong reasoning with efficient everyday serving
Open Qwen coding heavyweight for repository reasoning and agentic engineering
Smaller Qwen coder for efficient local agents and repo-level fixes
O-series reasoning model for hard analysis, math, coding, and planning
Cohere command model for multilingual enterprise agents, tools, and chat
Balanced Claude model for coding, analysis, agent workflows, and cost control
Smaller o-series reasoner for economical coding, math, and planning tasks
Earlier Gemini Flash workhorse for responsive multimodal apps and tool use
Low-latency Gemini model for high-volume multimodal and agent workloads
O-series reasoning model for hard analysis, math, coding, and planning
Flagship model for demanding analysis, coding, and production agent workflows
Efficient model for low-latency assistance, extraction, and routine automation
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Fast Claude model for responsive assistance, classification, and lightweight agents
Balanced Claude model for coding, analysis, agent workflows, and cost control
Research model for long-horizon investigation, synthesis, and analytical reports
Research model for long-horizon investigation, synthesis, and analytical reports
Mistral code model for completions, refactors, and developer IDE workflows
Legacy model retained for compatibility with older integrations
Qwen instruction model for multilingual chat, reasoning, and tool use
Deeper Sonar search model with broader retrieval and stronger synthesis
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Qwen Flashalibaba/qwen-flash | 1M | $0.05 | $0.4 | 2025-07-28 | |||
| Qwen3 Coder Plusalibaba/qwen3-coder-plus | 1.04858M | $1 | $5 | 2025-07-23 | |||
| Qwen3 235B-A22B Instruct 2507alibaba/qwen3-235b-a22b-instruct-2507 | 262.144K | $0.069 | $0.455 | 2025-07-21 | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1.04858M | $0.3 | $2.5 | 2025-06-17 | |||
| Gemini 2.5 Flash-Litegoogle/gemini-2.5-flash-lite | 1.04858M | $0.1 | $0.4 | 2025-06-17 | |||
| o3-proopenai/o3-pro | 200K | $20 | $80 | 2025-06-10 | |||
| Claude Sonnet 4anthropic/claude-sonnet-4-20250514 | 200K | $3 | $15 | 2025-05-22 | |||
| Claude Sonnet 4 (latest)anthropic/claude-sonnet-4-0 | 200K | $2.898 | $14.493 | 2025-05-22 | |||
| Claude Opus 4 (latest)anthropic/claude-opus-4-0 | 200K | — | — | 2025-05-22 | |||
| Claude Opus 4anthropic/claude-opus-4-20250514 | 200K | $15 | $75 | 2025-05-22 | |||
| Nova Premieramazon/nova-premier | 1M | — | — | 2025-04-30 | |||
| Palmyra X5writer/palmyra-x5 | 1M | $0.6 | $6 | 2025-04-28 | |||
| o3openai/o3 | 200K | $2 | $8 | 2025-04-16 | |||
| o4-miniopenai/o4-mini | 200K | $1.1 | $4.4 | 2025-04-16 | |||
| GPT-4.1openai/gpt-4.1 | 1.04758M | $2 | $8 | 2025-04-14 | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| Llama 4 Scout 17B Instructmeta/llama-4-scout-17b-instruct | 10M | $0.17 | $0.66 | 2025-04-05 | |||
| Llama 4 Maverick 17B Instructmeta/llama-4-maverick-17b-instruct | 1M | $0.14 | $0.59 | 2025-04-05 | |||
| Qwen3-Coder 480B-A35B Instructalibaba/qwen3-coder-480b-a35b-instruct | 262.144K | $1.5 | $7.5 | 2025-04 | |||
| Qwen3-Coder 30B-A3B Instructalibaba/qwen3-coder-30b-a3b-instruct | 262.144K | $0.45 | $2.25 | 2025-04 | |||
| o1-proopenai/o1-pro | 200K | $150 | $600 | 2025-03-19 | |||
| Command Acohere/command-a-03-2025 | 256K | $2.5 | $10 | 2025-03-13 | |||
| Claude Sonnet 3.7anthropic/claude-3-7-sonnet-20250219 | 200K | $3 | $15 | 2025-02-19 | |||
| o3-miniopenai/o3-mini | 200K | $1.1 | $4.4 | 2024-12-20 | |||
| Gemini 2.0 Flashgoogle/gemini-2.0-flash | 1.04858M | $0.1 | $0.42 | 2024-12-11 | |||
| Gemini 2.0 Flash-Litegoogle/gemini-2.0-flash-lite | 1.04858M | $0.052 | $0.21 | 2024-12-11 | |||
| o1openai/o1 | 200K | $15 | $60 | 2024-12-05 | |||
| Nova Proamazon/nova-pro | 300K | $0.8 | $3.2 | 2024-12-03 | |||
| Nova Liteamazon/nova-lite | 300K | $0.06 | $0.24 | 2024-12-03 | |||
| Mistral Large (latest)mistral/mistral-large-latest | 262.144K | $0.5 | $1.5 | 2024-11-01 | |||
| Qwen Turboalibaba/qwen-turbo | 1M | $0.05 | $0.2 | 2024-11-01 | |||
| Claude Haiku 3.5anthropic/claude-3-5-haiku-20241022 | 200K | $0.8 | $4 | 2024-10-22 | |||
| Claude Sonnet 3.5 v2anthropic/claude-3-5-sonnet-20241022 | 200K | — | — | 2024-10-22 | |||
| o3-deep-researchopenai/o3-deep-research | 200K | $9 | $36 | 2024-06-26 | |||
| o4-mini-deep-researchopenai/o4-mini-deep-research | 200K | $1.8 | $7.2 | 2024-06-26 | |||
| Codestral (latest)mistral/codestral-latest | 256K | $0.3 | $0.9 | 2024-05-29 | |||
| Claude Haiku 3anthropic/claude-3-haiku-20240307 | 200K | $0.25 | $1.25 | 2024-03-13 | |||
| Qwen Plusalibaba/qwen-plus | 1M | $0.4 | $1.2 | 2024-01-25 | |||
| Sonar Properplexity/sonar-pro | 200K | $3 | $15 | 2024-01-01 | |||
| us.amazon.nova-2-lite-v1:0bedrock_converse/us.amazon.nova-2-lite-v1:0 | 1M | $0.33 | $2.75 | — | |||
| eu.amazon.nova-2-pro-preview-20251202-v1:0bedrock_converse/eu.amazon.nova-2-pro-preview-20251202-v1:0 | 1M | $2.188 | $17.5 | — | |||
| FW-Kimi-K3azure_ai/fw-kimi-k3 | 1.04858M | $3.3 | $16.5 | — | |||
| eu.anthropic.claude-sonnet-5bedrock_converse/eu.anthropic.claude-sonnet-5 | 1M | $2.2 | $11 | — | |||
| us.anthropic.claude-sonnet-5bedrock_converse/us.anthropic.claude-sonnet-5 | 1M | $2.2 | $11 | — | |||
| DeepSeek: DeepSeek V4 Pro 0813 (batch)deepseek/deepseek-v4-pro-0813:batch | 1.04858M | $0.66 | $1.98 | — | |||
| eu.amazon.nova-2-lite-v1:0bedrock_converse/eu.amazon.nova-2-lite-v1:0 | 1M | $0.33 | $2.75 | — | |||
| us-gov-west-1/anthropic.claude-fable-5-1bedrock/us-gov-west-1/anthropic.claude-fable-5-1 | 1M | $12 | $60 | — | |||
| xai.grok-4.6bedrock_mantle/xai.grok-4.6 | 500K | $2.2 | $6.6 | — |