Balanced Claude model for coding, analysis, agent workflows, and cost control
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Qwen omni model for text, vision, audio, and multimodal agent tasks
Earlier Gemini Flash workhorse for responsive multimodal apps and tool use
Low-latency Gemini model for high-volume multimodal and agent workloads
O-series reasoning model for hard analysis, math, coding, and planning
Flagship model for demanding analysis, coding, and production agent workflows
Efficient model for low-latency assistance, extraction, and routine automation
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Mistral's larger vision model for document-heavy image understanding and chat
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Fast Claude model for responsive assistance, classification, and lightweight agents
Balanced Claude model for coding, analysis, agent workflows, and cost control
Open multimodal Llama model for image understanding, captioning, and visual QA
Mistral vision-language model for image understanding and multimodal chat
Qwen vision-language model for visual reasoning, documents, and agent tasks
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Small omni GPT for cheap multimodal assistance and production-scale traffic
Research model for long-horizon investigation, synthesis, and analytical reports
Research model for long-horizon investigation, synthesis, and analytical reports
Omni-era GPT for multimodal chat, practical coding, and general assistants
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Legacy model retained for compatibility with older integrations
Qwen vision-language model for visual reasoning, documents, and agent tasks
Web-grounded Sonar for multi-step research questions that need cited reasoning
Deeper Sonar search model with broader retrieval and stronger synthesis
Compact GPT model for low-latency assistance and high-volume workloads
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Claude Sonnet 3.7anthropic/claude-3-7-sonnet-20250219 | 200K | $3 | $15 | 2025-02-19 | |||
| Qwen-Omni Turboalibaba/qwen-omni-turbo | 32.768K | $0.07 | $0.27 | 2025-01-19 | |||
| Gemini 2.0 Flashgoogle/gemini-2.0-flash | 1.04858M | $0.1 | $0.42 | 2024-12-11 | |||
| Gemini 2.0 Flash-Litegoogle/gemini-2.0-flash-lite | 1.04858M | $0.052 | $0.21 | 2024-12-11 | |||
| o1openai/o1 | 200K | $15 | $60 | 2024-12-05 | |||
| Nova Proamazon/nova-pro | 300K | $0.8 | $3.2 | 2024-12-03 | |||
| Nova Liteamazon/nova-lite | 300K | $0.06 | $0.24 | 2024-12-03 | |||
| GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | 128K | $2.5 | $10 | 2024-11-20 | |||
| Pixtral Large (latest)mistral/pixtral-large-latest | 128K | $2 | $6 | 2024-11-01 | |||
| Mistral Large (latest)mistral/mistral-large-latest | 262.144K | $0.5 | $1.5 | 2024-11-01 | |||
| Claude Haiku 3.5anthropic/claude-3-5-haiku-20241022 | 200K | $0.8 | $4 | 2024-10-22 | |||
| Claude Sonnet 3.5 v2anthropic/claude-3-5-sonnet-20241022 | 200K | — | — | 2024-10-22 | |||
| Llama-3.2-11B-Vision-Instructmeta/llama-3.2-11b-vision-instruct | 128K | $0.055 | $0.055 | 2024-09-25 | |||
| Pixtral 12Bmistral/pixtral-12b | 128K | $0.15 | $0.15 | 2024-09-01 | |||
| Qwen2.5-VL 72B Instructalibaba/qwen2-5-vl-72b-instruct | 131.072K | $2.8 | $8.4 | 2024-09 | |||
| GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | 128K | $2.5 | $10 | 2024-08-06 | |||
| GPT-4o miniopenai/gpt-4o-mini | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| o3-deep-researchopenai/o3-deep-research | 200K | $9 | $36 | 2024-06-26 | |||
| o4-mini-deep-researchopenai/o4-mini-deep-research | 200K | $1.8 | $7.2 | 2024-06-26 | |||
| GPT-4oopenai/gpt-4o | 128K | $2.5 | $10 | 2024-05-13 | |||
| GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | 128K | $5 | $15 | 2024-05-13 | |||
| Qwen-VL Maxalibaba/qwen-vl-max | 131.072K | $0.8 | $3.2 | 2024-04-08 | |||
| Claude Haiku 3anthropic/claude-3-haiku-20240307 | 200K | $0.25 | $1.25 | 2024-03-13 | |||
| Qwen-VL Plusalibaba/qwen-vl-plus | 131.072K | $0.21 | $0.63 | 2024-01-25 | |||
| Sonar Reasoning Properplexity/sonar-reasoning-pro | 128K | $2 | $8 | 2024-01-01 | |||
| Sonar Properplexity/sonar-pro | 200K | $3 | $15 | 2024-01-01 | |||
| GPT-4 Turboopenai/gpt-4-turbo | 128K | $10 | $30 | 2023-11-06 | |||
| gemini-2.5-flash-lite-preview-06-17vertex_ai-language-models/gemini-2.5-flash-lite-preview-06-17 | 1.04858M | $0.1 | $0.4 | — | |||
| gemini-3-pro-previewvertex_ai/gemini-3-pro-preview | 1.04858M | $2 | $12 | — | |||
| gemini-2.5-flash-preview-09-2025vertex_ai-language-models/gemini-2.5-flash-preview-09-2025 | 1.04858M | $0.3 | $2.5 | — | |||
| gemini-2.5-flash-lite-preview-09-2025vertex_ai-language-models/gemini-2.5-flash-lite-preview-09-2025 | 1.04858M | $0.1 | $0.4 | — | |||
| kimi-k2.6azure_ai/kimi-k2.6 | 262.144K | $0.95 | $4 | — | |||
| us/gpt-4.1-mini-2025-04-14azure/us/gpt-4.1-mini-2025-04-14 | 1.04758M | $0.44 | $1.76 | — | |||
| gemini-2.5-flash-litevertex_ai-language-models/gemini-2.5-flash-lite | 1.04858M | $0.1 | $0.4 | — | |||
| gemini-3.5-flash-litevertex_ai-language-models/gemini-3.5-flash-lite | 1.04858M | $0.3 | $2.5 | — | |||
| kimi-k2.5azure_ai/kimi-k2.5 | 262.144K | $0.6 | $3 | — | |||
| moonshot-v1-32k-vision-previewmoonshot/moonshot-v1-32k-vision-preview | 32.768K | $1 | $3 | — | |||
| moonshot-v1-128k-vision-previewmoonshot/moonshot-v1-128k-vision-preview | 131.072K | $2 | $5 | — | |||
| gemini-3.1-flash-litevertex_ai-language-models/gemini-3.1-flash-lite | 1.04858M | $0.25 | $1.5 | — | |||
| kimi-thinking-previewmoonshot/kimi-thinking-preview | 131.072K | $0.6 | $2.5 | — | |||
| kimi-latest-8kmoonshot/kimi-latest-8k | 8.192K | $0.2 | $2 | — | |||
| gemini-3.1-flash-lite-previewvertex_ai-language-models/gemini-3.1-flash-lite-preview | 1.04858M | $0.25 | $1.5 | — | |||
| moonshot-v1-8k-vision-previewmoonshot/moonshot-v1-8k-vision-preview | 8.192K | $0.2 | $2 | — | |||
| google/gemma-3-27b-itnebius/google/gemma-3-27b-it | 128K | $0.06 | $0.2 | — | |||
| Qwen/Qwen2.5-VL-72B-Instructnebius/qwen/qwen2.5-vl-72b-instruct | 131.072K | $0.13 | $0.4 | — | |||
| Qwen/Qwen2-VL-72B-Instructnebius/qwen/qwen2-vl-72b-instruct | 131.072K | $0.13 | $0.4 | — | |||
| Qwen/Qwen2-VL-7B-Instructnebius/qwen/qwen2-vl-7b-instruct | 131.072K | $0.02 | $0.06 | — | |||
| nvidia.nemotron-nano-12b-v2bedrock_converse/nvidia.nemotron-nano-12b-v2 | 128K | $0.2 | $0.6 | — | |||
| o1-2024-12-17openai/o1-2024-12-17 | 200K | $15 | $60 | — | |||
| us/gpt-4.1-2025-04-14azure/us/gpt-4.1-2025-04-14 | 1.04758M | $2.2 | $8.8 | — |