Classic open reasoning model for transparent math, coding, and deliberate problem solving
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
R1 reasoning distilled into Qwen 2.5 32B for efficient open-weight step-by-step problem solving
Open DeepSeek MoE chat model for coding, math, and general reasoning
Smaller o-series reasoner for economical coding, math, and planning tasks
Compact Microsoft instruction model tuned for efficient coding assistance, reasoning, and low-latency agent tasks
Low-latency Gemini model for high-volume multimodal and agent workloads
Earlier Gemini Flash workhorse for responsive multimodal apps and tool use
Popular open Llama workhorse for multilingual chat, coding, and self-hosting
O-series reasoning model for hard analysis, math, coding, and planning
Efficient model for low-latency assistance, extraction, and routine automation
Flagship model for demanding analysis, coding, and production agent workflows
Efficient model for low-latency assistance, extraction, and routine automation
Cohere retrieval model for long-context chat and enterprise RAG workflows
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Open coding-focused Qwen model for code generation, repair, and repository reasoning
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Mistral's larger vision model for document-heavy image understanding and chat
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Open multilingual model optimized for generation across 23 languages
Fast Claude model for responsive assistance, classification, and lightweight agents
Balanced Claude model for coding, analysis, agent workflows, and cost control
Efficient open Mistral edge model for on-device chat and function calling
Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads
Enterprise language model for workflow automation, coding, data analysis, and tool use
Open multimodal Llama model for image understanding, captioning, and visual QA
Compact open Llama base model for lightweight and on-device use
Small open Llama base model for lightweight text generation and self-hosting
Mistral vision-language model for image understanding and multimodal chat
Qwen vision-language model for visual reasoning, documents, and agent tasks
Cohere's RAG workhorse for long-context enterprise search and tool use
Cohere retrieval model for long-context chat and enterprise RAG workflows
Compact Nemotron model for efficient reasoning and deployable AI agents
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Llama 3.1-based safety classifier for moderating prompts and model responses
Compact open Llama model for lightweight chat, drafting, and self-hosting
Open Llama instruction model for multilingual chat, reasoning, and coding
Small omni GPT for cheap multimodal assistance and production-scale traffic
Efficient Mistral-NVIDIA open model for multilingual chat and local deployment
Research model for long-horizon investigation, synthesis, and analytical reports
Research model for long-horizon investigation, synthesis, and analytical reports
Mistral code model for completions, refactors, and developer IDE workflows
Omni-era GPT for multimodal chat, practical coding, and general assistants
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Legacy model retained for compatibility with older integrations
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen vision-language model for visual reasoning, documents, and agent tasks
Web-grounded Sonar for multi-step research questions that need cited reasoning
Deeper Sonar search model with broader retrieval and stronger synthesis
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| DeepSeek-R1deepseek/deepseek-r1 | 128K | $0.7 | $2.5 | 2025-01-20 | |||
| DeepSeek-R1-Distill-Qwen-32Bdeepseek/deepseek-r1-distill-qwen-32b | 131.072K | $0.3 | $0.3 | 2025-01-20 | |||
| DeepSeek-V3deepseek/deepseek-v3 | 131.072K | $0.27 | $1.12 | 2024-12-26 | |||
| o3-miniopenai/o3-mini | 200K | $1.1 | $4.4 | 2024-12-20 | |||
| Phi-4-minimicrosoft/phi-4-mini | 128K | $0.075 | $0.3 | 2024-12-11 | |||
| Gemini 2.0 Flash-Litegoogle/gemini-2.0-flash-lite | 1.04858M | $0.052 | $0.21 | 2024-12-11 | |||
| Gemini 2.0 Flashgoogle/gemini-2.0-flash | 1.04858M | $0.1 | $0.42 | 2024-12-11 | |||
| Llama-3.3-70B-Instructmeta/llama-3.3-70b-instruct | 128K | $0.1 | $0.32 | 2024-12-06 | |||
| o1openai/o1 | 200K | $15 | $60 | 2024-12-05 | |||
| Nova Liteamazon/nova-lite | 300K | $0.06 | $0.24 | 2024-12-03 | |||
| Nova Proamazon/nova-pro | 300K | $0.8 | $3.2 | 2024-12-03 | |||
| Nova Microamazon/nova-micro | 128K | $0.035 | $0.14 | 2024-12-03 | |||
| Command R7Bcohere/command-r7b-12-2024 | 128K | $0.037 | $0.15 | 2024-12-02 | |||
| GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | 128K | $2.5 | $10 | 2024-11-20 | |||
| Mistral Large 2.1mistral/mistral-large-2411 | 131.072K | $2 | $6 | 2024-11-18 | |||
| Qwen2.5-Coder-32B-Instructalibaba/qwen2.5-coder-32b-instruct | 131.072K | $0.06 | $0.2 | 2024-11-12 | |||
| Qwen Turboalibaba/qwen-turbo | 1M | $0.05 | $0.2 | 2024-11-01 | |||
| Pixtral Large (latest)mistral/pixtral-large-latest | 128K | $2 | $6 | 2024-11-01 | |||
| Mistral Large (latest)mistral/mistral-large-latest | 262.144K | $0.5 | $1.5 | 2024-11-01 | |||
| Aya Expanse 32Bcohere/c4ai-aya-expanse-32b | 128K | — | — | 2024-10-24 | |||
| Claude Haiku 3.5anthropic/claude-3-5-haiku-20241022 | 200K | $0.8 | $4 | 2024-10-22 | |||
| Claude Sonnet 3.5 v2anthropic/claude-3-5-sonnet-20241022 | 200K | — | — | 2024-10-22 | |||
| Ministral 8B Instructmistral/ministral-8b-instruct-2410 | 131.072K | $0.15 | $0.15 | 2024-10-16 | |||
| Ministral 3Bmistral/ministral-3b | 128K | $0.04 | $0.04 | 2024-10-16 | |||
| Palmyra X4writer/palmyra-x4 | 128K | $2.5 | $10 | 2024-10-09 | |||
| Llama-3.2-11B-Vision-Instructmeta/llama-3.2-11b-vision-instruct | 128K | $0.055 | $0.055 | 2024-09-25 | |||
| Llama-3.2-1Bmeta/llama-3.2-1b | 131.072K | $0.1 | $0.1 | 2024-09-25 | |||
| Llama-3.2-3Bmeta/llama-3.2-3b | 131.072K | $0.1 | $0.1 | 2024-09-25 | |||
| Pixtral 12Bmistral/pixtral-12b | 128K | $0.15 | $0.15 | 2024-09-01 | |||
| Qwen2.5-VL 72B Instructalibaba/qwen2-5-vl-72b-instruct | 131.072K | $2.8 | $8.4 | 2024-09 | |||
| Command R+cohere/command-r-plus-08-2024 | 128K | $2.5 | $10 | 2024-08-30 | |||
| Command Rcohere/command-r-08-2024 | 128K | $0.15 | $0.6 | 2024-08-30 | |||
| Nemotron Mini 4B Instructnvidia/nemotron-mini-4b-instruct | 128K | — | — | 2024-08-21 | |||
| GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | 128K | $2.5 | $10 | 2024-08-06 | |||
| Llama-Guard-3-8Bmeta/llama-guard-3-8b | 128K | — | — | 2024-07-23 | |||
| Llama-3.1-8B-Instructmeta/llama-3.1-8b-instruct | 128K | $0.02 | $0.04 | 2024-07-23 | |||
| Llama-3.1-70B-Instructmeta/llama-3.1-70b-instruct | 128K | $0.4 | $0.4 | 2024-07-23 | |||
| GPT-4o miniopenai/gpt-4o-mini | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| Mistral Nemomistral/mistral-nemo | 128K | $0.15 | $0.15 | 2024-07-01 | |||
| o4-mini-deep-researchopenai/o4-mini-deep-research | 200K | $1.8 | $7.2 | 2024-06-26 | |||
| o3-deep-researchopenai/o3-deep-research | 200K | $9 | $36 | 2024-06-26 | |||
| Codestral (latest)mistral/codestral-latest | 256K | $0.3 | $0.9 | 2024-05-29 | |||
| GPT-4oopenai/gpt-4o | 128K | $2.5 | $10 | 2024-05-13 | |||
| GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | 128K | $5 | $15 | 2024-05-13 | |||
| Qwen-VL Maxalibaba/qwen-vl-max | 131.072K | $0.8 | $3.2 | 2024-04-08 | |||
| Claude Haiku 3anthropic/claude-3-haiku-20240307 | 200K | $0.25 | $1.25 | 2024-03-13 | |||
| Qwen Plusalibaba/qwen-plus | 1M | $0.4 | $1.2 | 2024-01-25 | |||
| Qwen-VL Plusalibaba/qwen-vl-plus | 131.072K | $0.21 | $0.63 | 2024-01-25 | |||
| Sonar Reasoning Properplexity/sonar-reasoning-pro | 128K | $2 | $8 | 2024-01-01 | |||
| Sonar Properplexity/sonar-pro | 200K | $3 | $15 | 2024-01-01 |