Long-lived GPT workhorse for coding, instruction following, and production apps
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
Affordable GPT-4.1 lane for fast coding help and structured extraction
Mistral vision-language model for image understanding and multimodal chat
Flagship Nemotron model for high-throughput reasoning and complex agents
Nemotron model for efficient reasoning, coding, and specialized AI agents
Open Llama with long-context vision for efficient multimodal agents
Open multimodal Llama for strong reasoning with efficient everyday serving
Dense open Qwen model for self-hosted chat, reasoning, and coding
Large open Qwen MoE for multilingual reasoning, coding, and tool use
Smaller Qwen coder for efficient local agents and repo-level fixes
Open Qwen coding heavyweight for repository reasoning and agentic engineering
March 2025 checkpoint of DeepSeek-V3 with improved reasoning and coding
O-series reasoning model for hard analysis, math, coding, and planning
Mistral reasoning model for transparent analysis, math, and complex decisions
Efficient multimodal model for instruction following, coding, reasoning, and function calling
Cohere command model for multilingual enterprise agents, tools, and chat
Open multimodal Gemma instruction model for multilingual text generation and image understanding
Largest open Gemma 3 instruction model for multilingual text generation and visual understanding
Open multimodal Gemma instruction model for efficient text generation and image understanding
Qwen reasoning model for deliberate problem solving, math, and coding
Open reasoning model from the Qwen team for math, coding, and step-by-step problem solving
Compact open multilingual vision model for OCR and visual question answering
Open multilingual vision model for OCR, visual reasoning, and image question answering
Open Command R model optimized for Arabic enterprise chat, RAG, and cultural knowledge
Balanced Claude model for coding, analysis, agent workflows, and cost control
MiniMax text-to-image generation model with reference-image support
Sonar search model for autonomous research and citation-backed long-form reports
ALLaM-2-7b instruction tuned model by SDAIA
R1 reasoning distilled into Qwen 2.5 32B for efficient open-weight step-by-step problem solving
Classic open reasoning model for transparent math, coding, and deliberate problem solving
Qwen omni model for text, vision, audio, and multimodal agent tasks
Open DeepSeek MoE chat model for coding, math, and general reasoning
Smaller o-series reasoner for economical coding, math, and planning tasks
Compact Microsoft instruction model tuned for efficient coding assistance, reasoning, and low-latency agent tasks
Earlier Gemini Flash workhorse for responsive multimodal apps and tool use
Low-latency Gemini model for high-volume multimodal and agent workloads
Popular open Llama workhorse for multilingual chat, coding, and self-hosting
O-series reasoning model for hard analysis, math, coding, and planning
Efficient model for low-latency assistance, extraction, and routine automation
Efficient model for low-latency assistance, extraction, and routine automation
Flagship model for demanding analysis, coding, and production agent workflows
Cohere retrieval model for long-context chat and enterprise RAG workflows
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Tiny open Qwen code model for lightweight completion and on-device coding
Open coding-focused Qwen model for code generation, repair, and repository reasoning
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Mistral's larger vision model for document-heavy image understanding and chat
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| GPT-4.1openai/gpt-4.1 | 1.04758M | $2 | $8 | 2025-04-14 | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| Pixtral Large (25.02)mistral/pixtral-large-2502 | 128K | $1.993 | $5.978 | 2025-04-08 | |||
| Llama 3.1 Nemotron Ultra 253Bnvidia/llama-3.1-nemotron-ultra-253b | 128K | — | — | 2025-04-07 | |||
| Llama 3.3 Nemotron Super 49B v1nvidia/llama-3.3-nemotron-super-49b-v1 | 131.072K | — | — | 2025-04-07 | |||
| Llama 4 Scout 17B Instructmeta/llama-4-scout-17b-instruct | 3.5M | $0.17 | $0.66 | 2025-04-05 | |||
| Llama 4 Maverick 17B Instructmeta/llama-4-maverick-17b-instruct | 1M | $0.14 | $0.59 | 2025-04-05 | |||
| Qwen3 32Balibaba/qwen3-32b | 131.072K | $0.7 | $2.8 | 2025-04 | |||
| Qwen3 235B-A22Balibaba/qwen3-235b-a22b | 131.072K | $0.7 | $2.8 | 2025-04 | |||
| Qwen3-Coder 30B-A3B Instructalibaba/qwen3-coder-30b-a3b-instruct | 262.144K | $0.45 | $2.25 | 2025-04 | |||
| Qwen3-Coder 480B-A35B Instructalibaba/qwen3-coder-480b-a35b-instruct | 262.144K | $1.5 | $7.5 | 2025-04 | |||
| DeepSeek V3 0324deepseek/deepseek-v3-0324 | 163.84K | $0.2 | $0.8 | 2025-03-24 | |||
| o1-proopenai/o1-pro | 200K | $150 | $600 | 2025-03-19 | |||
| Magistral Medium (latest)mistral/magistral-medium-latest | 128K | $2 | $5 | 2025-03-17 | |||
| Mistral Small 3.1 24Bmistral/mistral-small-3-1-24b-instruct-2503 | 128K | $0.106 | $0.318 | 2025-03-17 | |||
| Command Acohere/command-a-03-2025 | 256K | $2.5 | $10 | 2025-03-13 | |||
| Gemma 3 12B ITgoogle/gemma-3-12b-it | 131.072K | $0.05 | $0.1 | 2025-03-12 | |||
| Gemma 3 27B ITgoogle/gemma-3-27b-it | 131.072K | $0.08 | $0.16 | 2025-03-12 | |||
| Gemma 3 4B ITgoogle/gemma-3-4b-it | 131.072K | $0.04 | $0.08 | 2025-03-12 | |||
| QwQ Plusalibaba/qwq-plus | 131.072K | $0.8 | $2.4 | 2025-03-05 | |||
| QwQ 32Balibaba/qwq-32b | 131.072K | $0.287 | $0.861 | 2025-03-05 | |||
| Aya Vision 8Bcohere/c4ai-aya-vision-8b | 16K | — | — | 2025-03-04 | |||
| Aya Vision 32Bcohere/c4ai-aya-vision-32b | 16K | — | — | 2025-03-04 | |||
| Command R7B Arabiccohere/command-r7b-arabic-02-2025 | 128K | $0.037 | $0.15 | 2025-02-27 | |||
| Claude Sonnet 3.7anthropic/claude-3-7-sonnet-20250219 | 200K | $3 | $15 | 2025-02-19 | |||
| MiniMax image-01minimax/image-01 | Not documented | — | — | 2025-02-15 | |||
| Sonar Deep Researchperplexity/sonar-deep-research | 128K | $2 | $8 | 2025-02-01 | |||
| ALLaM-2-7bsdaia/allam-2-7b | 4.096K | — | — | 2025-01-23 | |||
| DeepSeek-R1-Distill-Qwen-32Bdeepseek/deepseek-r1-distill-qwen-32b | 131.072K | $0.3 | $0.3 | 2025-01-20 | |||
| DeepSeek-R1deepseek/deepseek-r1 | 128K | $0.7 | $2.5 | 2025-01-20 | |||
| Qwen-Omni Turboalibaba/qwen-omni-turbo | 32.768K | $0.07 | $0.27 | 2025-01-19 | |||
| DeepSeek-V3deepseek/deepseek-v3 | 131.072K | $0.27 | $1.12 | 2024-12-26 | |||
| o3-miniopenai/o3-mini | 200K | $1.1 | $4.4 | 2024-12-20 | |||
| Phi-4-minimicrosoft/phi-4-mini | 128K | $0.075 | $0.3 | 2024-12-11 | |||
| Gemini 2.0 Flashgoogle/gemini-2.0-flash | 1.04858M | $0.1 | $0.42 | 2024-12-11 | |||
| Gemini 2.0 Flash-Litegoogle/gemini-2.0-flash-lite | 1.04858M | $0.052 | $0.21 | 2024-12-11 | |||
| Llama-3.3-70B-Instructmeta/llama-3.3-70b-instruct | 128K | $0.1 | $0.32 | 2024-12-06 | |||
| o1openai/o1 | 200K | $15 | $60 | 2024-12-05 | |||
| Nova Microamazon/nova-micro | 128K | $0.035 | $0.14 | 2024-12-03 | |||
| Nova Liteamazon/nova-lite | 300K | $0.06 | $0.24 | 2024-12-03 | |||
| Nova Proamazon/nova-pro | 300K | $0.8 | $3.2 | 2024-12-03 | |||
| Command R7Bcohere/command-r7b-12-2024 | 128K | $0.037 | $0.15 | 2024-12-02 | |||
| GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | 128K | $2.5 | $10 | 2024-11-20 | |||
| Mistral Large 2.1mistral/mistral-large-2411 | 131.072K | $2 | $6 | 2024-11-18 | |||
| Qwen2.5-Coder-0.5Balibaba/qwen2.5-coder-0.5b | 32.768K | $0.1 | $0.1 | 2024-11-12 | |||
| Qwen2.5-Coder-32B-Instructalibaba/qwen2.5-coder-32b-instruct | 131.072K | $0.06 | $0.2 | 2024-11-12 | |||
| Mistral Large (latest)mistral/mistral-large-latest | 262.144K | $0.5 | $1.5 | 2024-11-01 | |||
| Qwen Turboalibaba/qwen-turbo | 1M | $0.05 | $0.2 | 2024-11-01 | |||
| Pixtral Large (latest)mistral/pixtral-large-latest | 128K | $2 | $6 | 2024-11-01 |