Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Balanced Claude model for coding, analysis, agent workflows, and cost control
DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass....
Open DeepSeek MoE chat model for coding, math, and general reasoning
OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model supports the `reasoning_effort` parameter, which can be set to...
Popular open Llama workhorse for multilingual chat, coding, and self-hosting
The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason...
Command R7B (12-2024) is a small, fast update of the Command R+ model, delivered in December 2024. It excels at RAG, tool use, agents, and similar tasks requiring complex reasoning...
The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improve relevance & readability. It’s also better at working with uploaded...
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Mistral's larger vision model for document-heavy image understanding and chat
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Enterprise language model for workflow automation, coding, data analysis, and tool use
command-r-08-2024 is an update of the [Command R](/models/cohere/command-r) with improved performance for multilingual retrieval-augmented generation (RAG) and tool use. More broadly, it is better at math, code and reasoning and...
command-r-plus-08-2024 is an update of the [Command R+](/models/cohere/command-r-plus) with roughly 50% higher throughput and 25% lower latencies as compared to the previous Command R+ version, while keeping the hardware footprint...
The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respone_format. Read more [here](https://openai.com/index/introducing-structured-outputs-in-the-api/). GPT-4o ("o" for "omni") is...
Compact open Llama model for lightweight chat, drafting, and self-hosting
Open Llama instruction model for multilingual chat, reasoning, and coding
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
Mistral code model for completions, refactors, and developer IDE workflows
GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as...
GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as...
Qwen instruction model for multilingual chat, reasoning, and tool use
The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.
OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous models due to its broader general knowledge and advanced reasoning...
GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Google: Gemma 3 12Bgoogle/gemma-3-12b-it | 131.072K | $0.05 | $0.15 | 2025-03-12 | |||
| Claude Sonnet 3.7anthropic/claude-3-7-sonnet-20250219 | 200K | $3 | $15 | 2025-02-19 | |||
| DeepSeek: R1deepseek/deepseek-r1 | 64K | $0.7 | $2.5 | 2025-01-20 | |||
| DeepSeek-V3deepseek/deepseek-v3 | 131.072K | $0.27 | $1.12 | 2024-12-26 | |||
| OpenAI: o3 Miniopenai/o3-mini | 200K | $1.1 | $4.4 | 2024-12-20 | |||
| Llama-3.3-70B-Instructmeta/llama-3.3-70b-instruct | 128K | $0.1 | $0.32 | 2024-12-06 | |||
| OpenAI: o1openai/o1 | 200K | $15 | $60 | 2024-12-05 | |||
| Cohere: Command R7B (12-2024)cohere/command-r7b-12-2024 | 128K | $0.037 | $0.15 | 2024-12-02 | |||
| OpenAI: GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | 128K | $2.5 | $10 | 2024-11-20 | |||
| Mistral Large 2.1mistral/mistral-large-2411 | 131.072K | $2 | $6 | 2024-11-18 | |||
| Pixtral Large (latest)mistral/pixtral-large-latest | 128K | $2 | $6 | 2024-11-01 | |||
| Mistral Large (latest)mistral/mistral-large-latest | 262.144K | $0.5 | $1.5 | 2024-11-01 | |||
| Palmyra X4writer/palmyra-x4 | 128K | $2.5 | $10 | 2024-10-09 | |||
| Cohere: Command R (08-2024)cohere/command-r-08-2024 | 128K | $0.15 | $0.6 | 2024-08-30 | |||
| Cohere: Command R+ (08-2024)cohere/command-r-plus-08-2024 | 128K | $2.5 | $10 | 2024-08-30 | |||
| OpenAI: GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | 128K | $2.5 | $10 | 2024-08-06 | |||
| Llama-3.1-8B-Instructmeta/llama-3.1-8b-instruct | 128K | $0.02 | $0.04 | 2024-07-23 | |||
| Llama-3.1-70B-Instructmeta/llama-3.1-70b-instruct | 128K | $0.4 | $0.4 | 2024-07-23 | |||
| OpenAI: GPT-4o-miniopenai/gpt-4o-mini | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| Codestral (latest)mistral/codestral-latest | 256K | $0.3 | $0.9 | 2024-05-29 | |||
| OpenAI: GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | 128K | $5 | $15 | 2024-05-13 | |||
| OpenAI: GPT-4oopenai/gpt-4o | 128K | $2.5 | $10 | 2024-05-13 | |||
| Qwen Plusalibaba/qwen-plus | 1M | $0.4 | $1.2 | 2024-01-25 | |||
| OpenAI: GPT-4 Turboopenai/gpt-4-turbo | 128K | $10 | $30 | 2023-11-06 | |||
| OpenAI: GPT-4openai/gpt-4 | 8.191K | $30 | $60 | 2023-11-06 | |||
| OpenAI: GPT-3.5 Turboopenai/gpt-3.5-turbo | 16.385K | $0.5 | $1.5 | 2023-03-01 |