Enterprise language model for workflow automation, coding, data analysis, and tool use
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Open multimodal Llama model for image understanding, captioning, and visual QA
Mistral vision-language model for image understanding and multimodal chat
Qwen vision-language model for visual reasoning, documents, and agent tasks
Cohere's RAG workhorse for long-context enterprise search and tool use
Cohere retrieval model for long-context chat and enterprise RAG workflows
Compact Nemotron model for efficient reasoning and deployable AI agents
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Compact open Llama model for lightweight chat, drafting, and self-hosting
Open Llama instruction model for multilingual chat, reasoning, and coding
Small omni GPT for cheap multimodal assistance and production-scale traffic
Efficient Mistral-NVIDIA open model for multilingual chat and local deployment
Research model for long-horizon investigation, synthesis, and analytical reports
Research model for long-horizon investigation, synthesis, and analytical reports
Mistral code model for completions, refactors, and developer IDE workflows
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Omni-era GPT for multimodal chat, practical coding, and general assistants
Qwen vision-language model for visual reasoning, documents, and agent tasks
Legacy model retained for compatibility with older integrations
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen instruction model for multilingual chat, reasoning, and tool use
Fast web-grounded Sonar for current answers, citations, and lightweight retrieval
Deeper Sonar search model with broader retrieval and stronger synthesis
Compact GPT model for low-latency assistance and high-volume workloads
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Palmyra X4writer/palmyra-x4 | 128K | $2.5 | $10 | 2024-10-09 | |||
| Llama-3.2-11B-Vision-Instructmeta/llama-3.2-11b-vision-instruct | 128K | $0.055 | $0.055 | 2024-09-25 | |||
| Pixtral 12Bmistral/pixtral-12b | 128K | $0.15 | $0.15 | 2024-09-01 | |||
| Qwen2.5-VL 72B Instructalibaba/qwen2-5-vl-72b-instruct | 131.072K | $2.8 | $8.4 | 2024-09 | |||
| Command R+cohere/command-r-plus-08-2024 | 128K | $2.5 | $10 | 2024-08-30 | |||
| Command Rcohere/command-r-08-2024 | 128K | $0.15 | $0.6 | 2024-08-30 | |||
| Nemotron Mini 4B Instructnvidia/nemotron-mini-4b-instruct | 128K | — | — | 2024-08-21 | |||
| GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | 128K | $2.5 | $10 | 2024-08-06 | |||
| Llama-3.1-8B-Instructmeta/llama-3.1-8b-instruct | 128K | $0.02 | $0.04 | 2024-07-23 | |||
| Llama-3.1-70B-Instructmeta/llama-3.1-70b-instruct | 128K | $0.4 | $0.4 | 2024-07-23 | |||
| GPT-4o miniopenai/gpt-4o-mini | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| Mistral Nemomistral/mistral-nemo | 128K | $0.15 | $0.15 | 2024-07-01 | |||
| o4-mini-deep-researchopenai/o4-mini-deep-research | 200K | $1.8 | $7.2 | 2024-06-26 | |||
| o3-deep-researchopenai/o3-deep-research | 200K | $9 | $36 | 2024-06-26 | |||
| Codestral (latest)mistral/codestral-latest | 256K | $0.3 | $0.9 | 2024-05-29 | |||
| GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | 128K | $5 | $15 | 2024-05-13 | |||
| GPT-4oopenai/gpt-4o | 128K | $2.5 | $10 | 2024-05-13 | |||
| Qwen-VL Maxalibaba/qwen-vl-max | 131.072K | $0.8 | $3.2 | 2024-04-08 | |||
| Claude Haiku 3anthropic/claude-3-haiku-20240307 | 200K | $0.25 | $1.25 | 2024-03-13 | |||
| Qwen-VL Plusalibaba/qwen-vl-plus | 131.072K | $0.21 | $0.63 | 2024-01-25 | |||
| Qwen Plusalibaba/qwen-plus | 1M | $0.4 | $1.2 | 2024-01-25 | |||
| Sonarperplexity/sonar | 128K | $1 | $1 | 2024-01-01 | |||
| Sonar Properplexity/sonar-pro | 200K | $3 | $15 | 2024-01-01 | |||
| GPT-4 Turboopenai/gpt-4-turbo | 128K | $10 | $30 | 2023-11-06 |