Efficient Qwen model for fast chat, extraction, and high-volume workloads
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Compact open multilingual model optimized for generation across 23 languages
Open multilingual model optimized for generation across 23 languages
Balanced Claude model for coding, analysis, agent workflows, and cost control
Fast Claude model for responsive assistance, classification, and lightweight agents
Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads
Efficient open Mistral edge model for on-device chat and function calling
Open Whisper checkpoint for robust multilingual transcription and captioning
Speech transcription model for accurate audio-to-text and captioning workflows
Small open Llama base model for lightweight text generation and self-hosting
Open multimodal Llama model for image understanding, captioning, and visual QA
Compact open Llama base model for lightweight and on-device use
Mistral vision-language model for image understanding and multimodal chat
Qwen vision-language model for visual reasoning, documents, and agent tasks
Cohere's RAG workhorse for long-context enterprise search and tool use
Cohere retrieval model for long-context chat and enterprise RAG workflows
Compact Nemotron model for efficient reasoning and deployable AI agents
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Llama 3.1-based safety classifier for moderating prompts and model responses
Open Llama instruction model for multilingual chat, reasoning, and coding
Compact open Llama model for lightweight chat, drafting, and self-hosting
Small omni GPT for cheap multimodal assistance and production-scale traffic
Efficient Mistral-NVIDIA open model for multilingual chat and local deployment
Research model for long-horizon investigation, synthesis, and analytical reports
Research model for long-horizon investigation, synthesis, and analytical reports
Mistral code model for completions, refactors, and developer IDE workflows
Open Mistral code model for fill-in-the-middle and 80+ programming languages
Omni-era GPT for multimodal chat, practical coding, and general assistants
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Flagship Qwen model for complex reasoning, coding, and agentic workflows
Legacy model retained for compatibility with older integrations
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen instruction model for multilingual chat, reasoning, and tool use
Web-grounded Sonar for multi-step research questions that need cited reasoning
Deeper Sonar search model with broader retrieval and stronger synthesis
Fast web-grounded Sonar for current answers, citations, and lightweight retrieval
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Compact GPT model for low-latency assistance and high-volume workloads
Compact GPT model for low-latency assistance and high-volume workloads
Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Qwen Turboalibaba/qwen-turbo | 1M | $0.05 | $0.2 | 2024-11-01 | |||
| Aya Expanse 8Bcohere/c4ai-aya-expanse-8b | 8K | — | — | 2024-10-24 | |||
| Aya Expanse 32Bcohere/c4ai-aya-expanse-32b | 128K | — | — | 2024-10-24 | |||
| Claude Sonnet 3.5 v2anthropic/claude-3-5-sonnet-20241022 | 200K | — | — | 2024-10-22 | |||
| Claude Haiku 3.5anthropic/claude-3-5-haiku-20241022 | 200K | $0.8 | $4 | 2024-10-22 | |||
| Ministral 3Bmistral/ministral-3b | 128K | $0.04 | $0.04 | 2024-10-16 | |||
| Ministral 8B Instructmistral/ministral-8b-instruct-2410 | 131.072K | $0.15 | $0.15 | 2024-10-16 | |||
| Whisper 3 Largeopenai/whisper-large-v3 | 448 | $0.002 | $0.002 | 2024-10-01 | |||
| Whisper Large v3 Turboopenai/whisper-large-v3-turbo | 448 | $0.002 | $0.002 | 2024-10-01 | |||
| Llama-3.2-3Bmeta/llama-3.2-3b | 131.072K | $0.1 | $0.1 | 2024-09-25 | |||
| Llama-3.2-11B-Vision-Instructmeta/llama-3.2-11b-vision-instruct | 128K | $0.055 | $0.055 | 2024-09-25 | |||
| Llama-3.2-1Bmeta/llama-3.2-1b | 131.072K | $0.1 | $0.1 | 2024-09-25 | |||
| Pixtral 12Bmistral/pixtral-12b | 128K | $0.15 | $0.15 | 2024-09-01 | |||
| Qwen2.5-VL 72B Instructalibaba/qwen2-5-vl-72b-instruct | 131.072K | $2.8 | $8.4 | 2024-09 | |||
| Command R+cohere/command-r-plus-08-2024 | 128K | $2.5 | $10 | 2024-08-30 | |||
| Command Rcohere/command-r-08-2024 | 128K | $0.15 | $0.6 | 2024-08-30 | |||
| Nemotron Mini 4B Instructnvidia/nemotron-mini-4b-instruct | 128K | — | — | 2024-08-21 | |||
| GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | 128K | $2.5 | $10 | 2024-08-06 | |||
| Llama-Guard-3-8Bmeta/llama-guard-3-8b | 128K | — | — | 2024-07-23 | |||
| Llama-3.1-70B-Instructmeta/llama-3.1-70b-instruct | 128K | $0.4 | $0.4 | 2024-07-23 | |||
| Llama-3.1-8B-Instructmeta/llama-3.1-8b-instruct | 128K | $0.02 | $0.04 | 2024-07-23 | |||
| GPT-4o miniopenai/gpt-4o-mini | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| Mistral Nemomistral/mistral-nemo | 128K | $0.15 | $0.15 | 2024-07-01 | |||
| o3-deep-researchopenai/o3-deep-research | 200K | $9 | $36 | 2024-06-26 | |||
| o4-mini-deep-researchopenai/o4-mini-deep-research | 200K | $1.8 | $7.2 | 2024-06-26 | |||
| Codestral (latest)mistral/codestral-latest | 256K | $0.3 | $0.9 | 2024-05-29 | |||
| Codestral-22B-v0.1mistral/codestral-22b-v0.1 | 32.768K | $0.3 | $0.9 | 2024-05-29 | |||
| GPT-4oopenai/gpt-4o | 128K | $2.5 | $10 | 2024-05-13 | |||
| GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | 128K | $5 | $15 | 2024-05-13 | |||
| Qwen-VL Maxalibaba/qwen-vl-max | 131.072K | $0.8 | $3.2 | 2024-04-08 | |||
| Qwen Maxalibaba/qwen-max | 32.768K | $1.6 | $6.4 | 2024-04-03 | |||
| Claude Haiku 3anthropic/claude-3-haiku-20240307 | 200K | $0.25 | $1.25 | 2024-03-13 | |||
| Qwen-VL Plusalibaba/qwen-vl-plus | 131.072K | $0.21 | $0.63 | 2024-01-25 | |||
| Qwen Plusalibaba/qwen-plus | 1M | $0.4 | $1.2 | 2024-01-25 | |||
| Sonar Reasoning Properplexity/sonar-reasoning-pro | 128K | $2 | $8 | 2024-01-01 | |||
| Sonar Properplexity/sonar-pro | 200K | $3 | $15 | 2024-01-01 | |||
| Sonarperplexity/sonar | 128K | $1 | $1 | 2024-01-01 | |||
| GPT-4openai/gpt-4 | 8.192K | $30 | $60 | 2023-11-06 | |||
| GPT-4 Turboopenai/gpt-4-turbo | 128K | $10 | $30 | 2023-11-06 | |||
| GPT-3.5-turboopenai/gpt-3.5-turbo | 16.385K | $0.5 | $1.5 | 2023-03-01 | |||
| Nous: Hermes 3 405B Instruct (free)nousresearch/hermes-3-llama-3.1-405b:free | 131.072K | Free | Free | — | |||
| meta-llama/CodeLlama-13b-hfmeta-llama/CodeLlama-13b-hf | Not documented | — | — | — | |||
| meta-llama/Llama-Guard-3-8Bmeta-llama/Llama-Guard-3-8B | Not documented | — | — | — | |||
| meta-llama/CodeLlama-13b-Python-hfmeta-llama/CodeLlama-13b-Python-hf | Not documented | — | — | — | |||
| eu.anthropic.claude-opus-4-6-v1bedrock_converse/eu.anthropic.claude-opus-4-6-v1 | 1M | $5.5 | $27.5 | — | |||
| meta-llama/Llama-3.1-8Bmeta-llama/Llama-3.1-8B | Not documented | — | — | — | |||
| us-gov-east-1/openai.gpt-oss-120bbedrock_mantle/us-gov-east-1/openai.gpt-oss-120b | 131.072K | $0.18 | $0.72 | — | |||
| meta-llama/CodeLlama-13b-Instruct-hfmeta-llama/CodeLlama-13b-Instruct-hf | Not documented | — | — | — | |||
| us.anthropic.claude-opus-4-6-v1bedrock_converse/us.anthropic.claude-opus-4-6-v1 | 1M | $5.5 | $27.5 | — | |||
| meta-llama/Llama-3.2-3B-Instructmeta-llama/Llama-3.2-3B-Instruct | Not documented | — | — | — |