Qwen vision-language instruct model for visual reasoning, documents, and agent tasks
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Efficient Qwen thinking model for local reasoning, math, and coding agents
Qwen instruction model for multilingual chat, reasoning, and tool use
GLM vision model for visual reasoning, documents, and multimodal agents
GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
Flagship Claude model for deep reasoning, coding, and long-horizon agents
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Hybrid-reasoning GLM release that made the 4.5 line broadly useful
Qwen coding model for software agents, repository edits, and code reasoning
Lighter GLM-4.5 variant for fast coding assistance and cheaper agents
Hosted Qwen coder for software agents, repo edits, and long-context code
Mistral coding agent model for repository tasks and software engineering workflows
Efficient Mistral model for fast chat, extraction, and production assistants
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...
OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning...
o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...
Flagship Nemotron model for high-throughput reasoning and complex agents
Open multimodal Llama for strong reasoning with efficient everyday serving
Open Llama with long-context vision for efficient multimodal agents
Smaller Qwen coder for efficient local agents and repo-level fixes
Dense open Qwen model for self-hosted chat, reasoning, and coding
Open Qwen coding heavyweight for repository reasoning and agentic engineering
Qwen reasoning model for deliberate problem solving, math, and coding
Qwen omni model for text, vision, audio, and multimodal agent tasks
OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model supports the `reasoning_effort` parameter, which can be set to...
Popular open Llama workhorse for multilingual chat, coding, and self-hosting
The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason...
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Mistral's largest general model for enterprise agents, coding, and multilingual reasoning
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Mistral's larger vision model for document-heavy image understanding and chat
Qwen vision-language model for visual reasoning, documents, and agent tasks
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as...
Qwen vision-language model for visual reasoning, documents, and agent tasks
Flagship Qwen model for complex reasoning, coding, and agentic workflows
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen instruction model for multilingual chat, reasoning, and tool use
Note: Sonar Pro pricing includes Perplexity search pricing. See [details here](https://docs.perplexity.ai/guides/pricing#detailed-pricing-breakdown-for-sonar-reasoning-pro-and-sonar-pro) Sonar Reasoning Pro is a premier reasoning model powered by DeepSeek R1 with Chain of Thought (CoT). Designed for...
Note: Sonar Pro pricing includes Perplexity search pricing. See [details here](https://docs.perplexity.ai/guides/pricing#detailed-pricing-breakdown-for-sonar-reasoning-pro-and-sonar-pro) For enterprises seeking more advanced capabilities, the Sonar Pro API can handle in-depth, multi-step queries with added extensibility, like...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Qwen3 VL 235B A22B Instructalibaba/qwen3-vl-235b-a22b-instruct | 131.072K | $0.2 | $0.88 | 2025-09-23 | |||
| Qwen3-Next 80B-A3B (Thinking)alibaba/qwen3-next-80b-a3b-thinking | 131.072K | $0.5 | $6 | 2025-09 | |||
| Qwen3-Next 80B-A3B Instructalibaba/qwen3-next-80b-a3b-instruct | 131.072K | $0.5 | $2 | 2025-09 | |||
| GLM-4.5Vzhipuai/glm-4.5v | 64K | $0.6 | $1.8 | 2025-08-11 | |||
| OpenAI: GPT-5 Nanoopenai/gpt-5-nano | 400K | $0.05 | $0.4 | 2025-08-07 | |||
| OpenAI: GPT-5 Miniopenai/gpt-5-mini | 400K | $0.25 | $2 | 2025-08-07 | |||
| OpenAI: GPT-5openai/gpt-5 | 400K | $1.25 | $10 | 2025-08-07 | |||
| GPT-5 Chat (latest)openai/gpt-5-chat-latest | 400K | $1.25 | $10 | 2025-08-07 | |||
| Claude Opus 4.1anthropic/claude-opus-4-1-20250805 | 200K | $15 | $75 | 2025-08-05 | |||
| OpenAI: gpt-oss-120bopenai/gpt-oss-120b | 131.072K | $0.037 | $0.17 | 2025-08-05 | |||
| OpenAI: gpt-oss-20bopenai/gpt-oss-20b | 131.072K | $0.03 | $0.13 | 2025-08-05 | |||
| Qwen Flashalibaba/qwen-flash | 1M | $0.05 | $0.4 | 2025-07-28 | |||
| GLM-4.5zhipuai/glm-4.5 | 131.072K | $0.6 | $2.2 | 2025-07-28 | |||
| Qwen3 Coder Flashalibaba/qwen3-coder-flash | 1M | $0.3 | $1.5 | 2025-07-28 | |||
| GLM-4.5-Airzhipuai/glm-4.5-air | 131.072K | $0.2 | $1.1 | 2025-07-28 | |||
| Qwen3 Coder Plusalibaba/qwen3-coder-plus | 1.04858M | $1 | $5 | 2025-07-23 | |||
| Devstral Smallmistral/devstral-small-2507 | 128K | $0.1 | $0.3 | 2025-07-10 | |||
| Mistral Small 3.2mistral/mistral-small-2506 | 128K | $0.1 | $0.3 | 2025-06-20 | |||
| Google: Gemini 2.5 Progoogle/gemini-2.5-pro | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| Google: Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1.04858M | $0.3 | $2.5 | 2025-06-17 | |||
| Google: Gemini 2.5 Flash Litegoogle/gemini-2.5-flash-lite | 1.04858M | $0.1 | $0.4 | 2025-06-17 | |||
| OpenAI: o4 Miniopenai/o4-mini | 200K | $1.1 | $4.4 | 2025-04-16 | |||
| OpenAI: o3openai/o3 | 200K | $2 | $8 | 2025-04-16 | |||
| OpenAI: GPT-4.1 Miniopenai/gpt-4.1-mini | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| OpenAI: GPT-4.1 Nanoopenai/gpt-4.1-nano | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| OpenAI: GPT-4.1openai/gpt-4.1 | 1.04758M | $2 | $8 | 2025-04-14 | |||
| Llama 3.1 Nemotron Ultra 253Bnvidia/llama-3.1-nemotron-ultra-253b | 128K | — | — | 2025-04-07 | |||
| Llama 4 Maverick 17B Instructmeta/llama-4-maverick-17b-instruct | 1M | $0.14 | $0.59 | 2025-04-05 | |||
| Llama 4 Scout 17B Instructmeta/llama-4-scout-17b-instruct | 3.5M | $0.17 | $0.66 | 2025-04-05 | |||
| Qwen3-Coder 30B-A3B Instructalibaba/qwen3-coder-30b-a3b-instruct | 262.144K | $0.45 | $2.25 | 2025-04 | |||
| Qwen3 32Balibaba/qwen3-32b | 131.072K | $0.7 | $2.8 | 2025-04 | |||
| Qwen3-Coder 480B-A35B Instructalibaba/qwen3-coder-480b-a35b-instruct | 262.144K | $1.5 | $7.5 | 2025-04 | |||
| QwQ Plusalibaba/qwq-plus | 131.072K | $0.8 | $2.4 | 2025-03-05 | |||
| Qwen-Omni Turboalibaba/qwen-omni-turbo | 32.768K | $0.07 | $0.27 | 2025-01-19 | |||
| OpenAI: o3 Miniopenai/o3-mini | 200K | $1.1 | $4.4 | 2024-12-20 | |||
| Llama-3.3-70B-Instructmeta/llama-3.3-70b-instruct | 128K | $0.1 | $0.32 | 2024-12-06 | |||
| OpenAI: o1openai/o1 | 200K | $15 | $60 | 2024-12-05 | |||
| Qwen Turboalibaba/qwen-turbo | 1M | $0.05 | $0.2 | 2024-11-01 | |||
| Mistral Large 3mistral/mistral-large-2512 | 262.144K | $0.5 | $1.5 | 2024-11-01 | |||
| Mistral Large (latest)mistral/mistral-large-latest | 262.144K | $0.5 | $1.5 | 2024-11-01 | |||
| Pixtral Large (latest)mistral/pixtral-large-latest | 128K | $2 | $6 | 2024-11-01 | |||
| Qwen2.5-VL 72B Instructalibaba/qwen2-5-vl-72b-instruct | 131.072K | $2.8 | $8.4 | 2024-09 | |||
| OpenAI: GPT-4o-miniopenai/gpt-4o-mini | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| OpenAI: GPT-4oopenai/gpt-4o | 128K | $2.5 | $10 | 2024-05-13 | |||
| Qwen-VL Maxalibaba/qwen-vl-max | 131.072K | $0.8 | $3.2 | 2024-04-08 | |||
| Qwen Maxalibaba/qwen-max | 32.768K | $1.6 | $6.4 | 2024-04-03 | |||
| Qwen-VL Plusalibaba/qwen-vl-plus | 131.072K | $0.21 | $0.63 | 2024-01-25 | |||
| Qwen Plusalibaba/qwen-plus | 1M | $0.4 | $1.2 | 2024-01-25 | |||
| Perplexity: Sonar Reasoning Properplexity/sonar-reasoning-pro | 128K | $2 | $8 | 2024-01-01 | |||
| Perplexity: Sonar Properplexity/sonar-pro | 200K | $3 | $15 | 2024-01-01 |