Qwen vision-language thinking model for visual reasoning, documents, and agent tasks
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Qwen vision-language instruct model for visual reasoning, documents, and agent tasks
Efficient Qwen thinking model for local reasoning, math, and coding agents
Qwen instruction model for multilingual chat, reasoning, and tool use
GLM vision model for visual reasoning, documents, and multimodal agents
Small GPT-5 for responsive agents, coding help, and everyday automation
Tiny GPT-5 lane for routing, extraction, classification, and bulk jobs
Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Open GPT reasoning model for self-hosted agents and controllable deployments
Open GPT reasoning model for self-hosted agents and controllable deployments
Lighter GLM-4.5 variant for fast coding assistance and cheaper agents
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Hybrid-reasoning GLM release that made the 4.5 line broadly useful
Qwen coding model for software agents, repository edits, and code reasoning
Hosted Qwen coder for software agents, repo edits, and long-context code
Mistral coding agent model for repository tasks and software engineering workflows
Efficient Mistral model for fast chat, extraction, and production assistants
Google's proven reasoning model for coding, math, and multimodal analysis
Fast Gemini workhorse for multimodal apps where latency and price matter
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents
Deliberate o-series reasoner for hard math, coding, and multi-step analysis
Fast o-series model for compact reasoning, coding, and tool use
Affordable GPT-4.1 lane for fast coding help and structured extraction
Long-lived GPT workhorse for coding, instruction following, and production apps
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
Flagship Nemotron model for high-throughput reasoning and complex agents
Open Llama with long-context vision for efficient multimodal agents
Open multimodal Llama for strong reasoning with efficient everyday serving
Smaller Qwen coder for efficient local agents and repo-level fixes
Dense open Qwen model for self-hosted chat, reasoning, and coding
Open Qwen coding heavyweight for repository reasoning and agentic engineering
Qwen reasoning model for deliberate problem solving, math, and coding
Qwen omni model for text, vision, audio, and multimodal agent tasks
Smaller o-series reasoner for economical coding, math, and planning tasks
Popular open Llama workhorse for multilingual chat, coding, and self-hosting
O-series reasoning model for hard analysis, math, coding, and planning
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Mistral's larger vision model for document-heavy image understanding and chat
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Qwen vision-language model for visual reasoning, documents, and agent tasks
Small omni GPT for cheap multimodal assistance and production-scale traffic
Omni-era GPT for multimodal chat, practical coding, and general assistants
Qwen vision-language model for visual reasoning, documents, and agent tasks
Flagship Qwen model for complex reasoning, coding, and agentic workflows
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen instruction model for multilingual chat, reasoning, and tool use
Deeper Sonar search model with broader retrieval and stronger synthesis
Fast web-grounded Sonar for current answers, citations, and lightweight retrieval
Compact GPT model for low-latency assistance and high-volume workloads
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Qwen3 VL 235B A22B Thinkingalibaba/qwen3-vl-235b-a22b-thinking | 131.072K | $0.4 | $4 | 2025-09-23 | |||
| Qwen3 VL 235B A22B Instructalibaba/qwen3-vl-235b-a22b-instruct | 131.072K | $0.2 | $0.88 | 2025-09-23 | |||
| Qwen3-Next 80B-A3B (Thinking)alibaba/qwen3-next-80b-a3b-thinking | 131.072K | $0.5 | $6 | 2025-09 | |||
| Qwen3-Next 80B-A3B Instructalibaba/qwen3-next-80b-a3b-instruct | 131.072K | $0.5 | $2 | 2025-09 | |||
| GLM-4.5Vzhipuai/glm-4.5v | 64K | $0.6 | $1.8 | 2025-08-11 | |||
| GPT-5 Miniopenai/gpt-5-mini | 400K | $0.25 | $2 | 2025-08-07 | |||
| GPT-5 Nanoopenai/gpt-5-nano | 400K | $0.05 | $0.4 | 2025-08-07 | |||
| GPT-5openai/gpt-5 | 400K | $1.25 | $10 | 2025-08-07 | |||
| Claude Opus 4.1anthropic/claude-opus-4-1-20250805 | 200K | $15 | $75 | 2025-08-05 | |||
| GPT OSS 120Bopenai/gpt-oss-120b | 131.072K | $0.03 | $0.17 | 2025-08-05 | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 131.072K | $0.02 | $0.1 | 2025-08-05 | |||
| GLM-4.5-Airzhipuai/glm-4.5-air | 131.072K | $0.2 | $1.1 | 2025-07-28 | |||
| Qwen Flashalibaba/qwen-flash | 1M | $0.05 | $0.4 | 2025-07-28 | |||
| GLM-4.5zhipuai/glm-4.5 | 131.072K | $0.6 | $2.2 | 2025-07-28 | |||
| Qwen3 Coder Flashalibaba/qwen3-coder-flash | 1M | $0.3 | $1.5 | 2025-07-28 | |||
| Qwen3 Coder Plusalibaba/qwen3-coder-plus | 1.04858M | $1 | $5 | 2025-07-23 | |||
| Devstral Smallmistral/devstral-small-2507 | 128K | $0.1 | $0.3 | 2025-07-10 | |||
| Mistral Small 3.2mistral/mistral-small-2506 | 128K | $0.1 | $0.3 | 2025-06-20 | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1.04858M | $0.3 | $2.5 | 2025-06-17 | |||
| Gemini 2.5 Flash-Litegoogle/gemini-2.5-flash-lite | 1.04858M | $0.1 | $0.4 | 2025-06-17 | |||
| o3openai/o3 | 200K | $2 | $8 | 2025-04-16 | |||
| o4-miniopenai/o4-mini | 200K | $1.1 | $4.4 | 2025-04-16 | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| GPT-4.1openai/gpt-4.1 | 1.04758M | $2 | $8 | 2025-04-14 | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| Llama 3.1 Nemotron Ultra 253Bnvidia/llama-3.1-nemotron-ultra-253b | 128K | — | — | 2025-04-07 | |||
| Llama 4 Scout 17B Instructmeta/llama-4-scout-17b-instruct | 10M | $0.17 | $0.66 | 2025-04-05 | |||
| Llama 4 Maverick 17B Instructmeta/llama-4-maverick-17b-instruct | 1M | $0.14 | $0.59 | 2025-04-05 | |||
| Qwen3-Coder 30B-A3B Instructalibaba/qwen3-coder-30b-a3b-instruct | 262.144K | $0.45 | $2.25 | 2025-04 | |||
| Qwen3 32Balibaba/qwen3-32b | 131.072K | $0.7 | $2.8 | 2025-04 | |||
| Qwen3-Coder 480B-A35B Instructalibaba/qwen3-coder-480b-a35b-instruct | 262.144K | $1.5 | $7.5 | 2025-04 | |||
| QwQ Plusalibaba/qwq-plus | 131.072K | $0.8 | $2.4 | 2025-03-05 | |||
| Qwen-Omni Turboalibaba/qwen-omni-turbo | 32.768K | $0.07 | $0.27 | 2025-01-19 | |||
| o3-miniopenai/o3-mini | 200K | $1.1 | $4.4 | 2024-12-20 | |||
| Llama-3.3-70B-Instructmeta/llama-3.3-70b-instruct | 128K | $0.1 | $0.32 | 2024-12-06 | |||
| o1openai/o1 | 200K | $15 | $60 | 2024-12-05 | |||
| Mistral Large (latest)mistral/mistral-large-latest | 262.144K | $0.5 | $1.5 | 2024-11-01 | |||
| Pixtral Large (latest)mistral/pixtral-large-latest | 128K | $2 | $6 | 2024-11-01 | |||
| Qwen Turboalibaba/qwen-turbo | 1M | $0.05 | $0.2 | 2024-11-01 | |||
| Qwen2.5-VL 72B Instructalibaba/qwen2-5-vl-72b-instruct | 131.072K | $2.8 | $8.4 | 2024-09 | |||
| GPT-4o miniopenai/gpt-4o-mini | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| GPT-4oopenai/gpt-4o | 128K | $2.5 | $10 | 2024-05-13 | |||
| Qwen-VL Maxalibaba/qwen-vl-max | 131.072K | $0.8 | $3.2 | 2024-04-08 | |||
| Qwen Maxalibaba/qwen-max | 32.768K | $1.6 | $6.4 | 2024-04-03 | |||
| Qwen-VL Plusalibaba/qwen-vl-plus | 131.072K | $0.21 | $0.63 | 2024-01-25 | |||
| Qwen Plusalibaba/qwen-plus | 1M | $0.4 | $1.2 | 2024-01-25 | |||
| Sonar Properplexity/sonar-pro | 200K | $3 | $15 | 2024-01-01 | |||
| Sonarperplexity/sonar | 128K | $1 | $1 | 2024-01-01 | |||
| GPT-4 Turboopenai/gpt-4-turbo | 128K | $10 | $30 | 2023-11-06 |