GPT model for general reasoning, writing, coding, and tool-assisted tasks
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Small omni GPT for cheap multimodal assistance and production-scale traffic
Mistral code model for completions, refactors, and developer IDE workflows
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Omni-era GPT for multimodal chat, practical coding, and general assistants
Fast web-grounded Sonar for current answers, citations, and lightweight retrieval
Deeper Sonar search model with broader retrieval and stronger synthesis
Compact GPT model for low-latency assistance and high-volume workloads
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...
GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks....
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It...
OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning...
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...
Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...
Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in...
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...
GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading...
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for...
Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances state-of-the-art reasoning and multimodal performance with 8× lower cost...
Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the...
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | 128K | $2.5 | $10 | 2024-08-06 | |||
| GPT-4o miniopenai/gpt-4o-mini | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| Codestral (latest)mistral/codestral-latest | 256K | $0.3 | $0.9 | 2024-05-29 | |||
| GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | 128K | $5 | $15 | 2024-05-13 | |||
| GPT-4oopenai/gpt-4o | 128K | $2.5 | $10 | 2024-05-13 | |||
| Sonarperplexity/sonar | 128K | $1 | $1 | 2024-01-01 | |||
| Sonar Properplexity/sonar-pro | 200K | $3 | $15 | 2024-01-01 | |||
| GPT-4 Turboopenai/gpt-4-turbo | 128K | $10 | $30 | 2023-11-06 | |||
| Qwen: Qwen3.8 27Bqwen/qwen3.8-27b | 262.144K | $0.214 | $2.55 | — | |||
| MiniMax: MiniMax M2.7minimax/minimax-m2.7 | 204.8K | $0.3 | $1.2 | — | |||
| Z.ai: GLM 5.3 Flashz-ai/glm-5.3-flash | 1.04858M | $0.15 | $0.5 | — | |||
| Qwen: Qwen3.8 2.4T A95Bqwen/qwen3.8-2.4t-a95b | 1M | $2 | $6 | — | |||
| OpenAI: GPT-5 Codex (batch)openai/gpt-5-codex:batch | 400K | $0.625 | $5 | — | |||
| Google: Gemini 3.5 Flash (batch)google/gemini-3.5-flash:batch | 1.04858M | $0.75 | $4.5 | — | |||
| OpenAI: GPT-4o-mini (batch)openai/gpt-4o-mini:batch | 128K | $0.075 | $0.3 | — | |||
| Mistral: Mistral Small 4 (batch)mistralai/mistral-small-2603:batch | 262.144K | $0.075 | $0.3 | — | |||
| Qwen: Qwen3 Maxqwen/qwen3-max | 262.144K | $0.78 | $3.9 | — | |||
| OpenAI: o4 Mini (batch)openai/o4-mini:batch | 200K | $0.55 | $2.2 | — | |||
| Z.ai: GLM 5.3 Flash (batch)z-ai/glm-5.3-flash:batch | 1.04858M | $0.075 | $0.25 | — | |||
| MiniMax: MiniMax M2.5minimax/minimax-m2.5 | 204.8K | $0.3 | $1.2 | — | |||
| Qwen: Qwen3.6 Plusqwen/qwen3.6-plus | 1M | $0.325 | $1.95 | — | |||
| OpenAI: GPT-4.1 (batch)openai/gpt-4.1:batch | 1.04758M | $1 | $4 | — | |||
| Z.ai: GLM 5.3 (batch)z-ai/glm-5.3:batch | 1.04858M | $0.7 | $2.2 | — | |||
| OpenAI: gpt-oss-120b (free)openai/gpt-oss-120b:free | 131.072K | Free | Free | — | |||
| Anthropic: Claude Haiku 4.5 (batch)anthropic/claude-haiku-4.5:batch | 200K | $0.5 | $2.5 | — | |||
| Qwen: Qwen3 Coder 480B A35Bqwen/qwen3-coder | 262.144K | $0.3 | $1 | — | |||
| Anthropic: Claude Opus 4.6 (batch)anthropic/claude-opus-4.6:batch | 1M | $2.5 | $12.5 | — | |||
| Anthropic: Claude Opus 4anthropic/claude-opus-4 | 200K | $15 | $75 | — | |||
| Google: Gemma 4 31B (batch)google/gemma-4-31b-it:batch | 262.144K | $0.39 | $0.97 | — | |||
| Anthropic: Claude Opus 4.8anthropic/claude-opus-4.8 | 1M | $5 | $25 | — | |||
| Z.ai: GLM 5z-ai/glm-5 | 198K | $0.6 | $1.92 | — | |||
| NVIDIA: Nemotron 3.5 Lightning (free)nvidia/nemotron-3.5-lightning:free | 1M | Free | Free | — | |||
| Mistral: Ministral 3 8B 2512mistralai/ministral-8b-2512 | 262.144K | $0.15 | $0.15 | — | |||
| MoonshotAI: Kimi K2 0711moonshotai/kimi-k2 | 131.072K | $0.57 | $2.3 | — | |||
| Mistral: Mistral Medium 3mistralai/mistral-medium-3 | 131.072K | $0.4 | $2 | — | |||
| Qwen: Qwen3 Coder 30B A3B Instructqwen/qwen3-coder-30b-a3b-instruct | 262.144K | $0.07 | $0.28 | — | |||
| OpenAI: GPT-5 Mini (batch)openai/gpt-5-mini:batch | 400K | $0.125 | $1 | — | |||
| DeepSeek: DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp | 163.84K | $0.27 | $0.41 | — | |||
| inclusionAI: Ling 3.0 Flash VL (free)inclusionai/ling-3.0-flash-vl:free | 262.144K | Free | Free | — | |||
| Meta: Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | 131.072K | $0.05 | $0.08 | — | |||
| DeepSeek: DeepSeek V3.1deepseek/deepseek-chat-v3.1 | 163.84K | $0.25 | $0.95 | — | |||
| Qwen: Qwen3.5-9Bqwen/qwen3.5-9b | 262.144K | $0.1 | $0.15 | — | |||
| MiniMax: MiniMax M3minimax/minimax-m3 | 524.288K | $0.3 | $1.2 | — | |||
| OpenAI: GPT-5.4 Nano (batch)openai/gpt-5.4-nano:batch | 400K | $0.1 | $0.625 | — | |||
| NVIDIA: Nemotron 3 Nano 30B A3B (free)nvidia/nemotron-3-nano-30b-a3b:free | 256K | Free | Free | — | |||
| DeepSeek: DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminus | 131.072K | $0.27 | $1 | — | |||
| OpenAI: GPT-5.4 Mini (batch)openai/gpt-5.4-mini:batch | 400K | $0.375 | $2.25 | — | |||
| Qwen: Qwen3 8Bqwen/qwen3-8b | 131.072K | $0.117 | $0.455 | — | |||
| inclusionAI: Ling 3.0 Flashinclusionai/ling-3.0-flash | 262.144K | $0.021 | $0.063 | — | |||
| DeepSeek: DeepSeek V3 0324deepseek/deepseek-chat-v3-0324 | 163.84K | $0.25 | $1 | — |