Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...
GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as...
DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...
o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....
Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in...
GPT-5.3 Chat is an update to ChatGPT's most-used model that makes everyday conversations smoother, more useful, and more directly helpful. It delivers more accurate answers with better contextualization and significantly...
This model always redirects to the latest model in the DeepSeek V4 Flash family.
NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and...
Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32...
Hermes 4 70B is a hybrid reasoning model from Nous Research, built on Meta-Llama-3.1-70B. It introduces the same hybrid mode as the larger 405B release, allowing the model to either...
Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features a 256k context window and can generate outputs of...
Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a 256K context window.
GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
Fast-mode variant of [Opus 5](/anthropic/claude-opus-5) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 5. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode Note: As of September...
Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...
GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization. Designed with the latest transformer architecture, it...
GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...
Fast-mode variant of [Opus 4.8](/anthropic/claude-opus-4.8) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 4.8. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode Note: As of September...
KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...
Auto Router (Beta) is a task-aware router from OpenRouter. It classifies each request, then routes it the [most popular model](/rankings#task-spend) for that task based on aggregate spend, filtered by your...
KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...
This model always redirects to the latest Grok model from xAI.
Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative generation process in which multiple specialized models each...
Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in which multiple specialized models each contribute...
North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...
Nex-N2-Mini is an open-source agentic mixture-of-experts model from Nex AGI, the smaller sibling in the Nex-N2 series. It accepts text and image input and is built for coding, tool use,...
GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...
Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fetch enabled, then a...
This model always redirects to the latest model in the Claude Fable family.
Fast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...
Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...
Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast execution and high efficiency at scale. It uses a “fast...
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...
Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...
Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....
The Pareto Router maintains a tiered shortlist of strong coding models, ranked by [Artificial Analysis](https://artificialanalysis.ai/) coding percentiles. Set min_coding_score between 0 and 1 on the [pareto-router plugin](https://openrouter.ai/docs/guides/routing/routers/pareto-router#the-min_coding_score-parameter) to control how...
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
[GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and...
This model always redirects to the latest model in the Claude Opus family.
KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration. It builds on the agentic coding strengths of earlier versions,...
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Anthropic: Claude Opus 4.7 (batch)anthropic/claude-opus-4.7:batch | 1M | $2.5 | $12.5 | — | |||
| Z.ai: GLM 5V Turboz-ai/glm-5v-turbo | 202.752K | $1.2 | $4 | — | |||
| OpenAI: GPT-4o (batch)openai/gpt-4o:batch | 128K | $1.25 | $5 | — | |||
| DeepSeek: DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp | 163.84K | $0.27 | $0.41 | — | |||
| Qwen: Qwen3 Coder 480B A35Bqwen/qwen3-coder | 262.144K | $0.3 | $1 | — | |||
| OpenAI: o3 (batch)openai/o3:batch | 200K | $1 | $4 | — | |||
| Qwen: Qwen3.6 Flashqwen/qwen3.6-flash | 1M | $0.188 | $1.125 | — | |||
| OpenAI: GPT-5.3 Chatopenai/gpt-5.3-chat | 128K | $1.75 | $14 | — | |||
| DeepSeek: DeepSeek V4 Flash Latest~deepseek/deepseek-v4-flash-latest | 1.04858M | $0.04 | $0.1 | — | |||
| NVIDIA: Nemotron Nano 9B V2 (free)nvidia/nemotron-nano-9b-v2:free | 128K | Free | Free | — | |||
| MoonshotAI: Kimi K2 0905moonshotai/kimi-k2-0905 | 262.144K | $0.6 | $2.5 | — | |||
| Nous: Hermes 4 70Bnousresearch/hermes-4-70b | 131.072K | $0.13 | $0.4 | — | |||
| ByteDance Seed: Seed 1.6 Flashbytedance-seed/seed-1.6-flash | 262.144K | $0.075 | $0.3 | — | |||
| ByteDance Seed: Seed 1.6bytedance-seed/seed-1.6 | 262.144K | $0.25 | $2 | — | |||
| OpenAI: GPT-5.6 Luna Proopenai/gpt-5.6-luna-pro | 1.05M | $0.2 | $1.2 | — | |||
| Qwen: Qwen3.7 Flashqwen/qwen3.7-flash | 1M | $0.03 | $0.13 | — | |||
| Ling-3.0-flash (free)inclusionai/ling-3.0-flash:free | 262.144K | Free | Free | — | |||
| Claude Opus 5 (Fast)anthropic/claude-opus-5-fast | 1M | $10 | $50 | — | |||
| Poolside: Laguna S 2.1 (free)poolside/laguna-s-2.1:free | 262.144K | Free | Free | — | |||
| OpenAI: GPT-5.6 Terra Proopenai/gpt-5.6-terra-pro | 1.05M | $2 | $12 | — | |||
| Meta: Llama 3.2 3B Instructmeta-llama/llama-3.2-3b-instruct | 131.072K | $0.05 | $0.33 | — | |||
| OpenAI: GPT-5.6 Sol Proopenai/gpt-5.6-sol-pro | 1.05M | $2 | $10 | — | |||
| Qwen: Qwen3 VL 8B Instructqwen/qwen3-vl-8b-instruct | 131.072K | $0.117 | $0.455 | — | |||
| Anthropic: Claude Opus 4.8 (Fast)anthropic/claude-opus-4.8-fast | 1M | $10 | $50 | — | |||
| Kwaipilot: KAT-Coder-Air V2.5kwaipilot/kat-coder-air-v2.5 | 256K | $0.15 | $0.6 | — | |||
| Auto Router (Beta)openrouter/auto-beta | 2M | — | — | — | |||
| Kwaipilot: KAT-Coder-Pro V2.5kwaipilot/kat-coder-pro-v2.5 | 262.144K | $0.74 | $2.96 | — | |||
| xAI: Grok Latest~x-ai/grok-latest | 500K | $2 | $6 | — | |||
| AionLabs: Aion-3.0-Miniaion-labs/aion-3.0-mini | 131.072K | $0.7 | $1.4 | — | |||
| AionLabs: Aion-3.0aion-labs/aion-3.0 | 131.072K | $3 | $6 | — | |||
| Cohere: North Mini Code (free)cohere/north-mini-code:free | 256K | Free | Free | — | |||
| Nex AGI: Nex-N2-Mininex-agi/nex-n2-mini | 262.144K | $0.025 | $0.1 | — | |||
| Z.ai: GLM 5.2z-ai/glm-5.2 | 1.04858M | $1.4 | $4.4 | — | |||
| OpenRouter: Fusionopenrouter/fusion | 1M | — | — | — | |||
| Anthropic: Claude Fable Latest~anthropic/claude-fable-latest | 1M | $10 | $50 | — | |||
| Anthropic: Claude Opus 4.7 (Fast)anthropic/claude-opus-4.7-fast | 1M | $30 | $150 | — | |||
| Qwen: Qwen3.7 Maxqwen/qwen3.7-max | 1M | $1.475 | $4.425 | — | |||
| SpaceXAI: Grok Build 0.1x-ai/grok-build-0.1 | 256K | $1 | $2 | — | |||
| inclusionAI: Ling-2.6-1Tinclusionai/ling-2.6-1t | 262.144K | $0.075 | $0.625 | — | |||
| Qwen: Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | 262.144K | $0.1 | $0.9 | — | |||
| Qwen: Qwen3.6 Max Previewqwen/qwen3.6-max-preview | 262.144K | $1.027 | $6.162 | — | |||
| Qwen: Qwen3.6 27Bqwen/qwen3.6-27b | 262.144K | $0.3 | $2 | — | |||
| inclusionAI: Ling-2.6-flashinclusionai/ling-2.6-flash | 262.144K | $0.01 | $0.03 | — | |||
| Pareto Code Routeropenrouter/pareto-code | 2M | — | — | — | |||
| Mistral: Mistral Small 4mistralai/mistral-small-2603 | 262.144K | $0.15 | $0.6 | — | |||
| OpenAI: GPT-5.4 Image 2openai/gpt-5.4-image-2 | 272K | $8 | $15 | — | |||
| Anthropic: Claude Opus Latest~anthropic/claude-opus-latest | 1M | $5 | $25 | — | |||
| Kwaipilot: KAT-Coder-Pro V2kwaipilot/kat-coder-pro-v2 | 262.144K | $0.3 | $1.2 | — | |||
| NVIDIA: Nemotron 3 Super (free)nvidia/nemotron-3-super-120b-a12b:free | 262.144K | Free | Free | — | |||
| Qwen: Qwen3.5-9Bqwen/qwen3.5-9b | 262.144K | $0.1 | $0.15 | — |