Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...
Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....
[GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and...
This model always redirects to the latest model in the Claude Opus family.
The Pareto Router maintains a tiered shortlist of strong coding models, ranked by [Artificial Analysis](https://artificialanalysis.ai/) coding percentiles. Set min_coding_score between 0 and 1 on the [pareto-router plugin](https://openrouter.ai/docs/guides/routing/routers/pareto-router#the-min_coding_score-parameter) to control how...
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration. It builds on the agentic coding strengths of earlier versions,...
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...
The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of...
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...
Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive roleplaying and storytelling. It is particularly strong at introducing tension, crises, and conflict into stories, making narratives feel more engaging....
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...
Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized...
The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced...
MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world...
GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...
GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on...
The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user request. In contrast to RAG, relace-search performs agentic...
Cogito v2.1 671B MoE represents one of the strongest open models globally, matching performance of frontier closed and open models. This model is trained using self play with reinforcement learning...
GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...
Transform your natural language requests into structured OpenRouter API request objects. Describe what you want to accomplish with AI models, and Body Builder will construct the appropriate API calls. Example:...
Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...
The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.
Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling custom models.
Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed for deeper reasoning and analysis. Pricing is based...
[GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,...
Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Thinking variant enhances reasoning in STEM, math, and complex tasks. It excels...
Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...
Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction following, multilingual understanding, and...
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...
Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in collaboration with Venice.ai. This model is designed as an “uncensored” instruct-tuned LLM, preserving...
Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B and support for reasoning via Chain-of-Thought. It offers competitive benchmark...
Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code>...
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, with 45.9 billion parameters activated per inference, and can handle a context...
Euryale L3.3 70B is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.2](/models/sao10k/l3-euryale-70b).
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output. Amazon Nova Lite...
Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost. With a context length...
Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This...
Euryale L3.1 70B v2.2 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.1](/models/sao10k/l3-euryale-70b).
Hermes 3 is a generalist language model with many improvements over [Hermes 2](/models/nousresearch/nous-hermes-2-mistral-7b-dpo), including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Qwen: Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | 262.144K | $0.1 | $0.9 | — | |||
| Qwen: Qwen3.6 Max Previewqwen/qwen3.6-max-preview | 262.144K | $1.027 | $6.162 | — | |||
| Qwen: Qwen3.6 27Bqwen/qwen3.6-27b | 262.144K | $0.3 | $2 | — | |||
| inclusionAI: Ling-2.6-flashinclusionai/ling-2.6-flash | 262.144K | $0.01 | $0.03 | — | |||
| OpenAI: GPT-5.4 Image 2openai/gpt-5.4-image-2 | 272K | $8 | $15 | — | |||
| Anthropic: Claude Opus Latest~anthropic/claude-opus-latest | 1M | $5 | $25 | — | |||
| Pareto Code Routeropenrouter/pareto-code | 2M | — | — | — | |||
| Mistral: Mistral Small 4mistralai/mistral-small-2603 | 262.144K | $0.15 | $0.6 | — | |||
| Kwaipilot: KAT-Coder-Pro V2kwaipilot/kat-coder-pro-v2 | 262.144K | $0.3 | $1.2 | — | |||
| NVIDIA: Nemotron 3 Super (free)nvidia/nemotron-3-super-120b-a12b:free | 262.144K | Free | Free | — | |||
| Qwen: Qwen3.5-9Bqwen/qwen3.5-9b | 262.144K | $0.1 | $0.15 | — | |||
| Qwen: Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | 262.144K | $0.163 | $1.3 | — | |||
| Qwen: Qwen3.5-27Bqwen/qwen3.5-27b | 262.144K | $0.195 | $1.56 | — | |||
| Qwen: Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | 262.144K | $0.26 | $2.08 | — | |||
| AionLabs: Aion-2.0aion-labs/aion-2.0 | 131.072K | $0.8 | $1.6 | — | |||
| Qwen: Qwen3.5 Plus 2026-02-15qwen/qwen3.5-plus-02-15 | 1M | $0.26 | $1.56 | — | |||
| Upstage: Solar Pro 3upstage/solar-pro-3 | 131.072K | $0.15 | $0.6 | — | |||
| OpenAI: GPT Audioopenai/gpt-audio | 128K | $2.5 | $10 | — | |||
| MiniMax: MiniMax M2.1minimax/minimax-m2.1 | 204.8K | $0.3 | $1.2 | — | |||
| Z.ai: GLM 4.7z-ai/glm-4.7 | 202.752K | $0.4 | $1.75 | — | |||
| OpenAI: GPT-5.2 Chatopenai/gpt-5.2-chat | 128K | $1.75 | $14 | — | |||
| Relace: Relace Searchrelace/relace-search | 256K | $1 | $3 | — | |||
| Deep Cogito: Cogito v2.1 671Bdeepcogito/cogito-v2.1-671b | 128K | $1.25 | $1.25 | — | |||
| Z.ai: GLM 4.6Vz-ai/glm-4.6v | 131.072K | $0.3 | $0.9 | — | |||
| Body Builder (beta)openrouter/bodybuilder | 128K | — | — | — | |||
| Amazon: Nova 2 Liteamazon/nova-2-lite-v1 | 1M | $0.3 | $2.5 | — | |||
| Mistral: Ministral 3 3B 2512mistralai/ministral-3b-2512 | 131.072K | $0.1 | $0.1 | — | |||
| Amazon: Nova Premier 1.0amazon/nova-premier-v1 | 1M | $2.5 | $12.5 | — | |||
| Perplexity: Sonar Pro Searchperplexity/sonar-pro-search | 200K | $3 | $15 | — | |||
| OpenAI: GPT-5 Imageopenai/gpt-5-image | 400K | $10 | $10 | — | |||
| Qwen: Qwen3 VL 30B A3B Thinkingqwen/qwen3-vl-30b-a3b-thinking | 131.072K | $0.2 | $2.4 | — | |||
| Qwen: Qwen3 VL 30B A3B Instructqwen/qwen3-vl-30b-a3b-instruct | 262.144K | $0.15 | $0.6 | — | |||
| Qwen: Qwen3 30B A3B Instruct 2507qwen/qwen3-30b-a3b-instruct-2507 | 128K | $0.048 | $0.193 | — | |||
| DeepSeek: DeepSeek V3.1deepseek/deepseek-chat-v3.1 | 163.84K | $0.25 | $0.95 | — | |||
| OpenAI: gpt-oss-20b (free)openai/gpt-oss-20b:free | 131.072K | Free | Free | — | |||
| Z.ai: GLM 4.5z-ai/glm-4.5 | 131.072K | $0.6 | $2.2 | — | |||
| Venice: Uncensoredcognitivecomputations/dolphin-mistral-24b-venice-edition | 128K | $0.2 | $0.9 | — | |||
| Tencent: Hunyuan A13B Instructtencent/hunyuan-a13b-instruct | 131.072K | $0.14 | $0.57 | — | |||
| Morph: Morph V3 Largemorph/morph-v3-large | 262.144K | $0.9 | $1.9 | — | |||
| Google: Gemini 2.5 Pro Preview 06-05google/gemini-2.5-pro-preview | 1.04858M | $1.25 | $10 | — | |||
| Meta: Llama 4 Maverickmeta-llama/llama-4-maverick | 128K | $0.188 | $0.652 | — | |||
| OpenAI: o3 Mini Highopenai/o3-mini-high | 200K | $1.1 | $4.4 | — | |||
| MiniMax: MiniMax-01minimax/minimax-01 | 1.00019M | $0.2 | $1.1 | — | |||
| Sao10K: Llama 3.3 Euryale 70Bsao10k/l3.3-euryale-70b | 131.072K | $0.65 | $0.75 | — | |||
| Meta: Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | 131.072K | $0.1 | $0.32 | — | |||
| Amazon: Nova Lite 1.0amazon/nova-lite-v1 | 300K | $0.06 | $0.24 | — | |||
| Amazon: Nova Micro 1.0amazon/nova-micro-v1 | 128K | $0.035 | $0.14 | — | |||
| Qwen: Qwen3.5 Plus 2026-04-20qwen/qwen3.5-plus-20260420 | 1M | $0.3 | $1.8 | — | |||
| Sao10K: Llama 3.1 Euryale 70B v2.2sao10k/l3.1-euryale-70b | 131.072K | $0.85 | $0.85 | — | |||
| Nous: Hermes 3 70B Instructnousresearch/hermes-3-llama-3.1-70b | 131.072K | $0.7 | $0.7 | — |