Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
No provider description is available for this model yet.
Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...
Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and multilingual text analysis. Its smaller size allows it to operate...
GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
No provider description is available for this model yet.
GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...
Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...
Fast-mode variant of [Opus 4.8](/anthropic/claude-opus-4.8) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 4.8. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode Note: As of September...
KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...
No provider description is available for this model yet.
KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...
This model always redirects to the latest Grok model from xAI.
Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative generation process in which multiple specialized models each...
Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in which multiple specialized models each contribute...
North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...
No provider description is available for this model yet.
GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...
Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fetch enabled, then a...
This model always redirects to the latest model in the Claude Fable family.
Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired with natural language queries, and produces detailed visual understanding...
Fast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode
No provider description is available for this model yet.
Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...
Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast execution and high efficiency at scale. It uses a “fast...
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...
Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...
Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....
The Pareto Router maintains a tiered shortlist of strong coding models, ranked by [Artificial Analysis](https://artificialanalysis.ai/) coding percentiles. Set min_coding_score between 0 and 1 on the [pareto-router plugin](https://openrouter.ai/docs/guides/routing/routers/pareto-router#the-min_coding_score-parameter) to control how...
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
[GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and...
This model always redirects to the latest model in the Claude Opus family.
KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration. It builds on the agentic coding strengths of earlier versions,...
Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimized specifically to deliver industry-leading performance in image understanding,...
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...
The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of...
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...
No provider description is available for this model yet.
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...
Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized...
MiniMax M2-her is a dialogue-first large language model built for immersive roleplay, character-driven chat, and expressive multi-turn conversations. Designed to stay consistent in tone and personality, it supports rich message...
No provider description is available for this model yet.
MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world...
GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...
No provider description is available for this model yet.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Qwen: Qwen3.7 Flashqwen/qwen3.7-flash | 1M | $0.03 | $0.13 | — | |||
| Ling-3.0-flash (free)inclusionai/ling-3.0-flash:free | 262.144K | Free | Free | — | |||
| qwen3-coder-flashqwencloud/qwen3-coder-flash | 997.952K | — | — | — | |||
| Poolside: Laguna S 2.1 (free)poolside/laguna-s-2.1:free | 262.144K | Free | Free | — | |||
| Meta: Llama 3.2 1B Instructmeta-llama/llama-3.2-1b-instruct | 60K | $0.027 | $0.201 | — | |||
| OpenAI: GPT-5.6 Terra Proopenai/gpt-5.6-terra-pro | 1.05M | $2 | $12 | — | |||
| zai-org/GLM-4.7-Flashdeepinfra/zai-org/glm-4.7-flash | 202.752K | $0.06 | $0.4 | — | |||
| OpenAI: GPT-5.6 Sol Proopenai/gpt-5.6-sol-pro | 1.05M | $2 | $10 | — | |||
| Qwen2.5 72B Instructqwen/qwen-2.5-72b-instruct | 32.768K | $0.36 | $0.4 | — | |||
| Qwen: Qwen3 VL 8B Instructqwen/qwen3-vl-8b-instruct | 131.072K | $0.117 | $0.455 | — | |||
| Anthropic: Claude Opus 4.8 (Fast)anthropic/claude-opus-4.8-fast | 1M | $10 | $50 | — | |||
| Kwaipilot: KAT-Coder-Air V2.5kwaipilot/kat-coder-air-v2.5 | 256K | $0.15 | $0.6 | — | |||
| mistralai/mistral-small-2603openrouter/mistralai/mistral-small-2603 | 262.144K | $0.15 | $0.6 | — | |||
| Kwaipilot: KAT-Coder-Pro V2.5kwaipilot/kat-coder-pro-v2.5 | 262.144K | $0.74 | $2.96 | — | |||
| xAI: Grok Latest~x-ai/grok-latest | 500K | $2 | $6 | — | |||
| AionLabs: Aion-3.0-Miniaion-labs/aion-3.0-mini | 131.072K | $0.7 | $1.4 | — | |||
| AionLabs: Aion-3.0aion-labs/aion-3.0 | 131.072K | $3 | $6 | — | |||
| Cohere: North Mini Code (free)cohere/north-mini-code:free | 256K | Free | Free | — | |||
| minimax/minimax-m2.7:freeopenrouter/minimax/minimax-m2.7:free | 196.608K | Free | Free | — | |||
| Z.ai: GLM 5.2z-ai/glm-5.2 | 202.752K | $0.6 | $2 | — | |||
| OpenRouter: Fusionopenrouter/fusion | 1M | — | — | — | |||
| Anthropic: Claude Fable Latest~anthropic/claude-fable-latest | 1M | $10 | $50 | — | |||
| Perceptron: Perceptron Mk1perceptron/perceptron-mk1 | 32.768K | $0.15 | $1.5 | — | |||
| Anthropic: Claude Opus 4.7 (Fast)anthropic/claude-opus-4.7-fast | 1M | $30 | $150 | — | |||
| qwen3-30b-a3bqwencloud/qwen3-30b-a3b | 129.024K | — | — | — | |||
| SpaceXAI: Grok Build 0.1x-ai/grok-build-0.1 | 256K | $1 | $2 | — | |||
| inclusionAI: Ling-2.6-1Tinclusionai/ling-2.6-1t | 262.144K | $0.075 | $0.625 | — | |||
| Qwen: Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | 262.144K | $0.1 | $0.9 | — | |||
| Qwen: Qwen3.6 Max Previewqwen/qwen3.6-max-preview | 262.144K | $1.027 | $6.162 | — | |||
| Qwen: Qwen3.6 27Bqwen/qwen3.6-27b | 262.144K | $0.3 | $2 | — | |||
| inclusionAI: Ling-2.6-flashinclusionai/ling-2.6-flash | 262.144K | $0.01 | $0.03 | — | |||
| Pareto Code Routeropenrouter/pareto-code | 2M | — | — | — | |||
| Mistral: Mistral Small 4mistralai/mistral-small-2603 | 262.144K | $0.15 | $0.6 | — | |||
| OpenAI: GPT-5.4 Image 2openai/gpt-5.4-image-2 | 272K | $8 | $15 | — | |||
| Anthropic: Claude Opus Latest~anthropic/claude-opus-latest | 1M | $5 | $25 | — | |||
| Kwaipilot: KAT-Coder-Pro V2kwaipilot/kat-coder-pro-v2 | 262.144K | $0.3 | $1.2 | — | |||
| Reka Edgerekaai/reka-edge | 16.384K | $0.1 | $0.1 | — | |||
| NVIDIA: Nemotron 3 Super (free)nvidia/nemotron-3-super-120b-a12b:free | 262.144K | Free | Free | — | |||
| Qwen: Qwen3.5-9Bqwen/qwen3.5-9b | 262.144K | $0.1 | $0.15 | — | |||
| Qwen: Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | 256K | $0.312 | $1.25 | — | |||
| Qwen: Qwen3.5-27Bqwen/qwen3.5-27b | 262.144K | $0.195 | $1.56 | — | |||
| Qwen: Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | 262.144K | $0.26 | $2.08 | — | |||
| minimax/minimax-m2.7openrouter/minimax/minimax-m2.7 | 204.8K | $0.3 | $1.2 | — | |||
| Qwen: Qwen3.5 Plus 2026-02-15qwen/qwen3.5-plus-02-15 | 1M | $0.26 | $1.56 | — | |||
| Upstage: Solar Pro 3upstage/solar-pro-3 | 131.072K | $0.15 | $0.6 | — | |||
| MiniMax: MiniMax M2-herminimax/minimax-m2-her | 65.536K | $0.3 | $1.2 | — | |||
| z-ai/glm-5v-turboopenrouter/z-ai/glm-5v-turbo | 202.752K | $1.2 | $4 | — | |||
| MiniMax: MiniMax M2.1minimax/minimax-m2.1 | 204.8K | $0.3 | $1.2 | — | |||
| Z.ai: GLM 4.7z-ai/glm-4.7 | 202.752K | $0.4 | $1.75 | — | |||
| qwen-turbo-latestqwencloud/qwen-turbo-latest | 1M | $0.05 | $0.2 | — |