Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...
Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing in autonomous programming via tool calling...
The preview GPT-4 model with improved instruction following, JSON mode, reproducible outputs, parallel function calling, and more. Training data: up to Dec 2023. **Note:** heavily rate limited by OpenAI while...
Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...
Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech transcription, translation and audio understanding. Input audio...
NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a hybrid Transformer-Mamba architecture, combining transformer-level accuracy with Mamba’s...
Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models. These models are the latest in a series of models released by IBM. They are fine-tuned for long...
Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual...
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model, designed for advanced visual and textual reasoning across complex scenes, documents, and temporal sequences. It integrates enhanced multimodal alignment and...
Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the...
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...
ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with 47B active per token. It is trained jointly on text and image data...
May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...
Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reasoning, multilingual support, and advanced agent tasks. Its unique...
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...
Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary...
Reka Flash 3 is a general-purpose, instruction-tuned large language model with 21 billion parameters, developed by Reka. It excels at general chat, coding tasks, instruction-following, and function calling. Featuring a...
The Auto Router automatically selects the best model for your prompt, powered by the wisdom of the market. It routes you based on what the OpenRouter community collectively spends on...
Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of tasks. As of December...
This model is a variant of GPT-3.5 Turbo tuned for instructional prompts and omitting chat-related optimizations. Training data: up to Sep 2021.
This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request at a higher cost. Training data: up...
This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement [here](https://mistral.ai/news/mistral-large-2407/)....
Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). Qwen2.5-Coder brings the following improvements upon CodeQwen1.5: - Significantly improvements in **code generation**, **code reasoning**...
UnslopNemo v4.1 is the latest addition from the creator of Rocinante, designed for adventure writing and role-play scenarios.
This is a series of models designed to replicate the prose quality of the Claude 3 models, specifically Sonnet(https://openrouter.ai/anthropic/claude-3.5-sonnet) and Opus(https://openrouter.ai/anthropic/claude-3-opus). The model is fine-tuned on top of [Qwen2.5 72B](https://openrouter.ai/qwen/qwen-2.5-72b-instruct).
Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...
An attempt to recreate Claude-style verbosity, but don't expect the same level of coherence or memory. Meant for use in roleplay/narrative situations.
Mistral's official instruct fine-tuned version of [Mixtral 8x22B](/models/mistralai/mixtral-8x22b). It uses 39B active parameters out of 141B, offering unparalleled cost efficiency for its size. Its strengths include: - strong math, coding,...
WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to leading proprietary models, and it consistently outperforms all existing state-of-the-art opensource models. It is...
GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.
KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...
No provider description is available for this model yet.
KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...
No provider description is available for this model yet.
Laguna M.1 is the flagship coding agent model from [Poolside](https://poolside.ai/), optimized for complex software engineering tasks. Designed for agentic coding workflows, it supports tool calling and reasoning, with a 256K...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Anthropic: Claude Opus 4.6anthropic/claude-opus-4.6 | 1M | $5 | $25 | — | |||
| NVIDIA: Nemotron 3 Nano 30B A3B (free)nvidia/nemotron-3-nano-30b-a3b:free | 256K | Free | Free | — | |||
| Mistral: Ministral 3 14B 2512mistralai/ministral-14b-2512 | 262.144K | $0.2 | $0.2 | — | |||
| Qwen: Qwen3 Coder Flashqwen/qwen3-coder-flash | 1M | $0.195 | $0.975 | — | |||
| OpenAI: GPT-4 Turbo Previewopenai/gpt-4-turbo-preview | 128K | $10 | $30 | — | |||
| Anthropic: Claude Opus 4.5anthropic/claude-opus-4.5 | 200K | $5 | $25 | — | |||
| Mistral: Voxtral Small 24B 2507mistralai/voxtral-small-24b-2507 | 32.768K | $0.1 | $0.3 | — | |||
| NVIDIA: Nemotron Nano 12B 2 VL (free)nvidia/nemotron-nano-12b-v2-vl:free | 128K | Free | Free | — | |||
| IBM: Granite 4.0 Microibm-granite/granite-4.0-h-micro | 131K | $0.017 | $0.112 | — | |||
| Qwen: Qwen3 Next 80B A3B Instructqwen/qwen3-next-80b-a3b-instruct | 262.144K | $0.09 | $1.1 | — | |||
| Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5 | 200K | $1 | $5 | — | |||
| Qwen: Qwen3 VL 8B Thinkingqwen/qwen3-vl-8b-thinking | 131.072K | $0.18 | $2.1 | — | |||
| Qwen: Qwen3 Coder 30B A3B Instructqwen/qwen3-coder-30b-a3b-instruct | 262.144K | $0.07 | $0.28 | — | |||
| Qwen: Qwen3 235B A22B Instruct 2507qwen/qwen3-235b-a22b-2507 | 262.144K | $0.087 | $0.35 | — | |||
| Baidu: ERNIE 4.5 VL 424B A47B baidu/ernie-4.5-vl-424b-a47b | 123K | $0.42 | $1.25 | — | |||
| DeepSeek: R1 0528deepseek/deepseek-r1-0528 | 163.84K | $0.5 | $2.15 | — | |||
| Qwen: Qwen3 30B A3Bqwen/qwen3-30b-a3b | 40.96K | $0.12 | $0.5 | — | |||
| Qwen: Qwen3 8Bqwen/qwen3-8b | 131.072K | $0.117 | $0.455 | — | |||
| Qwen: Qwen3 32Bqwen/qwen3-32b | 40.96K | $0.08 | $0.28 | — | |||
| Meta: Llama 4 Scoutmeta-llama/llama-4-scout | 327.68K | $0.1 | $0.3 | — | |||
| DeepSeek: DeepSeek V3 0324deepseek/deepseek-chat-v3-0324 | 163.84K | $0.25 | $1 | — | |||
| Cohere: Command Acohere/command-a | 256K | $2.5 | $10 | — | |||
| Reka Flash 3rekaai/reka-flash-3 | 65.536K | $0.1 | $0.2 | — | |||
| Auto Routeropenrouter/auto | 2M | — | — | — | |||
| Amazon: Nova Pro 1.0amazon/nova-pro-v1 | 300K | $0.8 | $3.2 | — | |||
| OpenAI: GPT-3.5 Turbo Instructopenai/gpt-3.5-turbo-instruct | 4.095K | $1.5 | $2 | — | |||
| OpenAI: GPT-3.5 Turbo 16kopenai/gpt-3.5-turbo-16k | 16.385K | $3 | $4 | — | |||
| Mistral Large 2407mistralai/mistral-large-2407 | 131.072K | $2 | $6 | — | |||
| Qwen2.5 Coder 32B Instructqwen/qwen-2.5-coder-32b-instruct | 32.768K | $0.66 | $1 | — | |||
| TheDrummer: UnslopNemo 12Bthedrummer/unslopnemo-12b | 1.024M | $0.4 | $0.4 | — | |||
| Magnum v4 72Banthracite-org/magnum-v4-72b | 32.768K | $2.5 | $5 | — | |||
| Qwen: Qwen2.5 7B Instructqwen/qwen-2.5-7b-instruct | 32.768K | $0.1 | $0.2 | — | |||
| Mancer: Weaver (alpha)mancer/weaver | 8K | $0.4 | $0.75 | — | |||
| Mistral: Mixtral 8x22B Instructmistralai/mixtral-8x22b-instruct | 65.536K | $2 | $6 | — | |||
| WizardLM-2 8x22Bmicrosoft/wizardlm-2-8x22b | 65.535K | $0.62 | $0.62 | — | |||
| OpenAI: GPT-3.5 Turbo (older v0613)openai/gpt-3.5-turbo-0613 | 4.095K | $1 | $2 | — | |||
| Kwaipilot: KAT-Coder-Pro V2.5 (free)kwaipilot/kat-coder-pro-v2.5:free | 256K | Free | Free | — | |||
| nvidia/Riva-Translate-4B-Instruct-v1.1nvidia/Riva-Translate-4B-Instruct-v1.1 | Not documented | — | — | — | |||
| Kwaipilot: KAT-Coder-Air V2.5 (free)kwaipilot/kat-coder-air-v2.5:free | 256K | Free | Free | — | |||
| nvidia/Nemotron-3-Labs-Ultra-Math-SFTnvidia/Nemotron-3-Labs-Ultra-Math-SFT | Not documented | — | — | — | |||
| ap-south-1/minimax.minimax-m2.5bedrock/ap-south-1/minimax.minimax-m2.5 | 1M | $0.36 | $1.44 | — | |||
| ap-south-1/moonshotai.kimi-k2-thinkingbedrock/ap-south-1/moonshotai.kimi-k2-thinking | 262.144K | $0.71 | $2.94 | — | |||
| nvidia/Nemotron-3-Labs-Ultra-Math-RLnvidia/Nemotron-3-Labs-Ultra-Math-RL | Not documented | — | — | — | |||
| OpenAI: GPT-6 Astra (batch)openai/gpt-6-astra:batch | 1.05M | $5 | $25 | — | |||
| Tencent: Hy3 (free)tencent/hy3:free | 262.144K | Free | Free | — | |||
| ap-south-1/moonshotai.kimi-k2.5bedrock/ap-south-1/moonshotai.kimi-k2.5 | 262.144K | $0.72 | $3.6 | — | |||
| Poolside: Laguna M.1 (free)poolside/laguna-m.1:free | 262.144K | Free | Free | — | |||
| ap-south-1/qwen.qwen3-coder-nextbedrock/ap-south-1/qwen.qwen3-coder-next | 262.144K | $0.6 | $1.44 | — | |||
| ap-southeast-2/minimax.minimax-m2.5bedrock/ap-southeast-2/minimax.minimax-m2.5 | 1M | $0.309 | $1.236 | — | |||
| ap-southeast-3/deepseek.v3.2bedrock/ap-southeast-3/deepseek.v3.2 | 163.84K | $0.74 | $2.22 | — |