No provider description is available for this model yet.
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
No provider description is available for this model yet.
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...
KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...
Auto Router (Beta) is a task-aware router from OpenRouter. It classifies each request, then routes it the [most popular model](/rankings#task-spend) for that task based on aggregate spend, filtered by your...
LFM2.5-1.2B-Thinking is a lightweight reasoning-focused model optimized for agentic tasks, data extraction, and RAG—while still running comfortably on edge devices. It supports long context (up to 32K tokens) and is...
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...
LFM2.5-1.2B-Instruct is a compact, high-performance instruction-tuned model built for fast on-device AI. It delivers strong chat quality in a 1.2B parameter footprint, with efficient edge inference and broad runtime support.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Fast-mode variant of [Opus 4.8](/anthropic/claude-opus-4.8) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 4.8. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode Note: As of September...
Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...
No provider description is available for this model yet.
No provider description is available for this model yet.
Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...
GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
No provider description is available for this model yet.
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as...
No provider description is available for this model yet.
Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization. Designed with the latest transformer architecture, it...
GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual...
No provider description is available for this model yet.
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
No provider description is available for this model yet.
Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and multilingual text analysis. Its smaller size allows it to operate...
No provider description is available for this model yet.
Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| eu-central-1/anthropic.claude-v2:1bedrock/eu-central-1/anthropic.claude-v2:1 | 100K | $8 | $24 | — | |||
| @cf/zai-org/glm-5.2cloudflare/@cf/zai-org/glm-5.2 | 262.144K | $1.4 | $4.4 | — | |||
| OpenAI: gpt-oss-120b (batch)openai/gpt-oss-120b:batch | 131.072K | $0.15 | $0.6 | — | |||
| eu-central-1/minimax.minimax-m2.1bedrock/eu-central-1/minimax.minimax-m2.1 | 196K | $0.36 | $1.44 | — | |||
| eu-central-1/minimax.minimax-m2.5bedrock/eu-central-1/minimax.minimax-m2.5 | 1M | $0.36 | $1.44 | — | |||
| eu-central-1/qwen.qwen3-coder-nextbedrock/eu-central-1/qwen.qwen3-coder-next | 262.144K | $0.6 | $1.44 | — | |||
| eu-west-1/minimax.minimax-m2.1bedrock/eu-west-1/minimax.minimax-m2.1 | 196K | $0.36 | $1.44 | — | |||
| eu-west-1/minimax.minimax-m2.5bedrock/eu-west-1/minimax.minimax-m2.5 | 1M | $0.36 | $1.44 | — | |||
| Google: Gemini 2.5 Flash Lite (batch)google/gemini-2.5-flash-lite:batch | 1.04858M | $0.05 | $0.2 | — | |||
| Kwaipilot: KAT-Coder-Air V2.5kwaipilot/kat-coder-air-v2.5 | 256K | $0.15 | $0.6 | — | |||
| Auto Router (Beta)openrouter/auto-beta | 2M | — | — | — | |||
| LiquidAI: LFM2.5-1.2B-Thinking (free)liquid/lfm-2.5-1.2b-thinking:free | 32.768K | Free | Free | — | |||
| Inception: Mercury 2.5inception/mercury-2.5 | 260K | $0.04 | $0.15 | — | |||
| LiquidAI: LFM2.5-1.2B-Instruct (free)liquid/lfm-2.5-1.2b-instruct:free | 32.768K | Free | Free | — | |||
| eu-west-1/qwen.qwen3-coder-nextbedrock/eu-west-1/qwen.qwen3-coder-next | 262.144K | $0.6 | $1.44 | — | |||
| eu-west-2/minimax.minimax-m2.1bedrock/eu-west-2/minimax.minimax-m2.1 | 196K | $0.47 | $1.86 | — | |||
| eu-west-2/minimax.minimax-m2.5bedrock/eu-west-2/minimax.minimax-m2.5 | 1M | $0.47 | $1.86 | — | |||
| eu-west-2/qwen.qwen3-coder-nextbedrock/eu-west-2/qwen.qwen3-coder-next | 262.144K | $0.78 | $1.86 | — | |||
| eu-west-3/mistral.mistral-7b-instruct-v0:2bedrock/eu-west-3/mistral.mistral-7b-instruct-v0:2 | 32K | $0.2 | $0.26 | — | |||
| deepseek-ai/DeepSeek-V4-Pro-0813deepinfra/deepseek-ai/deepseek-v4-pro-0813 | 1.04858M | $1.3 | $2.6 | — | |||
| eu-west-3/mistral.mistral-large-2402-v1:0bedrock/eu-west-3/mistral.mistral-large-2402-v1:0 | 32K | $10.4 | $31.2 | — | |||
| eu-west-3/mistral.mixtral-8x7b-instruct-v0:1bedrock/eu-west-3/mistral.mixtral-8x7b-instruct-v0:1 | 32K | $0.59 | $0.91 | — | |||
| eu-south-1/minimax.minimax-m2.1bedrock/eu-south-1/minimax.minimax-m2.1 | 196K | $0.36 | $1.44 | — | |||
| eu-south-1/minimax.minimax-m2.5bedrock/eu-south-1/minimax.minimax-m2.5 | 1M | $0.36 | $1.44 | — | |||
| eu-south-1/qwen.qwen3-coder-nextbedrock/eu-south-1/qwen.qwen3-coder-next | 262.144K | $0.6 | $1.44 | — | |||
| llama3.2-1bsnowflake/llama3.2-1b | 128K | — | — | — | |||
| Anthropic: Claude Opus 4.8 (Fast)anthropic/claude-opus-4.8-fast | 1M | $10 | $50 | — | |||
| Qwen: Qwen3 VL 8B Instructqwen/qwen3-vl-8b-instruct | 131.072K | $0.117 | $0.455 | — | |||
| moonshotai/Kimi-K2.6deepinfra/moonshotai/kimi-k2.6 | 262.144K | $0.75 | $3.5 | — | |||
| llama-3.1-8bllamagate/llama-3.1-8b | 131.072K | $0.03 | $0.05 | — | |||
| Qwen2.5 72B Instructqwen/qwen-2.5-72b-instruct | 32.768K | $0.36 | $0.4 | — | |||
| OpenAI: GPT-5.6 Sol Proopenai/gpt-5.6-sol-pro | 1.05M | $2 | $10 | — | |||
| zai-glm-4.6cerebras/zai-glm-4.6 | 128K | $2.25 | $2.75 | — | |||
| OpenAI: GPT-4o-mini (batch)openai/gpt-4o-mini:batch | 128K | $0.075 | $0.3 | — | |||
| OpenAI: GPT-4o (batch)openai/gpt-4o:batch | 128K | $1.25 | $5 | — | |||
| thinkingmachines/Inklingdeepinfra/thinkingmachines/inkling | 524.288K | $0.95 | $4.05 | — | |||
| Meta: Llama 3.2 3B Instructmeta-llama/llama-3.2-3b-instruct | 131.072K | $0.05 | $0.33 | — | |||
| OpenAI: GPT-5.6 Terra Proopenai/gpt-5.6-terra-pro | 1.05M | $2 | $12 | — | |||
| OpenAI: o3 Mini High (batch)openai/o3-mini-high:batch | 200K | $0.55 | $2.2 | — | |||
| SpaceXAI: Grok 4.3x-ai/grok-4.3 | 1M | $1.25 | $2.5 | — | |||
| meta-llama/llama-3.2-3b-instructnovita/meta-llama/llama-3.2-3b-instruct | 32.768K | $0.03 | $0.05 | — | |||
| gemini-robotics-er-1.6-previewgemini/gemini-robotics-er-1.6-preview | 131.072K | $1 | $5 | — | |||
| accounts/fireworks/models/qwen2p5-32bfireworks_ai/accounts/fireworks/models/qwen2p5-32b | 131.072K | $0.9 | $0.9 | — | |||
| Qwen: Qwen3 Next 80B A3B Instruct (free)qwen/qwen3-next-80b-a3b-instruct:free | 262.144K | Free | Free | — | |||
| accounts/fireworks/models/deepseek-r1-distill-llama-8bfireworks_ai/accounts/fireworks/models/deepseek-r1-distill-llama-8b | 131.072K | $0.2 | $0.2 | — | |||
| Mistral: Mistral Small 4 (batch)mistralai/mistral-small-2603:batch | 262.144K | $0.075 | $0.3 | — | |||
| grok-4.5-latestxai/grok-4.5-latest | 500K | $2 | $6 | — | |||
| Meta: Llama 3.2 1B Instructmeta-llama/llama-3.2-1b-instruct | 60K | $0.027 | $0.201 | — | |||
| sa-east-1/deepseek.v3.2bedrock/sa-east-1/deepseek.v3.2 | 163.84K | $0.74 | $2.22 | — | |||
| Poolside: Laguna S 2.1 (free)poolside/laguna-s-2.1:free | 262.144K | Free | Free | — |