1,808 models

Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...

amazon/nova-2-lite-v1 1M context $0.3/M input $2.5/M output

No provider description is available for this model yet.

azure/eu/gpt-5-mini-2025-08-07 272K context $0.275/M input $2.2/M output

No provider description is available for this model yet.

azure/eu/gpt-5.1 272K context $1.38/M input $11/M output

No provider description is available for this model yet.

azure/eu/gpt-5-nano-2025-08-07 272K context $0.055/M input $0.44/M output

No provider description is available for this model yet.

azure/eu/o1-2024-12-17 200K context $16.5/M input $66/M output

Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...

inclusionai/ling-3.0-flash-fin:free 262.144K context Free input Free output

No provider description is available for this model yet.

azure/us-gov/o3-mini 200K context $1.513/M input $6.05/M output

No provider description is available for this model yet.

openai/gpt-5.6-cyber 400K context $12.5/M input $75/M output

The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenRouter. The router smartly filters for models that...

openrouter/free 200K context Free input Free output

No provider description is available for this model yet.

gmi/zai-org/glm-4.7-fp8 202.752K context $0.4/M input $2/M output

Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total parameters and only 3B activated per...

qwen/qwen3-coder-next 262.144K context $0.12/M input $0.8/M output

The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user request. In contrast to RAG, relace-search performs agentic...

relace/relace-search 256K context $1/M input $3/M output

GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

openai/gpt-5.6-terra-pro 1.05M context $2/M input $12/M output

Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination.

qwen/qwen-plus 1M context $0.26/M input $0.78/M output

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...

deepseek/deepseek-v4-flash-vision-exp:batch 1.04858M context $0.11/M input $0.33/M output

No provider description is available for this model yet.

azure_ai/gpt-6-astra 922K context $10/M input $50/M output

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

anthropic/claude-opus-4.6 1M context $5/M input $25/M output

GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

openai/gpt-6-astra-pro 1.05M context $10/M input $50/M output

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...

minimax/minimax-m2.5 200K context $0.27/M input $1.08/M output

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...

z-ai/glm-4.7 202.752K context $0.4/M input $1.75/M output

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...

openai/gpt-5.1:batch 400K context $0.625/M input $5/M output

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost combination.

qwen/qwen-plus-2025-07-28:thinking 1M context $0.26/M input $0.78/M output

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

openai/gpt-5.6-luna:batch 1.05M context $0.1/M input $0.6/M output

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...

qwen/qwen3.5-397b-a17b 262.144K context $0.55/M input $3.5/M output

Ox Alpha is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon software engineering, complex reasoning, and workflows that combine text with...

stealth/ox-alpha 1.04858M context Input not listed Output not listed

No provider description is available for this model yet.

gmi/moonshotai/kimi-k2-thinking 262.144K context $0.8/M input $1.2/M output

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the...

qwen/qwen3.5-flash-02-23 1M context $0.065/M input $0.26/M output

MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world...

minimax/minimax-m2.1 204.8K context $0.3/M input $1.2/M output

Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fetch enabled, then a...

openrouter/fusion 1M context Input not listed Output not listed

No provider description is available for this model yet.

gmi/google/gemini-3-flash-preview 1.04858M context $0.5/M input $3/M output

No provider description is available for this model yet.

gmi/google/gemini-3-pro-preview 1.04858M context $2/M input $12/M output

GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text...

openai/gpt-5-image-mini 400K context $2.5/M input $2/M output

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for...

openai/gpt-5.5-pro:batch 1.05M context $15/M input $90/M output