1,808 models

Fast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

anthropic/claude-opus-4.7-fast 1M context $30/M input $150/M output

No provider description is available for this model yet.

snowflake/openai-gpt-4.1 300K context $2/M input $8/M output

Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...

x-ai/grok-build-0.1 256K context $1/M input $2/M output

Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast execution and high efficiency at scale. It uses a “fast...

inclusionai/ling-2.6-1t 262.144K context $0.075/M input $0.625/M output

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...

qwen/qwen3.6-35b-a3b 262.144K context $0.1/M input $0.9/M output

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...

qwen/qwen3.6-max-preview 262.144K context $1.027/M input $6.162/M output

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...

qwen/qwen3.6-27b 262.144K context $0.3/M input $2/M output

Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....

inclusionai/ling-2.6-flash 262.144K context $0.01/M input $0.03/M output

The Pareto Router maintains a tiered shortlist of strong coding models, ranked by [Artificial Analysis](https://artificialanalysis.ai/) coding percentiles. Set min_coding_score between 0 and 1 on the [pareto-router plugin](https://openrouter.ai/docs/guides/routing/routers/pareto-router#the-min_coding_score-parameter) to control how...

openrouter/pareto-code 2M context Input not listed Output not listed

No provider description is available for this model yet.

openrouter/x-ai/grok-4.3 1M context $1.25/M input $2.5/M output

[GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and...

openai/gpt-5.4-image-2 272K context $8/M input $15/M output

KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration. It builds on the agentic coding strengths of earlier versions,...

kwaipilot/kat-coder-pro-v2 262.144K context $0.3/M input $1.2/M output

Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...

qwen/qwen3.5-9b 262.144K context $0.1/M input $0.15/M output

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...

qwen/qwen3.5-35b-a3b 256K context $0.312/M input $1.25/M output

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of...

qwen/qwen3.5-27b 262.144K context $0.195/M input $1.56/M output

No provider description is available for this model yet.

deepinfra/bytedance/seed-2.0-pro 256K context $0.5/M input $3/M output

No provider description is available for this model yet.

openrouter/x-ai/grok-4.20 1M context $1.25/M input $2.5/M output

No provider description is available for this model yet.

openrouter/openai/o4-mini 200K context $1.1/M input $4.4/M output

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...

z-ai/glm-4.7 202.752K context $0.4/M input $1.75/M output

Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...

amazon/nova-2-lite-v1 1M context $0.3/M input $2.5/M output

The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user request. In contrast to RAG, relace-search performs agentic...

relace/relace-search 256K context $1/M input $3/M output

Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling custom models.

amazon/nova-premier-v1 1M context $2.5/M input $12.5/M output

Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed for deeper reasoning and analysis. Pricing is based...

perplexity/sonar-pro-search 200K context $3/M input $15/M output

Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction following, multilingual understanding, and...

qwen/qwen3-30b-a3b-instruct-2507 262.144K context $0.09/M input $0.3/M output

[GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,...

openai/gpt-5-image 400K context $10/M input $10/M output

Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code>...

morph/morph-v3-large 262.144K context $0.9/M input $1.9/M output

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

anthropic/claude-fable-5:batch 1M context $5/M input $25/M output

OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...

openai/o3-mini-high 200K context $1.1/M input $4.4/M output

No provider description is available for this model yet.

openrouter/openai/o3 200K context $2/M input $8/M output

Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output. Amazon Nova Lite...

amazon/nova-lite-v1 300K context $0.06/M input $0.24/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.6-terra 922K context $2/M input $12/M output

Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See the launch announcement and benchmark results [here](https://www.anthropic.com/news/claude-3-haiku) #multimodal

anthropic/claude-3-haiku 200K context $0.25/M input $1.25/M output

Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it...

qwen/qwen3-max-thinking 262.144K context $0.78/M input $3.9/M output

No provider description is available for this model yet.

deepinfra/tencent/hy3 262.144K context $0.14/M input $0.58/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.6-luna 922K context $0.2/M input $1.2/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.5 1.05M context $5/M input $30/M output

OpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high. OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining...

openai/o4-mini-high:batch 200K context $0.55/M input $2.2/M output

No provider description is available for this model yet.

snowflake/claude-3-7-sonnet 200K context $3/M input $15/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.4-nano 272K context $0.2/M input $1.25/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.4-mini 272K context $0.75/M input $4.5/M output

No provider description is available for this model yet.

deepinfra/bytedance/seed-1.8 256K context $0.25/M input $2/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.4 1.05M context $2.5/M input $15/M output

This model always redirects to the latest model in the Gemini Pro family.

~google/gemini-pro-latest 1.04858M context $2/M input $12/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.3-codex 272K context $1.75/M input $14/M output