Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

google/gemini-3.6-flash:batch 1.04858M context $0.375/M input $1.875/M output

Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....

inclusionai/ling-2.6-flash 262.144K context $0.01/M input $0.03/M output

KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

kwaipilot/kat-coder-air-v2.5 256K context $0.15/M input $0.6/M output

Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it...

qwen/qwen3-max-thinking 262.144K context $0.78/M input $3.9/M output

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...

qwen/qwen3.6-27b 262.144K context $0.3/M input $2/M output

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...

openai/gpt-5.4:batch 1.05M context $1.25/M input $7.5/M output

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

thinkingmachines/inkling:free 1.04858M context Free input Free output

Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See the launch announcement and benchmark results [here](https://www.anthropic.com/news/claude-3-haiku) #multimodal

anthropic/claude-3-haiku 200K context $0.25/M input $1.25/M output

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This...

qwen/qwen3.5-plus-20260420 1M context $0.3/M input $1.8/M output

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...

qwen/qwen3.6-max-preview 262.144K context $1.027/M input $6.162/M output

Auto Router (Beta) is a task-aware router from OpenRouter. It classifies each request, then routes it the [most popular model](/rankings#task-spend) for that task based on aggregate spend, filtered by your...

openrouter/auto-beta 2M context Input not listed Output not listed

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...

qwen/qwen3.7-flash 1M context $0.03/M input $0.13/M output

Laguna M.1 is the flagship coding agent model from [Poolside](https://poolside.ai/), optimized for complex software engineering tasks. Designed for agentic coding workflows, it supports tool calling and reasoning, with a 256K...

poolside/laguna-m.1:free 262.144K context Free input Free output

Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output. Amazon Nova Lite...

amazon/nova-lite-v1 300K context $0.06/M input $0.24/M output

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

thinkingmachines/inkling-small:batch 524.288K context $0.5/M input $1.2/M output

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...

tencent/hy3:free 262.144K context Free input Free output

MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, with 45.9 billion parameters activated per inference, and can handle a context...

minimax/minimax-01 1.00019M context $0.2/M input $1.1/M output

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...

qwen/qwen3.6-35b-a3b 262.144K context $0.1/M input $0.9/M output

OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...

openai/o3-mini-high 200K context $1.1/M input $4.4/M output

KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

kwaipilot/kat-coder-air-v2.5:free 256K context Free input Free output

KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

kwaipilot/kat-coder-pro-v2.5:free 256K context Free input Free output

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

google/gemini-2.5-pro-preview 1.04858M context $1.25/M input $10/M output

Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast execution and high efficiency at scale. It uses a “fast...

inclusionai/ling-2.6-1t 262.144K context $0.075/M input $0.625/M output

Fast-mode variant of [Opus 4.8](/anthropic/claude-opus-4.8) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 4.8. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode Note: As of September...

anthropic/claude-opus-4.8-fast 1M context $10/M input $50/M output

Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code>...

morph/morph-v3-large 262.144K context $0.9/M input $1.9/M output

No provider description is available for this model yet.

azure_ai/fw-kimi-k2.7-code 262.144K context $1.05/M input $4.4/M output

Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction following, multilingual understanding, and...

qwen/qwen3-30b-a3b-instruct-2507 262.144K context $0.09/M input $0.3/M output

Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...

x-ai/grok-build-0.1 256K context $1/M input $2/M output

UnslopNemo v4.1 is the latest addition from the creator of Rocinante, designed for adventure writing and role-play scenarios.

thedrummer/unslopnemo-12b 1.024M context $0.4/M input $0.4/M output

Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of tasks. As of December...

amazon/nova-pro-v1 300K context $0.8/M input $3.2/M output

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...

qwen/qwen3-vl-30b-a3b-instruct 262.144K context $0.15/M input $0.6/M output

No provider description is available for this model yet.

fireworks_ai/kimi-k3 1.04858M context $3/M input $15/M output

No provider description is available for this model yet.

fireworks_ai/glm-5p2-fast-us 1.04858M context $2.1/M input $6.6/M output

The Auto Router automatically selects the best model for your prompt, powered by the wisdom of the market. It routes you based on what the OpenRouter community collectively spends on...

openrouter/auto 2M context Input not listed Output not listed

No provider description is available for this model yet.

azure_ai/fw-kimi-k2.5 262.144K context $0.66/M input $3.3/M output

No provider description is available for this model yet.

fireworks_ai/glm-5p2-fast 1.04858M context $2.1/M input $6.6/M output

No provider description is available for this model yet.

fireworks_ai/kimi-k3-fast 1.04858M context $4.5/M input $22.5/M output

No provider description is available for this model yet.

azure_ai/fw-kimi-k2.6 262.144K context $1.045/M input $4.4/M output

Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary...

cohere/command-a 256K context $2.5/M input $10/M output

[GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,...

openai/gpt-5-image 400K context $10/M input $10/M output