Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

anthropic/claude-opus-4.7 1M context $5/M input $25/M output

No provider description is available for this model yet.

azure/global/gpt-4o-2024-11-20 128K context $2.5/M input $10/M output

No provider description is available for this model yet.

azure/global/gpt-5.1 272K context $1.25/M input $10/M output

No provider description is available for this model yet.

azure/global/gpt-5.1-chat 128K context $1.25/M input $10/M output

ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with 47B active per token. It is trained jointly on text and image data...

baidu/ernie-4.5-vl-424b-a47b 123K context $0.42/M input $1.25/M output

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...

meta-llama/llama-4-maverick 128K context $0.2/M input $0.696/M output

GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

openai/gpt-5.6-sol-pro:batch 1.05M context $1/M input $5/M output

No provider description is available for this model yet.

azure/gpt-4-turbo-2024-04-09 128K context $10/M input $30/M output

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...

qwen/qwen3.6-27b 262.144K context $0.3/M input $2/M output

No provider description is available for this model yet.

azure/gpt-4.1 1.04758M context $2/M input $8/M output

Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...

qwen/qwen3-vl-8b-instruct 131.072K context $0.117/M input $0.455/M output

GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

openai/gpt-5.6-luna-pro 1.05M context $0.2/M input $1.2/M output

No provider description is available for this model yet.

azure/gpt-4.1-mini 1.04758M context $0.4/M input $1.6/M output

No provider description is available for this model yet.

azure/gpt-4.1-mini-2025-04-14 1.04758M context $0.4/M input $1.6/M output

No provider description is available for this model yet.

azure/gpt-4.1-nano 1.04758M context $0.1/M input $0.4/M output

No provider description is available for this model yet.

azure/gpt-4.1-nano-2025-04-14 1.04758M context $0.1/M input $0.4/M output

No provider description is available for this model yet.

azure/gpt-4.5-preview 128K context $75/M input $150/M output

No provider description is available for this model yet.

azure/gpt-4o 128K context $2.5/M input $10/M output

No provider description is available for this model yet.

azure/gpt-4o-2024-05-13 128K context $5/M input $15/M output

No provider description is available for this model yet.

together_ai/minimaxai/minimax-m3 524.288K context $0.3/M input $1.2/M output

No provider description is available for this model yet.

azure/gpt-4o-2024-08-06 128K context $2.5/M input $10/M output

No provider description is available for this model yet.

together_ai/qwen/qwen3.5-9b 262.144K context $0.17/M input $0.25/M output

No provider description is available for this model yet.

azure/gpt-4o-mini 128K context $0.165/M input $0.66/M output

No provider description is available for this model yet.

azure/gpt-4o-mini-2024-07-18 128K context $0.165/M input $0.66/M output

GPT-5 Chat is designed for advanced, natural, multimodal, and context-aware conversations for enterprise applications.

openai/gpt-5-chat 128K context $1.25/M input $10/M output

No provider description is available for this model yet.

azure/gpt-5.1-chat-2025-11-13 128K context $1.25/M input $10/M output

Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model, designed for advanced visual and textual reasoning across complex scenes, documents, and temporal sequences. It integrates enhanced multimodal alignment and...

qwen/qwen3-vl-8b-thinking 131.072K context $0.18/M input $2.1/M output

No provider description is available for this model yet.

azure/gpt-5-chat 128K context $1.25/M input $10/M output

No provider description is available for this model yet.

azure/gpt-5-chat-latest 128K context $1.25/M input $10/M output

No provider description is available for this model yet.

together_ai/google/gemma-4-31b-it 262.144K context $0.39/M input $0.97/M output

No provider description is available for this model yet.

google/medgemma-1.5-4b-it Not documented context Input not listed Output not listed

No provider description is available for this model yet.

azure/gpt-5-mini 272K context $0.25/M input $2/M output

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

google/gemini-3.5-flash-lite:batch 1.04858M context $0.15/M input $1.25/M output

No provider description is available for this model yet.

azure/gpt-5-mini-2025-08-07 272K context $0.25/M input $2/M output

No provider description is available for this model yet.

azure/gpt-5-nano-2025-08-07 272K context $0.05/M input $0.4/M output

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

anthropic/claude-haiku-4.5 200K context $1/M input $5/M output