4,213 models

Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...

meta-llama/llama-guard-4-12b 163.84K context $0.18/M input $0.18/M output

No provider description is available for this model yet.

mistral/ministral-14b-2512 262.144K context $0.2/M input $0.2/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.6-sol 1.05M context $2/M input $10/M output

No provider description is available for this model yet.

openrouter/x-ai/grok-4.5 500K context $2/M input $6/M output

No provider description is available for this model yet.

openrouter/x-ai/grok-4.3 1M context $1.25/M input $2.5/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-01-25 129.024K context $0.4/M input $1.2/M output

No provider description is available for this model yet.

openrouter/x-ai/grok-4.20 1M context $1.25/M input $2.5/M output

No provider description is available for this model yet.

qwencloud/qwen-plus 129.024K context $0.4/M input $1.2/M output

This model always redirects to the latest model in the Kimi family.

~moonshotai/kimi-latest 1.04858M context $2.1/M input $10.95/M output

This model always redirects to the latest model in the Gemini Flash family.

~google/gemini-flash-latest 1.04858M context $0.75/M input $3.75/M output

No provider description is available for this model yet.

openrouter/openai/o4-mini 200K context $1.1/M input $4.4/M output

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...

poolside/laguna-xs-2.1:free 262.144K context Free input Free output

No provider description is available for this model yet.

openrouter/openai/o3 200K context $2/M input $8/M output

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

nvidia/nemotron-3.5-content-safety:free 128K context Free input Free output

No provider description is available for this model yet.

qwencloud/qwen-max 30.72K context $1.6/M input $6.4/M output

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

mistralai/mistral-medium-3-5 262.144K context $1.5/M input $7.5/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.6-terra 922K context $2/M input $12/M output

Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and operational efficiency. It is optimized for coding agents, tool...

inclusionai/ring-2.6-1t 262.144K context $0.075/M input $0.625/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.6-luna 922K context $0.2/M input $1.2/M output

Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal capabilities. It provides state-of-the-art performance in text-based reasoning and...

mistralai/mistral-small-3.1-24b-instruct 128K context $0.351/M input $0.555/M output

No provider description is available for this model yet.

qwencloud/qwen-flash-2025-07-28 997.952K context Input not listed Output not listed

GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text...

openai/gpt-5-image-mini 400K context $2.5/M input $2/M output

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the...

qwen/qwen3.5-flash-02-23 1M context $0.065/M input $0.26/M output

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...

qwen/qwen3.5-397b-a17b 262.144K context $0.55/M input $3.5/M output

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...

minimax/minimax-m2.5 200K context $0.27/M input $1.08/M output

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

anthropic/claude-opus-4.6 1M context $5/M input $25/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.5 1.05M context $5/M input $30/M output

The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenRouter. The router smartly filters for models that...

openrouter/free 200K context Free input Free output

No provider description is available for this model yet.

openrouter/openai/gpt-5.4-nano 272K context $0.2/M input $1.25/M output

The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...

mistralai/ministral-14b-2512 262.144K context $0.2/M input $0.2/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.4-mini 272K context $0.75/M input $4.5/M output

Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing in autonomous programming via tool calling...

qwen/qwen3-coder-flash 1M context $0.195/M input $0.975/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.4 1.05M context $2.5/M input $15/M output

Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech transcription, translation and audio understanding. Input audio...

mistralai/voxtral-small-24b-2507 32.768K context $0.1/M input $0.3/M output

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

meta/muse-spark-1.3-contributor 1.04858M context $0.1/M input $0.2/M output

Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models. These models are the latest in a series of models released by IBM. They are fine-tuned for long...

ibm-granite/granite-4.0-h-micro 131K context $0.017/M input $0.112/M output

No provider description is available for this model yet.

wandb/ibm-granite/granite-4.1-8b 131.072K context $0.05/M input $0.1/M output

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

anthropic/claude-haiku-4.5 200K context $1/M input $5/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.3-codex 272K context $1.75/M input $14/M output

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual...

qwen/qwen3-next-80b-a3b-instruct 262.144K context $0.09/M input $1.1/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.1 272K context $1.25/M input $10/M output

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...

qwen/qwen3-235b-a22b-2507 262.144K context $0.087/M input $0.35/M output

No provider description is available for this model yet.

qwencloud/qwen-flash 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

openrouter/openai/gpt-4o-mini 128K context $0.15/M input $0.6/M output