1,118 models

No provider description is available for this model yet.

openai/gpt-5.4-mini-2026-03-17 272K context $0.75/M input $4.5/M output

No provider description is available for this model yet.

openai/gpt-5.4-nano-2026-03-17 272K context $0.2/M input $1.25/M output

No provider description is available for this model yet.

openai/gpt-5-2025-08-07 272K context $1.25/M input $10/M output

No provider description is available for this model yet.

openai/gpt-5-mini-2025-08-07 272K context $0.25/M input $2/M output

No provider description is available for this model yet.

openai/gpt-5-nano-2025-08-07 272K context $0.05/M input $0.4/M output

No provider description is available for this model yet.

amazon_nova/nova-lite-v1 300K context $0.06/M input $0.24/M output

No provider description is available for this model yet.

amazon_nova/nova-premier-v1 1M context $2.5/M input $12.5/M output

No provider description is available for this model yet.

amazon_nova/nova-pro-v1 300K context $0.8/M input $3.2/M output

The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenRouter. The router smartly filters for models that...

openrouter/free 200K context Free input Free output

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

google/gemma-4-26b-a4b-it:free 262.144K context Free input Free output

No provider description is available for this model yet.

azure/gpt-5-nano 272K context $0.05/M input $0.4/M output

No provider description is available for this model yet.

azure/gpt-5-mini-2025-08-07 272K context $0.25/M input $2/M output

No provider description is available for this model yet.

azure/gpt-5-mini 272K context $0.25/M input $2/M output

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

google/gemini-3.5-flash-lite:batch 1.04858M context $0.15/M input $1.25/M output

No provider description is available for this model yet.

together_ai/google/gemma-4-31b-it 262.144K context $0.39/M input $0.97/M output

No provider description is available for this model yet.

azure/gpt-5-2025-08-07 272K context $1.25/M input $10/M output

No provider description is available for this model yet.

azure_ai/grok-4.3 200K context $1.25/M input $2.5/M output

No provider description is available for this model yet.

azure/gpt-5 272K context $1.25/M input $10/M output

No provider description is available for this model yet.

azure/gpt-5.1-2025-11-13 272K context $1.25/M input $10/M output

Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the Dots 3 family and is...

dots-studio/dots-3-note-preview:free 512K context Free input Free output

No provider description is available for this model yet.

azure_ai/fw-minimax-m3 512K context $0.33/M input $1.32/M output

No provider description is available for this model yet.

together_ai/qwen/qwen3.5-9b 262.144K context $0.17/M input $0.25/M output

No provider description is available for this model yet.

together_ai/minimaxai/minimax-m3 524.288K context $0.3/M input $1.2/M output

No provider description is available for this model yet.

azure/gpt-4.1-nano-2025-04-14 1.04758M context $0.1/M input $0.4/M output

No provider description is available for this model yet.

azure/gpt-4.1-nano 1.04758M context $0.1/M input $0.4/M output

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

openai/gpt-5.6-luna:batch 1.05M context $0.1/M input $0.6/M output

No provider description is available for this model yet.

azure_ai/fw-kimi-k3 1.04858M context $3.3/M input $16.5/M output

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

thinkingmachines/inkling:free 1.04858M context Free input Free output

No provider description is available for this model yet.

azure/gpt-4.1-mini-2025-04-14 1.04758M context $0.4/M input $1.6/M output

No provider description is available for this model yet.

azure/gpt-4.1-mini 1.04758M context $0.4/M input $1.6/M output

No provider description is available for this model yet.

azure/gpt-4.1-2025-04-14 1.04758M context $2/M input $8/M output

No provider description is available for this model yet.

azure/gpt-4.1 1.04758M context $2/M input $8/M output