1,564 models

No provider description is available for this model yet.

openrouter/openai/gpt-5.3-codex 272K context $1.75/M input $14/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.1 272K context $1.25/M input $10/M output

Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and improved function calling. Compared to the 3.1 release, version 3.2 significantly improves accuracy on...

mistralai/mistral-small-3.2-24b-instruct 128K context $0.075/M input $0.2/M output

No provider description is available for this model yet.

openrouter/openai/gpt-4o-mini 128K context $0.15/M input $0.6/M output

No provider description is available for this model yet.

openrouter/google/gemini-3.8-flash 1.04858M context $0.75/M input $3.75/M output

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...

google/gemini-3-flash-preview:batch 1.04858M context $0.25/M input $1.5/M output

No provider description is available for this model yet.

qwencloud/qwen3.7-plus 991.808K context Input not listed Output not listed

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...

anthropic/claude-opus-4.1 200K context $15/M input $75/M output

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

anthropic/claude-fable-5.1:batch 1M context $5/M input $25/M output

No provider description is available for this model yet.

openrouter/google/gemini-3.7-flash 1.04858M context $0.75/M input $3.75/M output

No provider description is available for this model yet.

openrouter/google/gemini-3.6-flash 1.04858M context $0.75/M input $3.75/M output

No provider description is available for this model yet.

openrouter/google/gemini-3.5-flash 1.04858M context $1.5/M input $9/M output

No provider description is available for this model yet.

qwencloud/qwen3.5-plus 991.808K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen3-vl-plus 260.096K context Input not listed Output not listed

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...

openai/gpt-5:batch 400K context $0.625/M input $5/M output

No provider description is available for this model yet.

anthropic/claude-mythos-5-1 1M context $10/M input $50/M output

No provider description is available for this model yet.

nebius/openbmb/minicpm-v-4_5 32K context $0.658/M input $1.11/M output

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

qwen/qwen3.6-plus 1M context $0.325/M input $1.95/M output

No provider description is available for this model yet.

qwencloud/qwen3-vl-32b-thinking 131.072K context $0.16/M input $2.87/M output

Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math....

qwen/qwen3-vl-235b-a22b-thinking 131.072K context $0.4/M input $4/M output

No provider description is available for this model yet.

deepinfra/moonshotai/kimi-k2.7-code 262.144K context $0.68/M input $3.4/M output

OpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high. OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining...

openai/o4-mini-high:batch 200K context $0.55/M input $2.2/M output

No provider description is available for this model yet.

nebius/nvidia/nemotron-3-nano-omni 262.144K context $0.06/M input $0.24/M output

No provider description is available for this model yet.

nebius/moonshotai/kimi-k3 1.024M context $3/M input $15/M output

No provider description is available for this model yet.

nebius/moonshotai/kimi-k2.7-code 262.144K context $0.95/M input $4/M output

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

minimax/minimax-m3:batch 524.288K context $0.3/M input $1.2/M output

No provider description is available for this model yet.

qwencloud/qwen3-vl-32b-instruct 131.072K context $0.16/M input $0.64/M output

No provider description is available for this model yet.

nebius/moonshotai/kimi-k2.6 262.144K context $0.95/M input $4/M output

No provider description is available for this model yet.

nebius/minimaxai/minimax-m3 1.04858M context $0.3/M input $1.2/M output

No provider description is available for this model yet.

databricks/databricks-inkling 1M context $1/M input $4.05/M output

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...

openai/gpt-4.1-nano:batch 1.04758M context $0.05/M input $0.2/M output

No provider description is available for this model yet.

deepinfra/google/gemma-4-31b-it 262.144K context $0.13/M input $0.38/M output

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...

openai/gpt-5-nano:batch 400K context $0.025/M input $0.2/M output

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....

openai/gpt-5-mini:batch 400K context $0.125/M input $1/M output

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...

qwen/qwen3-vl-32b-instruct 131.072K context $0.104/M input $0.416/M output