No provider description is available for this model yet.
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
No provider description is available for this model yet.
Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and improved function calling. Compared to the 3.1 release, version 3.2 significantly improves accuracy on...
No provider description is available for this model yet.
No provider description is available for this model yet.
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...
No provider description is available for this model yet.
Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
No provider description is available for this model yet.
Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math....
No provider description is available for this model yet.
No provider description is available for this model yet.
OpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high. OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...
No provider description is available for this model yet.
No provider description is available for this model yet.
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| openai/gpt-5.3-codexopenrouter/openai/gpt-5.3-codex | 272K | $1.75 | $14 | — | |||
| openai/gpt-5.1openrouter/openai/gpt-5.1 | 272K | $1.25 | $10 | — | |||
| Mistral: Mistral Small 3.2 24Bmistralai/mistral-small-3.2-24b-instruct | 128K | $0.075 | $0.2 | — | |||
| openai/gpt-4o-miniopenrouter/openai/gpt-4o-mini | 128K | $0.15 | $0.6 | — | |||
| google/gemini-3.8-flashopenrouter/google/gemini-3.8-flash | 1.04858M | $0.75 | $3.75 | — | |||
| Google: Gemini 3 Flash Preview (batch)google/gemini-3-flash-preview:batch | 1.04858M | $0.25 | $1.5 | — | |||
| qwen3.7-plusqwencloud/qwen3.7-plus | 991.808K | — | — | — | |||
| Anthropic: Claude Opus 4.1anthropic/claude-opus-4.1 | 200K | $15 | $75 | — | |||
| Anthropic: Claude Fable 5.1 (batch)anthropic/claude-fable-5.1:batch | 1M | $5 | $25 | — | |||
| google/gemini-3.7-flashopenrouter/google/gemini-3.7-flash | 1.04858M | $0.75 | $3.75 | — | |||
| google/gemini-3.6-flashopenrouter/google/gemini-3.6-flash | 1.04858M | $0.75 | $3.75 | — | |||
| google/gemini-3.5-flash-liteopenrouter/google/gemini-3.5-flash-lite | 1.04858M | $0.3 | $2.5 | — | |||
| google/gemini-3.5-flashopenrouter/google/gemini-3.5-flash | 1.04858M | $1.5 | $9 | — | |||
| qwen3.5-plusqwencloud/qwen3.5-plus | 991.808K | — | — | — | |||
| google/gemini-2.5-flash-liteopenrouter/google/gemini-2.5-flash-lite | 1.04858M | $0.1 | $0.4 | — | |||
| anthropic/claude-sonnet-5openrouter/anthropic/claude-sonnet-5 | 1M | $2 | $10 | — | |||
| anthropic/claude-opus-4.8openrouter/anthropic/claude-opus-4.8 | 1M | $5 | $25 | — | |||
| anthropic/claude-fable-5.1openrouter/anthropic/claude-fable-5.1 | 1M | $10 | $50 | — | |||
| qwen3-vl-plusqwencloud/qwen3-vl-plus | 260.096K | — | — | — | |||
| anthropic/claude-fable-5openrouter/anthropic/claude-fable-5 | 1M | $10 | $50 | — | |||
| deepseek-v4-flash-vision-expfireworks_ai/deepseek-v4-flash-vision-exp | 1.04858M | $0.22 | $0.66 | — | |||
| OpenAI: GPT-5 (batch)openai/gpt-5:batch | 400K | $0.625 | $5 | — | |||
| accounts/fireworks/models/deepseek-v4-flash-vision-expfireworks_ai/accounts/fireworks/models/deepseek-v4-flash-vision-exp | 1.04858M | $0.22 | $0.66 | — | |||
| claude-mythos-5-1anthropic/claude-mythos-5-1 | 1M | $10 | $50 | — | |||
| openbmb/MiniCPM-V-4_5nebius/openbmb/minicpm-v-4_5 | 32K | $0.658 | $1.11 | — | |||
| Qwen: Qwen3.6 Plusqwen/qwen3.6-plus | 1M | $0.325 | $1.95 | — | |||
| qwen3-vl-32b-thinkingqwencloud/qwen3-vl-32b-thinking | 131.072K | $0.16 | $2.87 | — | |||
| Qwen: Qwen3 VL 235B A22B Thinkingqwen/qwen3-vl-235b-a22b-thinking | 131.072K | $0.4 | $4 | — | |||
| nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55Bdeepinfra/nvidia/nvidia-nemotron-3-ultra-550b-a55b | 262.144K | $0.5 | $2.2 | — | |||
| moonshotai/Kimi-K2.7-Codedeepinfra/moonshotai/kimi-k2.7-code | 262.144K | $0.68 | $3.4 | — | |||
| OpenAI: o4 Mini High (batch)openai/o4-mini-high:batch | 200K | $0.55 | $2.2 | — | |||
| nvidia/Nemotron-3-Nano-Omninebius/nvidia/nemotron-3-nano-omni | 262.144K | $0.06 | $0.24 | — | |||
| nvidia/Cosmos3-Super-Reasonernebius/nvidia/cosmos3-super-reasoner | 262.144K | $0.1 | $0.3 | — | |||
| moonshotai/Kimi-K3nebius/moonshotai/kimi-k3 | 1.024M | $3 | $15 | — | |||
| moonshotai/Kimi-K2.7-Codenebius/moonshotai/kimi-k2.7-code | 262.144K | $0.95 | $4 | — | |||
| MiniMax: MiniMax M3 (batch)minimax/minimax-m3:batch | 524.288K | $0.3 | $1.2 | — | |||
| qwen3-vl-32b-instructqwencloud/qwen3-vl-32b-instruct | 131.072K | $0.16 | $0.64 | — | |||
| moonshotai/Kimi-K2.6nebius/moonshotai/kimi-k2.6 | 262.144K | $0.95 | $4 | — | |||
| MiniMaxAI/MiniMax-M3nebius/minimaxai/minimax-m3 | 1.04858M | $0.3 | $1.2 | — | |||
| databricks-inklingdatabricks/databricks-inkling | 1M | $1 | $4.05 | — | |||
| OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch | 1.04758M | $0.05 | $0.2 | — | |||
| qwen3-vl-235b-a22b-thinkingqwencloud/qwen3-vl-235b-a22b-thinking | 131.072K | $0.4 | $4 | — | |||
| google/gemma-4-31B-itdeepinfra/google/gemma-4-31b-it | 262.144K | $0.13 | $0.38 | — | |||
| databricks-gpt-5-6-lunadatabricks/databricks-gpt-5-6-luna | 922K | $1 | $6 | — | |||
| databricks-gpt-5-6-terradatabricks/databricks-gpt-5-6-terra | 922K | $2.5 | $15 | — | |||
| OpenAI: GPT-5 Nano (batch)openai/gpt-5-nano:batch | 400K | $0.025 | $0.2 | — | |||
| databricks-gpt-5-6-soldatabricks/databricks-gpt-5-6-sol | 922K | $4 | $20 | — | |||
| databricks-gemini-3-5-flash-litedatabricks/databricks-gemini-3-5-flash-lite | 1.04858M | $0.375 | $3.125 | — | |||
| OpenAI: GPT-5 Mini (batch)openai/gpt-5-mini:batch | 400K | $0.125 | $1 | — | |||
| Qwen: Qwen3 VL 32B Instructqwen/qwen3-vl-32b-instruct | 131.072K | $0.104 | $0.416 | — |