NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
This model always redirects to the latest model in the Google Gemini Pro family.
This model always redirects to the latest model in the MoonshotAI Kimi family.
This model always redirects to the latest model in the Google Gemini Flash family.
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the...
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...
NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a hybrid Transformer-Mamba architecture, combining transformer-level accuracy with Mamba’s...
The Auto Router automatically selects the best model for your prompt, powered by the wisdom of the market. It routes you based on what the OpenRouter community collectively spends on...
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...
Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark...
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
Ox Alpha is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon software engineering, complex reasoning, and workflows that combine text with...
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...
Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...
This model always redirects to the latest model in the GLM Flash family.
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Qwen3.8 Max (0803) is the August 3, 2026 checkpoint of Qwen3.8 Max, the flagship model in Alibaba's Qwen3.8 series and the general-availability successor to the Qwen3.8 Max Preview. It is...
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| NVIDIA: Nemotron 3 Nano Omni (free)nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | 256K | Free | Free | — | |||
| Google Gemini Pro Latest~google/gemini-pro-latest | 1.04858M | $2 | $12 | — | |||
| MoonshotAI Kimi Latest~moonshotai/kimi-latest | 1.04858M | $1.616 | $8.105 | — | |||
| Google Gemini Flash Latest~google/gemini-flash-latest | 1.04858M | $0.75 | $3.75 | — | |||
| Google: Gemma 4 26B A4B (free)google/gemma-4-26b-a4b-it:free | 262.144K | Free | Free | — | |||
| Google: Gemma 4 31B (free)google/gemma-4-31b-it:free | 262.144K | Free | Free | — | |||
| Qwen: Qwen3.5-Flashqwen/qwen3.5-flash-02-23 | 1M | $0.065 | $0.26 | — | |||
| Qwen: Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | 262.144K | $0.55 | $3.5 | — | |||
| NVIDIA: Nemotron Nano 12B 2 VL (free)nvidia/nemotron-nano-12b-v2-vl:free | 128K | Free | Free | — | |||
| Auto Routeropenrouter/auto | 2M | — | — | — | |||
| Google: Gemma 4 31B (batch)google/gemma-4-31b-it:batch | 262.144K | $0.39 | $0.97 | — | |||
| MiniMax: MiniMax M3minimax/minimax-m3 | 524.288K | $0.3 | $1.2 | — | |||
| Google: Gemini 2.5 Flash Lite (batch)google/gemini-2.5-flash-lite:batch | 1.04858M | $0.05 | $0.2 | — | |||
| Google: Gemini 3.6 Flash (batch)google/gemini-3.6-flash:batch | 1.04858M | $0.375 | $1.875 | — | |||
| Qwen: Qwen3.5-9B (batch)qwen/qwen3.5-9b:batch | 262.144K | $0.17 | $0.25 | — | |||
| Qwen: Qwen3.6 Plusqwen/qwen3.6-plus | 1M | $0.325 | $1.95 | — | |||
| MiniMax: MiniMax M3 (batch)minimax/minimax-m3:batch | 524.288K | $0.3 | $1.2 | — | |||
| Google: Gemini 3.5 Flash (batch)google/gemini-3.5-flash:batch | 1.04858M | $0.75 | $4.5 | — | |||
| Google: Gemini 3 Flash Preview (batch)google/gemini-3-flash-preview:batch | 1.04858M | $0.25 | $1.5 | — | |||
| Z.ai: GLM 5.3 Flashz-ai/glm-5.3-flash | 1.04858M | $0.15 | $0.5 | — | |||
| Qwen: Qwen3.8 Max (0902)qwen/qwen3.8-max-0902 | 1M | $2 | $6 | — | |||
| Meta: Muse Spark 1.2 Contributormeta/muse-spark-1.2-contributor | 1.04858M | $0.1 | $0.2 | — | |||
| inclusionAI: Ling 3.0 Flash VL (free)inclusionai/ling-3.0-flash-vl:free | 262.144K | Free | Free | — | |||
| Ox Alphastealth/ox-alpha | 1.04858M | — | — | — | |||
| Google: Gemini 2.5 Pro (batch)google/gemini-2.5-pro:batch | 1.04858M | $0.625 | $5 | — | |||
| Meta: Muse Spark 1.3 Contributormeta/muse-spark-1.3-contributor | 1.04858M | $0.1 | $0.2 | — | |||
| Qwen: Qwen3.8 Flashqwen/qwen3.8-flash | 1M | $0.15 | $0.47 | — | |||
| MiniMax: MiniMax M3 (free)minimax/minimax-m3:free | 1.04858M | Free | Free | — | |||
| Google: Gemini 3.1 Flash Lite (batch)google/gemini-3.1-flash-lite:batch | 1.04858M | $0.125 | $0.75 | — | |||
| Z.ai: GLM Flash Latest~z-ai/glm-flash-latest | 1.04858M | $0.075 | $0.25 | — | |||
| Google: Gemini 3.7 Flash (batch)google/gemini-3.7-flash:batch | 1.04858M | $0.375 | $1.875 | — | |||
| Qwen: Qwen3.8 27Bqwen/qwen3.8-27b | 1M | $0.42 | $3 | — | |||
| MoonshotAI: Kimi K3 (batch)moonshotai/kimi-k3:batch | 1.04858M | $3 | $15 | — | |||
| inclusionAI: Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl | 131.072K | $0.06 | $0.18 | — | |||
| Z.ai: GLM 5.3 Flash (batch)z-ai/glm-5.3-flash:batch | 1.04858M | $0.075 | $0.25 | — | |||
| Google: Gemini 3.1 Pro Preview (batch)google/gemini-3.1-pro-preview:batch | 1.04858M | $1 | $6 | — | |||
| Google: Gemini 3.8 Flash (batch)google/gemini-3.8-flash:batch | 1.04858M | $0.375 | $1.875 | — | |||
| Qwen: Qwen3.8 Max (0803)qwen/qwen3.8-max | 1M | $2 | $6 | — | |||
| Google: Gemini 3.5 Flash Lite (batch)google/gemini-3.5-flash-lite:batch | 1.04858M | $0.15 | $1.25 | — | |||
| ByteDance Seed: Seed 2.1 Turbobytedance-seed/seed-2-1-turbo | 262.144K | $0.5 | $2.5 | — |