NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a hybrid Transformer-Mamba architecture, combining transformer-level accuracy with Mamba’s...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...
This model always redirects to the latest model in the GLM Flash family.
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the...
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...
This model always redirects to the latest model in the Google Gemini Flash family.
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and...
Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in...
This model always redirects to the latest model in the MoonshotAI Kimi family.
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
Qwen3.8 Max (0803) is the August 3, 2026 checkpoint of Qwen3.8 Max, the flagship model in Alibaba's Qwen3.8 series and the general-availability successor to the Qwen3.8 Max Preview. It is...
This model always redirects to the latest model in the Google Gemini Pro family.
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This...
GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...
Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| NVIDIA: Nemotron Nano 12B 2 VL (free)nvidia/nemotron-nano-12b-v2-vl:free | 128K | Free | Free | — | |||
| inclusionAI: Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl | 131.072K | $0.06 | $0.18 | — | |||
| MoonshotAI: Kimi K3 (batch)moonshotai/kimi-k3:batch | 1.04858M | $3 | $15 | — | |||
| Qwen: Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | 262.144K | $0.55 | $3.5 | — | |||
| Z.ai: GLM Flash Latest~z-ai/glm-flash-latest | 1.04858M | $0.075 | $0.25 | — | |||
| inclusionAI: Ling 3.0 Flash VL (free)inclusionai/ling-3.0-flash-vl:free | 262.144K | Free | Free | — | |||
| Z.ai: GLM 5.3 Flashz-ai/glm-5.3-flash | 1.04858M | $0.15 | $0.5 | — | |||
| Qwen: Qwen3.5-Flashqwen/qwen3.5-flash-02-23 | 1M | $0.065 | $0.26 | — | |||
| Google: Gemini 3.7 Flash (batch)google/gemini-3.7-flash:batch | 1.04858M | $0.375 | $1.875 | — | |||
| Google: Gemma 4 31B (free)google/gemma-4-31b-it:free | 262.144K | Free | Free | — | |||
| Google: Gemini 2.5 Flash Lite (batch)google/gemini-2.5-flash-lite:batch | 1.04858M | $0.05 | $0.2 | — | |||
| MiniMax: MiniMax M3minimax/minimax-m3 | 524.288K | $0.3 | $1.2 | — | |||
| Google: Gemma 4 26B A4B (free)google/gemma-4-26b-a4b-it:free | 262.144K | Free | Free | — | |||
| Google: Gemini 3 Flash Preview (batch)google/gemini-3-flash-preview:batch | 1.04858M | $0.25 | $1.5 | — | |||
| Google: Gemma 4 31B (batch)google/gemma-4-31b-it:batch | 262.144K | $0.39 | $0.97 | — | |||
| Z.ai: GLM 5V Turboz-ai/glm-5v-turbo | 202.752K | $1.2 | $4 | — | |||
| Google: Gemini 2.5 Flash (batch)google/gemini-2.5-flash:batch | 1.04858M | $0.15 | $1.25 | — | |||
| Google Gemini Flash Latest~google/gemini-flash-latest | 1.04858M | $0.75 | $3.75 | — | |||
| Google: Gemini 2.5 Pro Preview 05-06google/gemini-2.5-pro-preview-05-06 | 1.04858M | $1.25 | $10 | — | |||
| ByteDance Seed: Seed 2.1 Turbobytedance-seed/seed-2-1-turbo | 262.144K | $0.5 | $2.5 | — | |||
| Qwen: Qwen3.6 Flashqwen/qwen3.6-flash | 1M | $0.188 | $1.125 | — | |||
| MoonshotAI Kimi Latest~moonshotai/kimi-latest | 1.04858M | $2.303 | $11.55 | — | |||
| Google: Gemini 3.5 Flash Lite (batch)google/gemini-3.5-flash-lite:batch | 1.04858M | $0.15 | $1.25 | — | |||
| Qwen: Qwen3.8 Max (0803)qwen/qwen3.8-max | 1M | $2 | $6 | — | |||
| Google Gemini Pro Latest~google/gemini-pro-latest | 1.04858M | $2 | $12 | — | |||
| NVIDIA: Nemotron 3 Nano Omni (free)nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | 256K | Free | Free | — | |||
| Google: Gemini 3.8 Flash (batch)google/gemini-3.8-flash:batch | 1.04858M | $0.375 | $1.875 | — | |||
| Qwen: Qwen3.5 Plus 2026-04-20qwen/qwen3.5-plus-20260420 | 1M | $0.3 | $1.8 | — | |||
| Z.ai: GLM 4.6Vz-ai/glm-4.6v | 131.072K | $0.3 | $0.9 | — | |||
| Qwen: Qwen3.8 27Bqwen/qwen3.8-27b | 262.144K | $0.214 | $2.55 | — | |||
| Amazon: Nova 2 Liteamazon/nova-2-lite-v1 | 1M | $0.3 | $2.5 | — | |||
| Qwen: Qwen3.5 Plus 2026-02-15qwen/qwen3.5-plus-02-15 | 1M | $0.26 | $1.56 | — | |||
| Google: Gemini 3.1 Flash Lite (batch)google/gemini-3.1-flash-lite:batch | 1.04858M | $0.125 | $0.75 | — | |||
| Qwen: Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | 262.144K | $0.26 | $2.08 | — | |||
| Z.ai: GLM 5.3 Flash (batch)z-ai/glm-5.3-flash:batch | 1.04858M | $0.075 | $0.25 | — | |||
| MiniMax: MiniMax M3 (batch)minimax/minimax-m3:batch | 524.288K | $0.3 | $1.2 | — |