Qwen vision-language model for visual reasoning, documents, and agent tasks
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Advanced Gemini model for complex reasoning, coding, and multimodal analysis
Reasoning-first Gemini preview for agentic coding and complex problem solving
Qwen vision-language model for visual reasoning, documents, and agent tasks
Large open Qwen multimodal MoE for visual agents and long technical tasks
Cost-efficient ByteDance Seed 2.0 model for production chat, analysis, and structured generation
Flagship ByteDance Seed 2.0 model for complex multimodal reasoning and long-horizon agent workflows
ByteDance Seed coding model for multimodal software engineering and long-running agents
Lightweight ByteDance Seed 2.0 model for low-latency multimodal reasoning and high-volume tasks
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs
Lightweight GLM vision model for visual reasoning, documents, and multimodal agents
GLM vision model for visual reasoning, documents, and multimodal agents
Multimodal reasoning model for visual analysis, planning, and tool use
Preview Gemini flagship for complex reasoning, coding, and rich multimodal prompts
Nemotron multimodal model for visual reasoning and agentic AI workflows
Video model for prompt-guided generation, editing, and motion workflows
Video model for prompt-guided generation, editing, and motion workflows
GLM vision model for visual reasoning, documents, and multimodal agents
Fast Gemini workhorse for multimodal apps where latency and price matter
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents
Google's proven reasoning model for coding, math, and multimodal analysis
Multimodal model for complex analysis, long-context understanding, tool use, and model distillation
Qwen omni model for text, vision, audio, and multimodal agent tasks
Low-latency Gemini model for high-volume multimodal and agent workloads
Earlier Gemini Flash workhorse for responsive multimodal apps and tool use
Efficient model for low-latency assistance, extraction, and routine automation
Flagship model for demanding analysis, coding, and production agent workflows
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...
This model always redirects to the latest model in the MoonshotAI Kimi family.
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
This model always redirects to the latest model in the Google Gemini Flash family.
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This...
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...
GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...
Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...
Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
This model always redirects to the latest model in the GLM Flash family.
Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Qwen3.5 27Balibaba/qwen3.5-27b | 262.144K | $0.3 | $2.4 | 2026-02-23 | |||
| Qwen3.5 122B-A10Balibaba/qwen3.5-122b-a10b | 262.144K | $0.4 | $3.2 | 2026-02-23 | |||
| Qwen3.5 Flashalibaba/qwen3.5-flash | 1M | $0.029 | $0.287 | 2026-02-23 | |||
| Gemini 3.1 Pro Preview Custom Toolsgoogle/gemini-3.1-pro-preview-customtools | 1.04858M | $2 | $12 | 2026-02-19 | |||
| Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | 1.04858M | $2 | $12 | 2026-02-19 | |||
| Qwen3.5 Plusalibaba/qwen3.5-plus | 1M | $0.4 | $2.4 | 2026-02-16 | |||
| Qwen3.5 397B-A17Balibaba/qwen3.5-397b-a17b | 262.144K | $0.6 | $3.6 | 2026-02-15 | |||
| Seed 2.0 Litebytedance-seed/seed-2.0-lite | 256K | $0.089 | $0.534 | 2026-02-14 | |||
| Seed 2.0 Probytedance-seed/seed-2.0-pro | 256K | $0.475 | $2.375 | 2026-02-14 | |||
| Seed 2.0 Codebytedance-seed/seed-2.0-code | 262.144K | $0.4 | $2.4 | 2026-02-14 | |||
| Seed 2.0 Minibytedance-seed/seed-2.0-mini | 256K | $0.03 | $0.297 | 2026-02-14 | |||
| Kimi K2.5moonshotai/kimi-k2.5 | 262.144K | $0.3 | $1.9 | 2026-01 | |||
| Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | 1.04858M | $0.5 | $3 | 2025-12-17 | |||
| GLM-4.6V-Flashzhipuai/glm-4.6v-flash | 128K | $0.3 | $0.9 | 2025-12-08 | |||
| GLM-4.6Vzhipuai/glm-4.6v | 128K | $0.3 | $0.9 | 2025-12-08 | |||
| Nova 2 Liteamazon/nova-2-lite | 1M | $0.3 | $2.5 | 2025-12-02 | |||
| Gemini 3 Pro Previewgoogle/gemini-3-pro-preview | 1.04858M | $0.57 | $3.43 | 2025-11-18 | |||
| Nemotron Nano 12B v2 VLnvidia/nemotron-nano-12b-v2-vl | 128K | $0.2 | $0.6 | 2025-10-28 | |||
| Veo 3.1 Previewgoogle/veo-3.1-generate-preview | 1.024K | — | — | 2025-10-15 | |||
| Veo 3.1 Fast Previewgoogle/veo-3.1-fast-generate-preview | 1.024K | — | — | 2025-10-15 | |||
| GLM-4.5Vzhipuai/glm-4.5v | 64K | $0.6 | $1.8 | 2025-08-11 | |||
| Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1.04858M | $0.3 | $2.5 | 2025-06-17 | |||
| Gemini 2.5 Flash-Litegoogle/gemini-2.5-flash-lite | 1.04858M | $0.1 | $0.4 | 2025-06-17 | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| Nova Premieramazon/nova-premier | 1M | — | — | 2025-04-30 | |||
| Qwen-Omni Turboalibaba/qwen-omni-turbo | 32.768K | $0.07 | $0.27 | 2025-01-19 | |||
| Gemini 2.0 Flash-Litegoogle/gemini-2.0-flash-lite | 1.04858M | $0.052 | $0.21 | 2024-12-11 | |||
| Gemini 2.0 Flashgoogle/gemini-2.0-flash | 1.04858M | $0.1 | $0.42 | 2024-12-11 | |||
| Nova Liteamazon/nova-lite | 300K | $0.06 | $0.24 | 2024-12-03 | |||
| Nova Proamazon/nova-pro | 300K | $0.8 | $3.2 | 2024-12-03 | |||
| Google: Gemini 3.7 Flash (batch)google/gemini-3.7-flash:batch | 1.04858M | $0.375 | $1.875 | — | |||
| MoonshotAI Kimi Latest~moonshotai/kimi-latest | 1.04858M | $2.34 | $11.7 | — | |||
| MiniMax: MiniMax M3 (batch)minimax/minimax-m3:batch | 524.288K | $0.3 | $1.2 | — | |||
| Google Gemini Flash Latest~google/gemini-flash-latest | 1.04858M | $0.75 | $3.75 | — | |||
| Qwen: Qwen3.5-9B (batch)qwen/qwen3.5-9b:batch | 262.144K | $0.17 | $0.25 | — | |||
| Z.ai: GLM 5.3 Flashz-ai/glm-5.3-flash | 1.04858M | $0.15 | $0.5 | — | |||
| Qwen: Qwen3.5 Plus 2026-04-20qwen/qwen3.5-plus-20260420 | 1M | $0.3 | $1.8 | — | |||
| Google: Gemini 3.6 Flash (batch)google/gemini-3.6-flash:batch | 1.04858M | $0.375 | $1.875 | — | |||
| Z.ai: GLM 4.6Vz-ai/glm-4.6v | 131.072K | $0.3 | $0.9 | — | |||
| Amazon: Nova 2 Liteamazon/nova-2-lite-v1 | 1M | $0.3 | $2.5 | — | |||
| Qwen: Qwen3.8 Flashqwen/qwen3.8-flash | 1M | $0.15 | $0.47 | — | |||
| Google: Gemini 3.5 Flash (batch)google/gemini-3.5-flash:batch | 1.04858M | $0.75 | $4.5 | — | |||
| Z.ai: GLM Flash Latest~z-ai/glm-flash-latest | 1.04858M | $0.075 | $0.25 | — | |||
| Meta: Muse Spark 1.3 Contributormeta/muse-spark-1.3-contributor | 1.04858M | $0.1 | $0.2 | — | |||
| Qwen: Qwen3.5 Plus 2026-02-15qwen/qwen3.5-plus-02-15 | 1M | $0.26 | $1.56 | — | |||
| Qwen: Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | 256K | $0.312 | $1.25 | — | |||
| Google: Gemma 4 26B A4B (free)google/gemma-4-26b-a4b-it:free | 262.144K | Free | Free | — | |||
| Qwen: Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | 262.144K | $0.26 | $2.08 | — | |||
| inclusionAI: Ling 3.0 Flash VL (free)inclusionai/ling-3.0-flash-vl:free | 262.144K | Free | Free | — | |||
| Qwen: Qwen3.5-27Bqwen/qwen3.5-27b | 262.144K | $0.195 | $1.56 | — |