This model always redirects to the latest model in the MoonshotAI Kimi family.
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
This model always redirects to the latest model in the Google Gemini Pro family.
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...
Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...
Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This...
GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...
Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...
The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of...
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...
Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired with natural language queries, and produces detailed visual understanding...
Auto Router (Beta) is a task-aware router from OpenRouter. It classifies each request, then routes it the [most popular model](/rankings#task-spend) for that task based on aggregate spend, filtered by your...
Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...
Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a 256K context window.
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features a 256k context window and can generate outputs of...
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and...
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
Qwen3.8 Max (0803) is the August 3, 2026 checkpoint of Qwen3.8 Max, the flagship model in Alibaba's Qwen3.8 series and the general-availability successor to the Qwen3.8 Max Preview. It is...
This model always redirects to the latest model in the GLM Flash family.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| MoonshotAI Kimi Latest~moonshotai/kimi-latest | 1.04858M | $2.303 | $11.55 | — | |||
| Google Gemini Pro Latest~google/gemini-pro-latest | 1.04858M | $2 | $12 | — | |||
| Google: Gemini 3 Flash Preview (batch)google/gemini-3-flash-preview:batch | 1.04858M | $0.25 | $1.5 | — | |||
| NVIDIA: Nemotron 3 Nano Omni (free)nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | 256K | Free | Free | — | |||
| Meta: Muse Spark 1.3 Contributormeta/muse-spark-1.3-contributor | 1.04858M | $0.1 | $0.2 | — | |||
| Qwen: Qwen3.5 Plus 2026-04-20qwen/qwen3.5-plus-20260420 | 1M | $0.3 | $1.8 | — | |||
| Z.ai: GLM 4.6Vz-ai/glm-4.6v | 131.072K | $0.3 | $0.9 | — | |||
| Amazon: Nova 2 Liteamazon/nova-2-lite-v1 | 1M | $0.3 | $2.5 | — | |||
| Qwen: Qwen3.6 Plusqwen/qwen3.6-plus | 1M | $0.325 | $1.95 | — | |||
| MiniMax: MiniMax M3 (free)minimax/minimax-m3:free | 1.04858M | Free | Free | — | |||
| Qwen: Qwen3.5 Plus 2026-02-15qwen/qwen3.5-plus-02-15 | 1M | $0.26 | $1.56 | — | |||
| Qwen: Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | 256K | $0.312 | $1.25 | — | |||
| Qwen: Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | 262.144K | $0.26 | $2.08 | — | |||
| Qwen: Qwen3.5-27Bqwen/qwen3.5-27b | 262.144K | $0.195 | $1.56 | — | |||
| Qwen: Qwen3.5-9Bqwen/qwen3.5-9b | 262.144K | $0.1 | $0.15 | — | |||
| Qwen: Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | 262.144K | $0.1 | $0.9 | — | |||
| Qwen: Qwen3.6 27Bqwen/qwen3.6-27b | 262.144K | $0.3 | $2 | — | |||
| Perceptron: Perceptron Mk1perceptron/perceptron-mk1 | 32.768K | $0.15 | $1.5 | — | |||
| Auto Router (Beta)openrouter/auto-beta | 2M | — | — | — | |||
| Qwen: Qwen3.7 Flashqwen/qwen3.7-flash | 1M | $0.03 | $0.13 | — | |||
| ByteDance Seed: Seed 1.6bytedance-seed/seed-1.6 | 262.144K | $0.25 | $2 | — | |||
| MiniMax: MiniMax M3minimax/minimax-m3 | 524.288K | $0.3 | $1.2 | — | |||
| ByteDance Seed: Seed 1.6 Flashbytedance-seed/seed-1.6-flash | 262.144K | $0.075 | $0.3 | — | |||
| Google: Gemma 4 31B (batch)google/gemma-4-31b-it:batch | 262.144K | $0.39 | $0.97 | — | |||
| Z.ai: GLM 5V Turboz-ai/glm-5v-turbo | 202.752K | $1.2 | $4 | — | |||
| Google: Gemini 2.5 Flash (batch)google/gemini-2.5-flash:batch | 1.04858M | $0.15 | $1.25 | — | |||
| Google: Gemini 3.1 Pro Preview (batch)google/gemini-3.1-pro-preview:batch | 1.04858M | $1 | $6 | — | |||
| Google: Gemini 2.5 Pro Preview 05-06google/gemini-2.5-pro-preview-05-06 | 1.04858M | $1.25 | $10 | — | |||
| ByteDance Seed: Seed 2.1 Turbobytedance-seed/seed-2-1-turbo | 262.144K | $0.5 | $2.5 | — | |||
| Google: Gemini 2.5 Flash Lite (batch)google/gemini-2.5-flash-lite:batch | 1.04858M | $0.05 | $0.2 | — | |||
| inclusionAI: Ling 3.0 Flash VL (free)inclusionai/ling-3.0-flash-vl:free | 262.144K | Free | Free | — | |||
| Google: Gemini 3.5 Flash Lite (batch)google/gemini-3.5-flash-lite:batch | 1.04858M | $0.15 | $1.25 | — | |||
| Qwen: Qwen3.8 Max (0803)qwen/qwen3.8-max | 1M | $2 | $6 | — | |||
| Z.ai: GLM Flash Latest~z-ai/glm-flash-latest | 1.04858M | $0.075 | $0.25 | — |