131 models
Reasoning Tools JSON

Advanced Gemini model for complex reasoning, coding, and multimodal analysis

google/gemini-3.1-pro-preview-customtools 2026-02-19 1.04858M context $2/M input $12/M output
11 providers
Reasoning Tools JSON

Reasoning-first Gemini preview for agentic coding and complex problem solving

google/gemini-3.1-pro-preview 2026-02-19 1.04858M context $2/M input $12/M output
31 providers

Qwen vision-language model for visual reasoning, documents, and agent tasks

alibaba/qwen3.5-plus 2026-02-16 1M context $0.4/M input $2.4/M output
13 providers
Reasoning Tools JSON Open weights

Large open Qwen multimodal MoE for visual agents and long technical tasks

alibaba/qwen3.5-397b-a17b 2026-02-15 262.144K context $0.6/M input $3.6/M output
18 providers
Reasoning Tools JSON

Flagship ByteDance Seed 2.0 model for complex multimodal reasoning and long-horizon agent workflows

bytedance-seed/seed-2.0-pro 2026-02-14 256K context $0.475/M input $2.375/M output
6 providers
JSON

Cost-efficient ByteDance Seed 2.0 model for production chat, analysis, and structured generation

bytedance-seed/seed-2.0-lite 2026-02-14 256K context $0.089/M input $0.534/M output
8 providers
Reasoning Tools JSON

ByteDance Seed coding model for multimodal software engineering and long-running agents

bytedance-seed/seed-2.0-code 2026-02-14 262.144K context $0.4/M input $2.4/M output
10 providers
Reasoning Tools JSON

Lightweight ByteDance Seed 2.0 model for low-latency multimodal reasoning and high-volume tasks

bytedance-seed/seed-2.0-mini 2026-02-14 256K context $0.03/M input $0.297/M output
9 providers
Tools Open weights

Earlier Kimi frontier model for long-context agents, coding, and multimodal work

moonshotai/kimi-k2.5 2026-01 262.144K context $0.3/M input $1.9/M output
46 providers
Reasoning Tools JSON

New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs

google/gemini-3-flash-preview 2025-12-17 1.04858M context $0.5/M input $3/M output
25 providers
Reasoning Tools Open weights

Lightweight GLM vision model for visual reasoning, documents, and multimodal agents

zhipuai/glm-4.6v-flash 2025-12-08 128K context $0.3/M input $0.9/M output
2 providers
Reasoning Tools Open weights

GLM vision model for visual reasoning, documents, and multimodal agents

zhipuai/glm-4.6v 2025-12-08 128K context $0.3/M input $0.9/M output
9 providers
Reasoning

Multimodal reasoning model for visual analysis, planning, and tool use

amazon/nova-2-lite 2025-12-02 1M context $0.3/M input $2.5/M output
4 providers
Reasoning Tools JSON

Preview Gemini flagship for complex reasoning, coding, and rich multimodal prompts

google/gemini-3-pro-preview 2025-11-18 1.04858M context $0.57/M input $3.43/M output
10 providers
Reasoning Tools Open weights

Nemotron multimodal model for visual reasoning and agentic AI workflows

nvidia/nemotron-nano-12b-v2-vl 2025-10-28 128K context $0.2/M input $0.6/M output
3 providers
Reasoning Tools JSON

Fast Gemini workhorse for multimodal apps where latency and price matter

google/gemini-2.5-flash 2025-06-17 1.04858M context $0.3/M input $2.5/M output
30 providers
Reasoning Tools JSON

Google's proven reasoning model for coding, math, and multimodal analysis

google/gemini-2.5-pro 2025-06-17 1.04858M context $1.25/M input $10/M output
29 providers
Reasoning Tools JSON

Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents

google/gemini-2.5-flash-lite 2025-06-17 1.04858M context $0.1/M input $0.4/M output
18 providers

Multimodal model for complex analysis, long-context understanding, tool use, and model distillation

amazon/nova-premier 2025-04-30 1M context Input not listed Output not listed
Tools

Earlier Gemini Flash workhorse for responsive multimodal apps and tool use

google/gemini-2.0-flash 2024-12-11 1.04858M context $0.1/M input $0.42/M output
2 providers
Tools

Low-latency Gemini model for high-volume multimodal and agent workloads

google/gemini-2.0-flash-lite 2024-12-11 1.04858M context $0.052/M input $0.21/M output
2 providers
Tools

Flagship model for demanding analysis, coding, and production agent workflows

amazon/nova-pro 2024-12-03 300K context $0.8/M input $3.2/M output
3 providers
Tools

Efficient model for low-latency assistance, extraction, and routine automation

amazon/nova-lite 2024-12-03 300K context $0.06/M input $0.24/M output
3 providers

This model always redirects to the latest model in the Google Gemini Pro family.

~google/gemini-pro-latest 1.04858M context $2/M input $12/M output

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

minimax/minimax-m3:batch 524.288K context $0.3/M input $1.2/M output

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

z-ai/glm-5.3-flash 1.04858M context $0.15/M input $0.5/M output

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

google/gemini-3.6-flash:batch 1.04858M context $0.375/M input $1.875/M output

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

qwen/qwen3.8-flash 1M context $0.15/M input $0.47/M output

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

meta/muse-spark-1.3-contributor 1.04858M context $0.1/M input $0.2/M output

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

google/gemini-3.1-pro-preview:batch 1.04858M context $1/M input $6/M output

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

minimax/minimax-m3:free 1.04858M context Free input Free output

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in...

qwen/qwen3.6-flash 1M context $0.188/M input $1.125/M output

Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features a 256k context window and can generate outputs of...

bytedance-seed/seed-1.6-flash 262.144K context $0.075/M input $0.3/M output

Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a 256K context window.

bytedance-seed/seed-1.6 262.144K context $0.25/M input $2/M output

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...

qwen/qwen3.7-flash 1M context $0.03/M input $0.13/M output

Auto Router (Beta) is a task-aware router from OpenRouter. It classifies each request, then routes it the [most popular model](/rankings#task-spend) for that task based on aggregate spend, filtered by your...

openrouter/auto-beta 2M context Input not listed Output not listed

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...

qwen/qwen3.6-35b-a3b 262.144K context $0.1/M input $0.9/M output

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...

qwen/qwen3.6-27b 262.144K context $0.3/M input $2/M output

Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...

qwen/qwen3.5-9b 262.144K context $0.1/M input $0.15/M output

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...

qwen/qwen3.5-35b-a3b 256K context $0.312/M input $1.25/M output

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of...

qwen/qwen3.5-27b 262.144K context $0.195/M input $1.56/M output

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...

qwen/qwen3.5-122b-a10b 262.144K context $0.26/M input $2.08/M output

The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...

qwen/qwen3.5-plus-02-15 1M context $0.26/M input $1.56/M output

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...

z-ai/glm-4.6v 131.072K context $0.3/M input $0.9/M output

Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...

amazon/nova-2-lite-v1 1M context $0.3/M input $2.5/M output

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This...

qwen/qwen3.5-plus-20260420 1M context $0.3/M input $1.8/M output

NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free 256K context Free input Free output

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

google/gemini-3.7-flash:batch 1.04858M context $0.375/M input $1.875/M output

This model always redirects to the latest model in the MoonshotAI Kimi family.

~moonshotai/kimi-latest 1.04858M context $2.34/M input $11.7/M output

This model always redirects to the latest model in the Google Gemini Flash family.

~google/gemini-flash-latest 1.04858M context $0.75/M input $3.75/M output