Fast variant of GPT-6 Astra for low-latency assistance and high-volume workloads.
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...
2026-09-02 upgraded snapshot of Qwen3.8 Max with stronger coding, collaborative agents, and multimodal document understanding
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through...
Claude model for demanding reasoning and long-horizon agentic work
Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...
Qwen vision-language model for visual reasoning, documents, and agent tasks
DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...
High-efficiency Gemini model for agentic workflows, coding, and multimodal reasoning
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...
xAI's frontier model for long-running agents, coding, knowledge work, and visual projects
Microsoft coding model with native vision support, optimized for fast and efficient software development
Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive...
Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context...
Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts. It is suited for Japanese instruction following,...
2.4-trillion-parameter MoE flagship for coding, professional work, multimodal understanding, and long-horizon agentic workflows
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
Preview Qwen flagship for million-token multimodal reasoning and long-horizon agentic workflows
Lightweight multimodal Qwen model for high-throughput text, image, and video tasks
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...
xAI's Grok model for chat, coding, agentic tools, and lower hallucination risk
Realtime speech-to-speech model with configurable reasoning, tool use, and robust voice-agent behavior
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...
LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic...
Video generation and editing model for fast, conversational text- and image-to-video workflows
ByteDance Seed model optimized for character-driven dialogue and consistent conversational behavior
Rolling ByteDance Seed model for rapidly updated reasoning, coding, and agent capabilities
Faster ByteDance Seed 2.1 model for multimodal reasoning and latency-sensitive agent workflows
Flagship ByteDance Seed 2.1 model for complex multimodal reasoning, coding, and agents
Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...
Multi-agent model for routing expert agents across complex analytical tasks
Restricted Claude model for advanced cybersecurity and biology research workflows
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...
Low-latency audio-to-audio model for real-time speech translation across 70+ languages
Microsoft coding model built for fast, efficient assistance in everyday developer workflows
Multimodal Qwen workhorse for long-context agents, visual inputs, and coding
Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents
Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced...
Qwen frontier model tuned for agent frameworks, coding assistants, and long tasks
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...
Compact GPT model for low-latency assistance and high-volume workloads
Qwen vision-language model for visual reasoning, documents, and agent tasks
GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| GPT-6 Astra (Fast)openai/gpt-6-astra-fast | 1.05M | $20 | $100 | 2026-09-04 | |||
| OpenAI: GPT-6 Astraopenai/gpt-6-astra | 1.05M | $10 | $50 | 2026-09-04 | |||
| Qwen3.8 Max 0902alibaba/qwen3.8-max-0902 | 1M | $1.71 | $5.14 | 2026-09-02 | |||
| Google: Gemini 3.8 Flashgoogle/gemini-3.8-flash | 1.04858M | $0.75 | $3.75 | 2026-09-02 | |||
| Meta: Muse Spark 1.3meta/muse-spark-1.3 | 1.04858M | $1.25 | $4.25 | 2026-09-02 | |||
| Claude Fable 5.1anthropic/claude-fable-5-1 | 1M | $10 | $50 | 2026-09-01 | |||
| inclusionAI: Ling 3.0 Flash Fininclusionai/ling-3.0-flash-fin | 262.144K | $0.06 | $0.18 | 2026-08-27 | |||
| Qwen3.8 Flashalibaba/qwen3.8-flash | 1M | $0.15 | $0.47 | 2026-08-26 | |||
| DeepSeek: DeepSeek V4 Flash Vision Expdeepseek/deepseek-v4-flash-vision-exp | 1.04858M | $0.22 | $0.66 | 2026-08-21 | |||
| Gemini Flash Latestgoogle/gemini-flash-latest | 1.04858M | $0.75 | $3.75 | 2026-08-13 | |||
| Google: Gemini 3.7 Flashgoogle/gemini-3.7-flash | 1.04858M | $0.75 | $3.75 | 2026-08-13 | |||
| Grok 4.6xai/grok-4.6 | 500K | $2 | $6 | 2026-08-12 | |||
| MAI-Code-1.1-Flashmicrosoft/mai-code-1.1-flash | 256K | $0.2 | $1.2 | 2026-08-11 | |||
| Upstage: Solar Pro 4upstage/solar-pro4 | 524.288K | $0.09 | $0.36 | 2026-08-06 | |||
| Meta: Muse Spark 1.2meta/muse-spark-1.2 | 1.04858M | $1.25 | $4.25 | 2026-08-05 | |||
| Sakana: Sakana Namazusakana/sakana-namazu | 262.144K | $0.95 | $4 | 2026-08-03 | |||
| Qwen3.8 Maxalibaba/qwen3.8-max | 1M | $2 | $6 | 2026-08-03 | |||
| Claude Opus 5anthropic/claude-opus-5 | 1M | $5 | $25 | 2026-07-24 | |||
| Google: Gemini 3.6 Flashgoogle/gemini-3.6-flash | 1.04858M | $0.75 | $3.75 | 2026-07-21 | |||
| Gemini Flash-Lite Latestgoogle/gemini-flash-lite-latest | 1.04858M | $0.3 | $2.5 | 2026-07-21 | |||
| Google: Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 1.04858M | $0.3 | $2.5 | 2026-07-21 | |||
| Qwen3.8 Max Previewalibaba/qwen3.8-max-preview | 1M | $2 | $6 | 2026-07-19 | |||
| Qwen3.7 Flashalibaba/qwen3.7-flash | 1M | $0.028 | $0.113 | 2026-07-15 | |||
| OpenAI: GPT-5.6 Terraopenai/gpt-5.6-terra | 1.05M | $2 | $12 | 2026-07-09 | |||
| OpenAI: GPT-5.6 Lunaopenai/gpt-5.6-luna | 1.05M | $0.2 | $1.2 | 2026-07-09 | |||
| OpenAI: GPT-5.6 Solopenai/gpt-5.6-sol | 1.05M | $2 | $10 | 2026-07-09 | |||
| Grok 4.5xai/grok-4.5 | 500K | $2 | $6 | 2026-07-08 | |||
| GPT-Realtime-2.1openai/gpt-realtime-2.1 | 128K | $4 | $24 | 2026-07-06 | |||
| Anthropic: Claude Sonnet 5anthropic/claude-sonnet-5 | 1M | $2 | $10 | 2026-06-30 | |||
| Meituan: LongCat 2.0meituan/longcat-2.0 | 1.04876M | $0.3 | $1.2 | 2026-06-30 | |||
| Gemini Omni Flash Previewgoogle/gemini-omni-flash-preview | 1.04858M | $1.5 | $17.5 | 2026-06-30 | |||
| Seed Characterbytedance-seed/seed-character | 256K | $0.119 | $0.297 | 2026-06-23 | |||
| Seed Evolvingbytedance-seed/seed-evolving | 256K | $0.884 | $4.42 | 2026-06-23 | |||
| Seed 2.1 Turbobytedance-seed/seed-2.1-turbo | 256K | $0.354 | $1.77 | 2026-06-23 | |||
| Seed 2.1 Probytedance-seed/seed-2.1-pro | 256K | $0.707 | $3.536 | 2026-06-23 | |||
| Sakana: Fugu Ultrasakana/fugu-ultra | 1M | $5 | $30 | 2026-06-15 | |||
| Fugusakana/fugu | 1M | — | — | 2026-06-15 | |||
| Claude Mythos 5anthropic/claude-mythos-5 | 1M | $10 | $50 | 2026-06-09 | |||
| Anthropic: Claude Fable 5anthropic/claude-fable-5 | 1M | $10 | $50 | 2026-06-09 | |||
| Gemini 3.5 Live Translate Previewgoogle/gemini-3.5-live-translate-preview | 131.072K | $3.5 | $21 | 2026-06-09 | |||
| MAI-Code-1-Flashmicrosoft/mai-code-1-flash | 256K | $0.75 | $4.5 | 2026-06-02 | |||
| Qwen3.7 Plusalibaba/qwen3.7-plus | 1M | $0.5 | $3 | 2026-06-02 | |||
| Claude Opus 4.8anthropic/claude-opus-4-8 | 1M | $5 | $25 | 2026-05-28 | |||
| Google: Nano Banana 2 (Gemini 3.1 Flash Image)google/gemini-3.1-flash-image | 131.072K | $0.5 | $3 | 2026-05-28 | |||
| Qwen3.7 Maxalibaba/qwen3.7-max | 1M | $2.5 | $7.5 | 2026-05-21 | |||
| Google: Gemini 3.5 Flashgoogle/gemini-3.5-flash | 1.04858M | $1.5 | $9 | 2026-05-19 | |||
| Google: Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-lite | 1.04858M | $0.25 | $1.5 | 2026-05-07 | |||
| GPT-5.5 Instantopenai/gpt-5.5-instant | 400K | $5 | $30 | 2026-05-05 | |||
| Qwen3.6 Flashalibaba/qwen3.6-flash | 1M | $0.188 | $1.125 | 2026-04-27 | |||
| OpenAI: GPT-5.5openai/gpt-5.5 | 1.05M | $5 | $30 | 2026-04-23 |