No provider description is available for this model yet.
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...
Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...
No provider description is available for this model yet.
GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to...
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| claude-opus-5vertex_ai-anthropic_models/claude-opus-5 | 1M | $5 | $25 | — | |||
| claude-opus-5@defaultvertex_ai-anthropic_models/claude-opus-5@default | 1M | $5 | $25 | — | |||
| claude-opus-4-8vertex_ai-anthropic_models/claude-opus-4-8 | 1M | $5 | $25 | — | |||
| claude-opus-4-8@defaultvertex_ai-anthropic_models/claude-opus-4-8@default | 1M | $5 | $25 | — | |||
| claude-sonnet-5vertex_ai-anthropic_models/claude-sonnet-5 | 1M | $2 | $10 | — | |||
| claude-sonnet-4-6vertex_ai-anthropic_models/claude-sonnet-4-6 | 1M | $3 | $15 | — | |||
| claude-sonnet-4vertex_ai-anthropic_models/claude-sonnet-4 | 1M | $3 | $15 | — | |||
| claude-sonnet-4@20250514vertex_ai-anthropic_models/claude-sonnet-4@20250514 | 1M | $3 | $15 | — | |||
| meta/llama-4-maverick-17b-128e-instruct-maasvertex_ai-llama_models/meta/llama-4-maverick-17b-128e-instruct-maas | 1M | $0.35 | $1.15 | — | |||
| meta/llama-4-maverick-17b-16e-instruct-maasvertex_ai-llama_models/meta/llama-4-maverick-17b-16e-instruct-maas | 1M | $0.35 | $1.15 | — | |||
| meta/llama-4-scout-17b-128e-instruct-maasvertex_ai-llama_models/meta/llama-4-scout-17b-128e-instruct-maas | 10M | $0.25 | $0.7 | — | |||
| meta/llama-4-scout-17b-16e-instruct-maasvertex_ai-llama_models/meta/llama-4-scout-17b-16e-instruct-maas | 10M | $0.25 | $0.7 | — | |||
| xai/grok-4.1-fast-non-reasoningvertex_ai/xai/grok-4.1-fast-non-reasoning | 2M | $0.2 | $0.5 | — | |||
| xai/grok-4.1-fast-reasoningvertex_ai/xai/grok-4.1-fast-reasoning | 2M | $0.2 | $0.5 | — | |||
| Qwen: Qwen3.6 Flashqwen/qwen3.6-flash | 1M | $0.188 | $1.125 | — | |||
| xai/grok-4.20-reasoningvertex_ai/xai/grok-4.20-reasoning | 2M | $2 | $6 | — | |||
| grok-4-fast-reasoningxai/grok-4-fast-reasoning | 2M | $0.2 | $0.5 | — | |||
| grok-4-fast-non-reasoningxai/grok-4-fast-non-reasoning | 2M | $0.2 | $0.5 | — | |||
| grok-4-1-fastxai/grok-4-1-fast | 2M | $0.2 | $0.5 | — | |||
| grok-4-1-fast-reasoningxai/grok-4-1-fast-reasoning | 2M | $0.2 | $0.5 | — | |||
| grok-4-1-fast-reasoning-latestxai/grok-4-1-fast-reasoning-latest | 2M | $0.2 | $0.5 | — | |||
| grok-4-1-fast-non-reasoningxai/grok-4-1-fast-non-reasoning | 2M | $0.2 | $0.5 | — | |||
| grok-4-1-fast-non-reasoning-latestxai/grok-4-1-fast-non-reasoning-latest | 2M | $0.2 | $0.5 | — | |||
| grok-4.20-multi-agent-beta-0309xai/grok-4.20-multi-agent-beta-0309 | 2M | $2 | $6 | — | |||
| grok-4.20-beta-0309-reasoningxai/grok-4.20-beta-0309-reasoning | 2M | $2 | $6 | — | |||
| grok-4.20-beta-0309-non-reasoningxai/grok-4.20-beta-0309-non-reasoning | 2M | $2 | $6 | — | |||
| grok-4.3-latestxai/grok-4.3-latest | 1M | $1.25 | $2.5 | — | |||
| minimaxai/minimax-m1-80knovita/minimaxai/minimax-m1-80k | 1M | $0.55 | $2.2 | — | |||
| meta-llama/llama-4-maverick-17b-128e-instruct-fp8novita/meta-llama/llama-4-maverick-17b-128e-instruct-fp8 | 1.04858M | $0.27 | $0.85 | — | |||
| gemini-2.0-flash-lite-001gemini/gemini-2.0-flash-lite-001 | 1.04858M | $0.075 | $0.3 | — | |||
| gemini-2.5-flash-native-audio-latestgemini/gemini-2.5-flash-native-audio-latest | 1.04858M | $0.3 | $2.5 | — | |||
| gemini-2.5-flash-native-audio-preview-09-2025gemini/gemini-2.5-flash-native-audio-preview-09-2025 | 1.04858M | $0.3 | $2.5 | — | |||
| gemini-2.5-flash-native-audio-preview-12-2025gemini/gemini-2.5-flash-native-audio-preview-12-2025 | 1.04858M | $0.3 | $2.5 | — | |||
| gemini-pro-latestgemini/gemini-pro-latest | 1.04858M | $1.25 | $10 | — | |||
| claude-sonnet-5@defaultvertex_ai-anthropic_models/claude-sonnet-5@default | 1M | $2 | $10 | — | |||
| claude-sonnet-4-6@defaultvertex_ai-anthropic_models/claude-sonnet-4-6@default | 1M | $3 | $15 | — | |||
| openai-gpt-5-minisnowflake/openai-gpt-5-mini | 1M | $0.3 | $1.2 | — | |||
| openai-gpt-5-nanosnowflake/openai-gpt-5-nano | 5M | $0.15 | $0.6 | — | |||
| deepseek-v4-protencent/deepseek-v4-pro | 1M | $0.435 | $0.87 | — | |||
| deepseek-v4-flashtencent/deepseek-v4-flash | 1M | $0.14 | $0.28 | — | |||
| ps/minimax-m2.7pinstripes/ps/minimax-m2.7 | 1.00019M | $0.255 | $0.55 | — | |||
| Google: Gemini 3.7 Flash (batch)google/gemini-3.7-flash:batch | 1.04858M | $0.375 | $1.875 | — | |||
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | 1M | Free | Free | — | |||
| Anthropic: Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | 1M | $3 | $15 | — | |||
| Sakana: Fugu Maxsakana/fugu-max | 1M | $2 | $6 | — | |||
| openai/gpt-5.6-solopenrouter/openai/gpt-5.6-sol | 1.05M | $2 | $10 | — | |||
| OpenAI: GPT-5.5 (batch)openai/gpt-5.5:batch | 1.05M | $2.5 | $15 | — | |||
| NVIDIA: Nemotron 3.5 Lightning (free)nvidia/nemotron-3.5-lightning:free | 1M | Free | Free | — | |||
| Sakana: Fugu Ultra v2sakana/fugu-ultra-v2 | 1M | $5 | $30 | — | |||
| OpenAI: GPT-6 Astra (batch)openai/gpt-6-astra:batch | 1.05M | $5 | $25 | — |