Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
Google's proven reasoning model for coding, math, and multimodal analysis
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...
GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...
Long-lived GPT workhorse for coding, instruction following, and production apps
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
Affordable GPT-4.1 lane for fast coding help and structured extraction
Fast Gemini workhorse for multimodal apps where latency and price matter
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...
DeepSeek chat model for instruction following, coding, and analysis
Low-latency Gemini model for high-volume multimodal and agent workloads
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flashdeepseek/deepseek-v4-flash | 1219.0 | 1M | $0.15 | $0.6 | 2026-04-24 | |||
| SpaceXAI: Grok 4.20x-ai/grok-4.20 | 1216.0 | 2M | $1.25 | $2.5 | — | |||
| SpaceXAI: Grok 4.3x-ai/grok-4.3 | 1197.0 | 1M | $1.25 | $2.5 | — | |||
| SpaceXAI: Grok 4.3 (batch)x-ai/grok-4.3:batch | 1197.0 | 1M | $1 | $2 | — | |||
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | 1189.0 | 1M | $3 | $15 | — | |||
| Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | 1189.0 | 1.04858M | $0.5 | $3 | 2025-12-17 | |||
| Google: Gemini 3 Flash Preview (batch)google/gemini-3-flash-preview:batch | 1189.0 | 1.04858M | $0.25 | $1.5 | — | |||
| Anthropic: Claude Sonnet 4.5 (batch)anthropic/claude-sonnet-4.5:batch | 1189.0 | 1M | $1.5 | $7.5 | — | |||
| Inklingthinkingmachines/inkling | 1176.0 | 1.04858M | $1.87 | $4.68 | 2026-07-15 | |||
| Thinking Machines: Inkling (free)thinkingmachines/inkling:free | 1176.0 | 1.04858M | Free | Free | — | |||
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | 1153.0 | 1M | Free | Free | — | |||
| Nemotron 3 Ultra 550B A55Bnvidia/nemotron-3-ultra-550b-a55b | 1153.0 | 1M | $0.5 | $2.5 | 2026-06-04 | |||
| Google: Gemini 2.5 Pro (batch)google/gemini-2.5-pro:batch | 1134.0 | 1.04858M | $0.625 | $5 | — | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 1134.0 | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| Qwen: Qwen3.5 Plus 2026-02-15qwen/qwen3.5-plus-02-15 | 1128.0 | 1M | $0.26 | $1.56 | — | |||
| OpenAI: GPT-4.1 (batch)openai/gpt-4.1:batch | 1102.0 | 1.04758M | $1 | $4 | — | |||
| GPT-4.1openai/gpt-4.1 | 1102.0 | 1.04758M | $2 | $8 | 2025-04-14 | |||
| OpenAI: GPT-4.1 Mini (batch)openai/gpt-4.1-mini:batch | 1093.0 | 1.04758M | $0.2 | $0.8 | — | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1093.0 | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1088.0 | 1.04858M | $0.3 | $2.5 | 2025-06-17 | |||
| Google: Gemini 2.5 Flash (batch)google/gemini-2.5-flash:batch | 1088.0 | 1.04858M | $0.15 | $1.25 | — | |||
| DeepSeek Chatdeepseek/deepseek-chat | 1076.0 | 1M | $0.147 | $0.295 | 2025-12-01 | |||
| Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 1052.0 | 1.04858M | $0.25 | $1.5 | 2026-03-03 | |||
| OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch | 993.0 | 1.04758M | $0.05 | $0.2 | — | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 993.0 | 1.04758M | $0.1 | $0.4 | 2025-04-14 |