Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Open MoE flagship with million-token context for coding and long agent runs
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...
Reasoning-first Gemini preview for agentic coding and complex problem solving
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
Multimodal MoE reasoning model (276B total, 12B active) for text, image, and audio
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Open MiMo model for multimodal coding agents and long-context automation
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Meituan LongCat-2.0, a reasoning model with tool calling and a 1M-token context window
Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...
Google's proven reasoning model for coding, math, and multimodal analysis
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
Low-latency Gemini model for high-volume multimodal and agent workloads
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
Affordable GPT-4.1 lane for fast coding help and structured extraction
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| Google: Gemini 3.5 Flash (batch)google/gemini-3.5-flash:batch | 33.0 | 1.04858M | $0.75 | $4.5 | — | |||
| Gemini 3.5 Flashgoogle/gemini-3.5-flash | 33.0 | 1.04858M | $1.5 | $9 | 2026-05-19 | |||
| DeepSeek V4 Prodeepseek/deepseek-v4-pro | 30.9 | 1M | $0.435 | $0.87 | 2026-04-24 | |||
| Anthropic: Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | 30.5 | 1M | $3 | $15 | — | |||
| Anthropic: Claude Sonnet 4.6 (batch)anthropic/claude-sonnet-4.6:batch | 30.5 | 1M | $1.5 | $7.5 | — | |||
| Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | 30.4 | 1.04858M | $2 | $12 | 2026-02-19 | |||
| Google: Gemini 3.1 Pro Preview (batch)google/gemini-3.1-pro-preview:batch | 30.4 | 1.04858M | $1 | $6 | — | |||
| Qwen: Qwen3.7 Maxqwen/qwen3.7-max | 29.9 | 1M | $1.475 | $4.425 | — | |||
| MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 26.4 | 1.04858M | $0.435 | $0.87 | 2026-04-22 | |||
| Inkling Smallthinkingmachines/inkling-small | 26.1 | 1.04858M | $0.45 | $1.2 | 2026-07-30 | |||
| Thinking Machines: Inkling Small (free)thinkingmachines/inkling-small:free | 26.1 | 1.04858M | Free | Free | — | |||
| Qwen: Qwen3.7 Plusqwen/qwen3.7-plus | 25.8 | 1M | $0.32 | $1.28 | — | |||
| Inklingthinkingmachines/inkling | 25.5 | 1.04858M | $1.87 | $4.68 | 2026-07-15 | |||
| Thinking Machines: Inkling (free)thinkingmachines/inkling:free | 25.5 | 1.04858M | Free | Free | — | |||
| SpaceXAI: Grok 4.3 (batch)x-ai/grok-4.3:batch | 25.4 | 1M | $1 | $2 | — | |||
| SpaceXAI: Grok 4.3x-ai/grok-4.3 | 25.4 | 1M | $1.25 | $2.5 | — | |||
| DeepSeek V4 Flashdeepseek/deepseek-v4-flash | 24.8 | 1M | $0.15 | $0.6 | 2026-04-24 | |||
| Nemotron 3 Ultra 550B A55Bnvidia/nemotron-3-ultra-550b-a55b | 23.4 | 1M | $0.5 | $2.5 | 2026-06-04 | |||
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | 23.4 | 1M | Free | Free | — | |||
| Google: Gemini 3.5 Flash Lite (batch)google/gemini-3.5-flash-lite:batch | 22.7 | 1.04858M | $0.15 | $1.25 | — | |||
| Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 22.7 | 1.04858M | $0.3 | $2.5 | 2026-07-21 | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 22.3 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | 21.2 | 1M | $3 | $15 | — | |||
| Anthropic: Claude Sonnet 4.5 (batch)anthropic/claude-sonnet-4.5:batch | 21.2 | 1M | $1.5 | $7.5 | — | |||
| LongCat-2.0meituan/longcat-2.0 | 19.7 | 1M | $0.3 | $1.2 | 2026-06-30 | |||
| Amazon: Nova 2 Liteamazon/nova-2-lite-v1 | 18.4 | 1M | $0.3 | $2.5 | — | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 16.7 | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| Google: Gemini 2.5 Pro (batch)google/gemini-2.5-pro:batch | 16.7 | 1.04858M | $0.625 | $5 | — | |||
| Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 16.0 | 1.04858M | $0.25 | $1.5 | 2026-03-03 | |||
| OpenAI: GPT-4.1 Mini (batch)openai/gpt-4.1-mini:batch | 14.8 | 1.04758M | $0.2 | $0.8 | — | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 14.8 | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| NVIDIA: Nemotron 3.5 Lightning (free)nvidia/nemotron-3.5-lightning:free | 13.6 | 1M | Free | Free | — | |||
| OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch | 9.6 | 1.04758M | $0.05 | $0.2 | — | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 9.6 | 1.04758M | $0.1 | $0.4 | 2025-04-14 |