Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
Open MoE flagship with million-token context for coding and long agent runs
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
Open MiMo model for multimodal coding agents and long-context automation
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Multimodal MoE reasoning model (276B total, 12B active) for text, image, and audio
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Meituan LongCat-2.0, a reasoning model with tool calling and a 1M-token context window
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Low-latency Gemini model for high-volume multimodal and agent workloads
Google's proven reasoning model for coding, math, and multimodal analysis
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...
Affordable GPT-4.1 lane for fast coding help and structured extraction
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| Google: Gemini 3.1 Pro Preview (batch)google/gemini-3.1-pro-preview:batch | 68.8 | 1.04858M | $1 | $6 | — | |||
| Qwen: Qwen3.8 27Bqwen/qwen3.8-27b | 68.1 | 1M | $0.42 | $3 | — | |||
| Qwen: Qwen3.7 Maxqwen/qwen3.7-max | 66.0 | 1M | $1.475 | $4.425 | — | |||
| Anthropic: Claude Sonnet 4.6 (batch)anthropic/claude-sonnet-4.6:batch | 63.0 | 1M | $1.5 | $7.5 | — | |||
| Anthropic: Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | 63.0 | 1M | $3 | $15 | — | |||
| MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 60.2 | 1.04858M | $0.435 | $0.87 | 2026-04-22 | |||
| DeepSeek V4 Prodeepseek/deepseek-v4-pro | 59.4 | 1M | $0.435 | $0.87 | 2026-04-24 | |||
| MiniMax: MiniMax M3 (free)minimax/minimax-m3:free | 58.6 | 1.04858M | Free | Free | — | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 56.8 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| Qwen: Qwen3.7 Plusqwen/qwen3.7-plus | 55.9 | 1M | $0.32 | $1.28 | — | |||
| Qwen: Qwen3.6 Plusqwen/qwen3.6-plus | 54.5 | 1M | $0.325 | $1.95 | — | |||
| Thinking Machines: Inkling Small (free)thinkingmachines/inkling-small:free | 52.9 | 1.04858M | Free | Free | — | |||
| Inkling Smallthinkingmachines/inkling-small | 52.9 | 1.04858M | $0.45 | $1.2 | 2026-07-30 | |||
| Thinking Machines: Inkling (free)thinkingmachines/inkling:free | 52.1 | 1.04858M | Free | Free | — | |||
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | 52.1 | 1M | $3 | $15 | — | |||
| Inklingthinkingmachines/inkling | 52.1 | 1.04858M | $1.87 | $4.68 | 2026-07-15 | |||
| Anthropic: Claude Sonnet 4.5 (batch)anthropic/claude-sonnet-4.5:batch | 52.1 | 1M | $1.5 | $7.5 | — | |||
| DeepSeek V4 Flashdeepseek/deepseek-v4-flash | 52.0 | 1M | $0.15 | $0.6 | 2026-04-24 | |||
| Nemotron 3 Ultra 550B A55Bnvidia/nemotron-3-ultra-550b-a55b | 49.3 | 1M | $0.5 | $2.5 | 2026-06-04 | |||
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | 49.3 | 1M | Free | Free | — | |||
| Google: Gemini 3.5 Flash Lite (batch)google/gemini-3.5-flash-lite:batch | 49.3 | 1.04858M | $0.15 | $1.25 | — | |||
| Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 49.3 | 1.04858M | $0.3 | $2.5 | 2026-07-21 | |||
| LongCat-2.0meituan/longcat-2.0 | 45.3 | 1M | $0.3 | $1.2 | 2026-06-30 | |||
| SpaceXAI: Grok 4.3 (batch)x-ai/grok-4.3:batch | 42.2 | 1M | $1 | $2 | — | |||
| SpaceXAI: Grok 4.3x-ai/grok-4.3 | 42.2 | 1M | $1.25 | $2.5 | — | |||
| Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 34.7 | 1.04858M | $0.25 | $1.5 | 2026-03-03 | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 33.3 | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| Google: Gemini 2.5 Pro (batch)google/gemini-2.5-pro:batch | 33.3 | 1.04858M | $0.625 | $5 | — | |||
| NVIDIA: Nemotron 3.5 Lightning (free)nvidia/nemotron-3.5-lightning:free | 26.8 | 1M | Free | Free | — | |||
| Amazon: Nova 2 Liteamazon/nova-2-lite-v1 | 23.0 | 1M | $0.3 | $2.5 | — | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 20.2 | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| OpenAI: GPT-4.1 Mini (batch)openai/gpt-4.1-mini:batch | 20.2 | 1.04758M | $0.2 | $0.8 | — | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 11.1 | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch | 11.1 | 1.04758M | $0.05 | $0.2 | — |