NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...
Cost-efficient GPT-5.6 model for fast, high-volume workloads
Sharper GPT-5 generation for coding, product work, and tool-assisted tasks
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
DeepSeek V4 Pro snapshot with million-token context and support for thinking and non-thinking modes
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...
Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows
Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding
GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...
KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration. It builds on the agentic coding strengths of earlier versions,...
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
Hybrid-reasoning DeepSeek model with thinking and non-thinking modes, sparse attention, and tool-use
Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and operational efficiency. It is optimized for coding agents, tool...
Open MoE flagship with million-token context for coding and long agent runs
Newer StepFun flash model for faster agents, coding, and multimodal prompts
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...
Reasoning-first Gemini preview for agentic coding and complex problem solving
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Open Gemma instruction model for efficient chat and self-hosted deployments
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Multimodal MoE reasoning model (276B total, 12B active) for text, image, and audio
Tencent Hy reasoning model for coding, instruction following, and agent tasks
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| NVIDIA: Nemotron 3 Ultra (batch)nvidia/nemotron-3-ultra-550b-a55b:batch | 38.3 | 512.288K | $0.6 | $3.6 | — | |||
| OpenAI: GPT-5.6 Luna (batch)openai/gpt-5.6-luna:batch | 37.5 | 1.05M | $0.1 | $0.6 | — | |||
| GPT-5.6 Lunaopenai/gpt-5.6-luna | 37.5 | 1.05M | $0.2 | $1.2 | 2026-07-09 | |||
| GPT-5.1openai/gpt-5.1 | 37.5 | 400K | $1.25 | $10 | 2025-11-13 | |||
| OpenAI: GPT-5.1 (batch)openai/gpt-5.1:batch | 37.5 | 400K | $0.625 | $5 | — | |||
| Ling-3.0-flash (free)inclusionai/ling-3.0-flash:free | 37.4 | 262.144K | Free | Free | — | |||
| DeepSeek: DeepSeek V4 Pro 0813 (batch)deepseek/deepseek-v4-pro-0813:batch | 36.3 | 1.04858M | $0.66 | $1.98 | — | |||
| DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 | 36.3 | 1M | $0.442 | $0.884 | 2026-08-12 | |||
| Kimi K2.5moonshotai/kimi-k2.5 | 36.0 | 262.144K | $0.3 | $1.9 | 2026-01 | |||
| MiniMax: MiniMax M3 (free)minimax/minimax-m3:free | 35.7 | 1.04858M | Free | Free | — | |||
| OpenAI: GPT-5 (batch)openai/gpt-5:batch | 35.3 | 400K | $0.625 | $5 | — | |||
| GPT-5openai/gpt-5 | 35.3 | 400K | $1.25 | $10 | 2025-08-07 | |||
| DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | 34.5 | 1M | $0.05 | $0.16 | 2026-07-31 | |||
| Z.ai: GLM 4.7z-ai/glm-4.7 | 34.5 | 202.752K | $0.4 | $1.75 | — | |||
| DeepSeek: DeepSeek V4 Flash 0731 (batch)deepseek/deepseek-v4-flash-0731:batch | 34.5 | 1.04858M | $0.11 | $0.33 | — | |||
| Gemini 3.6 Flashgoogle/gemini-3.6-flash | 34.3 | 1.04858M | $0.75 | $3.75 | 2026-07-21 | |||
| Muse Spark 1.1meta/muse-spark-1.1 | 34.3 | 1.04858M | $1.25 | $4.25 | 2026-04-08 | |||
| Google: Gemini 3.6 Flash (batch)google/gemini-3.6-flash:batch | 34.3 | 1.04858M | $0.375 | $1.875 | — | |||
| Qwen: Qwen3.8 27Bqwen/qwen3.8-27b | 33.9 | 1M | $0.42 | $3 | — | |||
| Kwaipilot: KAT-Coder-Pro V2kwaipilot/kat-coder-pro-v2 | 33.7 | 262.144K | $0.3 | $1.2 | — | |||
| Gemini 3.5 Flashgoogle/gemini-3.5-flash | 33.0 | 1.04858M | $1.5 | $9 | 2026-05-19 | |||
| Google: Gemini 3.5 Flash (batch)google/gemini-3.5-flash:batch | 33.0 | 1.04858M | $0.75 | $4.5 | — | |||
| DeepSeek V3.2deepseek/deepseek-v3.2 | 32.6 | 128K | $0.18 | $0.35 | 2025-12-01 | |||
| inclusionAI: Ring-2.6-1Tinclusionai/ring-2.6-1t | 31.7 | 262.144K | $0.075 | $0.625 | — | |||
| DeepSeek V4 Prodeepseek/deepseek-v4-pro | 30.9 | 1M | $0.435 | $0.87 | 2026-04-24 | |||
| Step 3.7 Flashstepfun/step-3.7-flash | 30.9 | 256K | $0.185 | $1.11 | 2026-05-29 | |||
| Anthropic: Claude Sonnet 4.6 (batch)anthropic/claude-sonnet-4.6:batch | 30.5 | 1M | $1.5 | $7.5 | — | |||
| Anthropic: Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | 30.5 | 1M | $3 | $15 | — | |||
| Google: Gemini 3.1 Pro Preview (batch)google/gemini-3.1-pro-preview:batch | 30.4 | 1.04858M | $1 | $6 | — | |||
| Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | 30.4 | 1.04858M | $2 | $12 | 2026-02-19 | |||
| Qwen: Qwen3.7 Maxqwen/qwen3.7-max | 29.9 | 1M | $1.475 | $4.425 | — | |||
| Anthropic: Claude Sonnet 4anthropic/claude-sonnet-4 | 29.8 | 200K | $3 | $15 | — | |||
| MiniMax: MiniMax M3minimax/minimax-m3 | 29.6 | 524.288K | $0.3 | $1.2 | — | |||
| MiniMax: MiniMax M3 (batch)minimax/minimax-m3:batch | 29.6 | 524.288K | $0.3 | $1.2 | — | |||
| Z.ai: GLM 4.6z-ai/glm-4.6 | 29.3 | 198K | $0.43 | $1.75 | — | |||
| inclusionAI: Ling 3.0 Flashinclusionai/ling-3.0-flash | 27.4 | 262.144K | $0.021 | $0.063 | — | |||
| Z.ai: GLM 5.1z-ai/glm-5.1 | 26.4 | 200K | $0.966 | $3.036 | — | |||
| MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 26.4 | 1.04858M | $0.435 | $0.87 | 2026-04-22 | |||
| Kimi K2.7 Codemoonshotai/kimi-k2.7-code | 26.3 | 262.144K | $0.95 | $4 | 2026-06-12 | |||
| Thinking Machines: Inkling Small (batch)thinkingmachines/inkling-small:batch | 26.1 | 524.288K | $0.5 | $1.2 | — | |||
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 26.1 | 262.144K | $0.042 | $0.22 | 2026-04-02 | |||
| Google: Gemma 4 26B A4B (free)google/gemma-4-26b-a4b-it:free | 26.1 | 262.144K | Free | Free | — | |||
| Thinking Machines: Inkling Small (free)thinkingmachines/inkling-small:free | 26.1 | 1.04858M | Free | Free | — | |||
| Inkling Smallthinkingmachines/inkling-small | 26.1 | 1.04858M | $0.45 | $1.2 | 2026-07-30 | |||
| Hy3 previewtencent/hy3-preview | 25.8 | 256K | $0.066 | $0.26 | 2026-04-20 | |||
| Qwen: Qwen3.7 Plusqwen/qwen3.7-plus | 25.8 | 1M | $0.32 | $1.28 | — | |||
| Inklingthinkingmachines/inkling | 25.5 | 1.04858M | $1.87 | $4.68 | 2026-07-15 | |||
| Thinking Machines: Inkling (batch)thinkingmachines/inkling:batch | 25.5 | 524.288K | $1 | $4.05 | — | |||
| Thinking Machines: Inkling (free)thinkingmachines/inkling:free | 25.5 | 1.04858M | Free | Free | — | |||
| SpaceXAI: Grok 4.3x-ai/grok-4.3 | 25.4 | 1M | $1.25 | $2.5 | — |