Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...
Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts text and image input and produces...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...
Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...
GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...
Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to...
GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...
MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....
MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...
Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...
Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and operational efficiency. It is optimized for coding agents, tool...
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...
Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...
DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| Anthropic: Claude Sonnet 4.6 (batch)anthropic/claude-sonnet-4.6:batch | 33.1 | 1M | $1.5 | $7.5 | — | |||
| Anthropic: Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | 33.1 | 1M | $3 | $15 | — | |||
| Nex AGI: Nex-N2-Pronex-agi/nex-n2-pro | 31.2 | 262.144K | $0.25 | $1 | — | |||
| MiniMax: MiniMax M3 (free)minimax/minimax-m3:free | 31.0 | 1.04858M | Free | Free | — | |||
| MiniMax: MiniMax M3minimax/minimax-m3 | 30.8 | 524.288K | $0.3 | $1.2 | — | |||
| MiniMax: MiniMax M3 (batch)minimax/minimax-m3:batch | 30.8 | 524.288K | $0.3 | $1.2 | — | |||
| MoonshotAI: Kimi K2.7 Code (batch)moonshotai/kimi-k2.7-code:batch | 30.3 | 262.144K | $0.95 | $4 | — | |||
| Google: Gemini 3.6 Flashgoogle/gemini-3.6-flash | 30.2 | 1.04858M | $0.75 | $3.75 | 2026-07-21 | |||
| Google: Gemini 3.6 Flash (batch)google/gemini-3.6-flash:batch | 30.2 | 1.04858M | $0.375 | $1.875 | — | |||
| inclusionAI: Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl | 30.0 | 131.072K | $0.06 | $0.18 | — | |||
| inclusionAI: Ling 3.0 Flash VL (free)inclusionai/ling-3.0-flash-vl:free | 30.0 | 262.144K | Free | Free | — | |||
| Ling-3.0-flash (free)inclusionai/ling-3.0-flash:free | 29.6 | 262.144K | Free | Free | — | |||
| Qwen: Qwen3.6 Plusqwen/qwen3.6-plus | 29.0 | 1M | $0.325 | $1.95 | — | |||
| SpaceXAI: Grok Build 0.1x-ai/grok-build-0.1 | 28.9 | 256K | $1 | $2 | — | |||
| DeepSeek: DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | 27.9 | 1.024M | $0.085 | $0.171 | 2026-04-24 | |||
| DeepSeek: DeepSeek V4 Pro 0423deepseek/deepseek-v4-pro | 27.7 | 1.024M | $0.86 | $1.72 | 2026-04-24 | |||
| Meta: Muse Spark 1.1meta/muse-spark-1.1 | 27.5 | 1.04858M | $1.25 | $4.25 | 2026-04-08 | |||
| NVIDIA: Nemotron 3 Ultra (batch)nvidia/nemotron-3-ultra-550b-a55b:batch | 27.5 | 512.288K | $0.6 | $3.6 | — | |||
| Google: Gemini 3.5 Flash (batch)google/gemini-3.5-flash:batch | 27.3 | 1.04858M | $0.75 | $4.5 | — | |||
| Google: Gemini 3.5 Flashgoogle/gemini-3.5-flash | 27.3 | 1.04858M | $1.5 | $9 | 2026-05-19 | |||
| OpenAI: GPT-5openai/gpt-5 | 26.5 | 400K | $1.25 | $10 | 2025-08-07 | |||
| OpenAI: GPT-5 (batch)openai/gpt-5:batch | 26.5 | 400K | $0.625 | $5 | — | |||
| Z.ai: GLM 4.7z-ai/glm-4.7 | 26.2 | 202.752K | $0.4 | $1.75 | — | |||
| MiniMax: MiniMax M2.7 (free)minimax/minimax-m2.7:free | 25.9 | 196.608K | Free | Free | — | |||
| Tencent: Hy3 previewtencent/hy3-preview | 25.6 | 262.144K | $0.18 | $0.6 | 2026-04-20 | |||
| Z.ai: GLM 5.1z-ai/glm-5.1 | 25.2 | 200K | $0.966 | $3.036 | — | |||
| Thinking Machines: Inkling Smallthinkingmachines/inkling-small | 25.0 | 524.288K | $0.45 | $1.2 | 2026-07-30 | |||
| Thinking Machines: Inkling Small (free)thinkingmachines/inkling-small:free | 25.0 | 1.04858M | Free | Free | — | |||
| Thinking Machines: Inkling Small (batch)thinkingmachines/inkling-small:batch | 25.0 | 524.288K | $0.5 | $1.2 | — | |||
| Thinking Machines: Inklingthinkingmachines/inkling | 24.3 | 1.04858M | $1 | $4.05 | 2026-07-15 | |||
| Thinking Machines: Inkling (batch)thinkingmachines/inkling:batch | 24.3 | 524.288K | $1 | $4.05 | — | |||
| Thinking Machines: Inkling (free)thinkingmachines/inkling:free | 24.3 | 1.04858M | Free | Free | — | |||
| Qwen: Qwen3.7 Maxqwen/qwen3.7-max | 23.9 | 1M | $1.475 | $4.425 | — | |||
| Xiaomi: MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 22.7 | 1.04858M | $0.435 | $0.87 | 2026-04-22 | |||
| MoonshotAI: Kimi K2.7 Codemoonshotai/kimi-k2.7-code | 22.5 | 262.144K | $0.71 | $3.5 | 2026-06-12 | |||
| MoonshotAI: Kimi K2.6moonshotai/kimi-k2.6 | 22.1 | 262.144K | $0.95 | $4 | 2026-04-21 | |||
| StepFun: Step 3.7 Flashstepfun/step-3.7-flash | 21.7 | 256K | $0.2 | $1.15 | 2026-05-29 | |||
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | 21.7 | 1M | Free | Free | — | |||
| MoonshotAI: Kimi K2.5moonshotai/kimi-k2.5 | 21.7 | 262.144K | $0.45 | $2.25 | 2026-01 | |||
| NVIDIA: Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b | 21.7 | 256K | $0.625 | $3.125 | 2026-06-04 | |||
| OpenAI: GPT-5.1openai/gpt-5.1 | 21.6 | 400K | $1.25 | $10 | 2025-11-13 | |||
| OpenAI: GPT-5.1 (batch)openai/gpt-5.1:batch | 21.6 | 400K | $0.625 | $5 | — | |||
| inclusionAI: Ling 3.0 Flashinclusionai/ling-3.0-flash | 21.0 | 262.144K | $0.021 | $0.063 | — | |||
| Qwen: Qwen3.6 27Bqwen/qwen3.6-27b | 20.1 | 262.144K | $0.3 | $2 | — | |||
| inclusionAI: Ring-2.6-1Tinclusionai/ring-2.6-1t | 20.1 | 262.144K | $0.075 | $0.625 | — | |||
| OpenAI: GPT-5.4 Mini (batch)openai/gpt-5.4-mini:batch | 19.7 | 400K | $0.375 | $2.25 | — | |||
| Qwen: Qwen3.7 Plusqwen/qwen3.7-plus | 19.7 | 1M | $0.32 | $1.28 | — | |||
| OpenAI: GPT-5.4 Miniopenai/gpt-5.4-mini | 19.7 | 400K | $0.75 | $4.5 | 2026-03-17 | |||
| Z.ai: GLM 4.6z-ai/glm-4.6 | 18.6 | 198K | $0.43 | $1.75 | — | |||
| DeepSeek: DeepSeek V3.2deepseek/deepseek-v3.2 | 18.3 | 163.84K | $0.269 | $0.4 | 2025-12-01 |