GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Reasoning-first Gemini preview for agentic coding and complex problem solving
GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...
DeepSeek V4 Pro snapshot with million-token context and support for thinking and non-thinking modes
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...
Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration. It builds on the agentic coding strengths of earlier versions,...
Open MoE flagship with million-token context for coding and long agent runs
Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts text and image input and produces...
Tencent Hy reasoning model for coding, instruction following, and agent tasks
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
Open MiMo model for multimodal coding agents and long-context automation
Strong small GPT for coding subagents, quick tool use, and high-volume work
Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...
GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Multimodal MoE reasoning model (276B total, 12B active) for text, image, and audio
Upstage's flagship model, specialized for agentic use
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
Sharper GPT-5 generation for coding, product work, and tool-assisted tasks
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
Fast Gemini model balancing multimodal reasoning, tool use, and cost
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| Z.ai: GLM 5.2z-ai/glm-5.2 | 68.8 | 1M | $0.966 | $3.036 | — | |||
| Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | 68.8 | 1.04858M | $2 | $12 | 2026-02-19 | |||
| Z.ai: GLM 5.2 (batch)z-ai/glm-5.2:batch | 68.8 | 1.04858M | $0.7 | $2.2 | — | |||
| DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 | 68.8 | 1M | $0.442 | $0.884 | 2026-08-12 | |||
| Google: Gemini 3.1 Pro Preview (batch)google/gemini-3.1-pro-preview:batch | 68.8 | 1.04858M | $1 | $6 | — | |||
| Qwen: Qwen3.8 27Bqwen/qwen3.8-27b | 68.1 | 1M | $0.42 | $3 | — | |||
| Qwen: Qwen3.7 Maxqwen/qwen3.7-max | 66.0 | 1M | $1.475 | $4.425 | — | |||
| Anthropic: Claude Sonnet 4.6 (batch)anthropic/claude-sonnet-4.6:batch | 63.0 | 1M | $1.5 | $7.5 | — | |||
| Anthropic: Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | 63.0 | 1M | $3 | $15 | — | |||
| Kimi K2.6moonshotai/kimi-k2.6 | 61.8 | 262.144K | $0.95 | $4 | 2026-04-21 | |||
| MoonshotAI: Kimi K2.7 Code (batch)moonshotai/kimi-k2.7-code:batch | 60.8 | 262.144K | $0.95 | $4 | — | |||
| Kimi K2.7 Codemoonshotai/kimi-k2.7-code | 60.8 | 262.144K | $0.95 | $4 | 2026-06-12 | |||
| MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 60.2 | 1.04858M | $0.435 | $0.87 | 2026-04-22 | |||
| Kwaipilot: KAT-Coder-Pro V2kwaipilot/kat-coder-pro-v2 | 59.5 | 262.144K | $0.3 | $1.2 | — | |||
| DeepSeek V4 Prodeepseek/deepseek-v4-pro | 59.4 | 1M | $0.435 | $0.87 | 2026-04-24 | |||
| Nex AGI: Nex-N2-Pronex-agi/nex-n2-pro | 59.1 | 262.144K | $0.25 | $1 | — | |||
| Hy3 previewtencent/hy3-preview | 58.8 | 256K | $0.066 | $0.26 | 2026-04-20 | |||
| MiniMax: MiniMax M3minimax/minimax-m3 | 58.6 | 524.288K | $0.3 | $1.2 | — | |||
| MiniMax: MiniMax M3 (batch)minimax/minimax-m3:batch | 58.6 | 524.288K | $0.3 | $1.2 | — | |||
| MiniMax: MiniMax M3 (free)minimax/minimax-m3:free | 58.6 | 1.04858M | Free | Free | — | |||
| inclusionAI: Ling 3.0 Flash VL (free)inclusionai/ling-3.0-flash-vl:free | 57.0 | 262.144K | Free | Free | — | |||
| inclusionAI: Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl | 57.0 | 131.072K | $0.06 | $0.18 | — | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 56.8 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| GPT-5.4 miniopenai/gpt-5.4-mini | 56.1 | 400K | $0.75 | $4.5 | 2026-03-17 | |||
| GPT-5.4 nanoopenai/gpt-5.4-nano | 56.1 | 400K | $0.2 | $1.25 | 2026-03-17 | |||
| OpenAI: GPT-5.4 Mini (batch)openai/gpt-5.4-mini:batch | 56.1 | 400K | $0.375 | $2.25 | — | |||
| OpenAI: GPT-5.4 Nano (batch)openai/gpt-5.4-nano:batch | 56.1 | 400K | $0.1 | $0.625 | — | |||
| Qwen: Qwen3.7 Plusqwen/qwen3.7-plus | 55.9 | 1M | $0.32 | $1.28 | — | |||
| Z.ai: GLM 5.1z-ai/glm-5.1 | 55.8 | 200K | $0.966 | $3.036 | — | |||
| Qwen: Qwen3.6 Plusqwen/qwen3.6-plus | 54.5 | 1M | $0.325 | $1.95 | — | |||
| Qwen: Qwen3.6 27Bqwen/qwen3.6-27b | 53.7 | 262.144K | $0.3 | $2 | — | |||
| Thinking Machines: Inkling Small (batch)thinkingmachines/inkling-small:batch | 52.9 | 524.288K | $0.5 | $1.2 | — | |||
| Thinking Machines: Inkling Small (free)thinkingmachines/inkling-small:free | 52.9 | 1.04858M | Free | Free | — | |||
| Inkling Smallthinkingmachines/inkling-small | 52.9 | 1.04858M | $0.45 | $1.2 | 2026-07-30 | |||
| Solar Pro 4upstage/solar-pro4 | 52.7 | 524.288K | $0.3 | $1.2 | 2026-08-06 | |||
| MiniMax: MiniMax M2.7 (free)minimax/minimax-m2.7:free | 52.6 | 196.608K | Free | Free | — | |||
| MiniMax: MiniMax M2.7minimax/minimax-m2.7 | 52.6 | 204.8K | $0.3 | $1.2 | — | |||
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | 52.1 | 1M | $3 | $15 | — | |||
| Inklingthinkingmachines/inkling | 52.1 | 1.04858M | $1.87 | $4.68 | 2026-07-15 | |||
| Thinking Machines: Inkling (batch)thinkingmachines/inkling:batch | 52.1 | 524.288K | $1 | $4.05 | — | |||
| Thinking Machines: Inkling (free)thinkingmachines/inkling:free | 52.1 | 1.04858M | Free | Free | — | |||
| Anthropic: Claude Sonnet 4.5 (batch)anthropic/claude-sonnet-4.5:batch | 52.1 | 1M | $1.5 | $7.5 | — | |||
| DeepSeek V4 Flashdeepseek/deepseek-v4-flash | 52.0 | 1M | $0.15 | $0.6 | 2026-04-24 | |||
| SpaceXAI: Grok Build 0.1x-ai/grok-build-0.1 | 51.5 | 256K | $1 | $2 | — | |||
| inclusionAI: Ling 3.0 Flashinclusionai/ling-3.0-flash | 50.6 | 262.144K | $0.021 | $0.063 | — | |||
| Ling-3.0-flash (free)inclusionai/ling-3.0-flash:free | 50.6 | 262.144K | Free | Free | — | |||
| GPT-5.1openai/gpt-5.1 | 49.4 | 400K | $1.25 | $10 | 2025-11-13 | |||
| OpenAI: GPT-5.1 (batch)openai/gpt-5.1:batch | 49.4 | 400K | $0.625 | $5 | — | |||
| Nemotron 3 Ultra 550B A55Bnvidia/nemotron-3-ultra-550b-a55b | 49.3 | 1M | $0.5 | $2.5 | 2026-06-04 | |||
| Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 49.3 | 1.04858M | $0.3 | $2.5 | 2026-07-21 |