GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...
Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...
MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...
Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed...
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...
GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...
Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...
GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...
Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive...
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...
GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...
MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...
MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world...
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...
Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...
DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...
Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...
Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| Z.ai: GLM 5z-ai/glm-5 | 1257.0 | 198K | $0.6 | $1.92 | — | |||
| Anthropic: Claude Opus 4.8anthropic/claude-opus-4.8 | 1252.0 | 1M | $5 | $25 | — | |||
| Anthropic: Claude Opus 4.8 (batch)anthropic/claude-opus-4.8:batch | 1252.0 | 1M | $2.5 | $12.5 | — | |||
| MiniMax: MiniMax M3 (free)minimax/minimax-m3:free | 1243.0 | 1.04858M | Free | Free | — | |||
| Anthropic: Claude Opus 4.5anthropic/claude-opus-4.5 | 1242.0 | 200K | $5 | $25 | — | |||
| Anthropic: Claude Opus 4.5 (batch)anthropic/claude-opus-4.5:batch | 1242.0 | 200K | $2.5 | $12.5 | — | |||
| Xiaomi: MiMo-V2.5xiaomi/mimo-v2.5 | 1241.0 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| MiniMax: MiniMax M3minimax/minimax-m3 | 1239.0 | 524.288K | $0.3 | $1.2 | — | |||
| MiniMax: MiniMax M3 (batch)minimax/minimax-m3:batch | 1239.0 | 524.288K | $0.3 | $1.2 | — | |||
| DeepSeek: DeepSeek V4 Flash 0731 (batch)deepseek/deepseek-v4-flash-0731:batch | 1238.0 | 1.04858M | $0.11 | $0.33 | — | |||
| DeepSeek: DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | 1238.0 | 1.04858M | $0.065 | $0.18 | 2026-07-31 | |||
| Z.ai: GLM 5V Turboz-ai/glm-5v-turbo | 1238.0 | 202.752K | $1.2 | $4 | — | |||
| MoonshotAI: Kimi K2.5moonshotai/kimi-k2.5 | 1234.0 | 262.144K | $0.45 | $2.25 | 2026-01 | |||
| Qwen: Qwen3.6 Plusqwen/qwen3.6-plus | 1230.0 | 1M | $0.325 | $1.95 | — | |||
| OpenAI: GPT-5.5openai/gpt-5.5 | 1228.0 | 1.05M | $5 | $30 | 2026-04-23 | |||
| OpenAI: GPT-5.5 (batch)openai/gpt-5.5:batch | 1228.0 | 1.05M | $2.5 | $15 | — | |||
| MiniMax: MiniMax M2.7 (free)minimax/minimax-m2.7:free | 1225.0 | 196.608K | Free | Free | — | |||
| SpaceXAI: Grok 4.20x-ai/grok-4.20 | 1222.0 | 2M | $1.25 | $2.5 | — | |||
| MiniMax: MiniMax M2.7minimax/minimax-m2.7 | 1222.0 | 204.8K | $0.3 | $1.2 | — | |||
| Z.ai: GLM 4.7z-ai/glm-4.7 | 1219.0 | 202.752K | $0.4 | $1.75 | — | |||
| DeepSeek: DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | 1217.0 | 1.024M | $0.086 | $0.171 | 2026-04-24 | |||
| Google: Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | 1212.0 | 1.04858M | $0.5 | $3 | 2025-12-17 | |||
| Google: Gemini 3 Flash Preview (batch)google/gemini-3-flash-preview:batch | 1212.0 | 1.04858M | $0.25 | $1.5 | — | |||
| Upstage: Solar Pro 4upstage/solar-pro4 | 1207.0 | 524.288K | $0.09 | $0.36 | 2026-08-06 | |||
| Tencent: Hy3tencent/hy3 | 1205.0 | 262.144K | $0.083 | $0.33 | 2026-07-06 | |||
| Z.ai: GLM 4.5z-ai/glm-4.5 | 1199.0 | 131.072K | $0.6 | $2.2 | — | |||
| MiniMax: MiniMax M2.5minimax/minimax-m2.5 | 1198.0 | 200K | $0.27 | $1.08 | — | |||
| MiniMax: MiniMax M2.1minimax/minimax-m2.1 | 1194.0 | 204.8K | $0.3 | $1.2 | — | |||
| Qwen: Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | 1190.0 | 262.144K | $0.55 | $3.5 | — | |||
| Thinking Machines: Inkling (free)thinkingmachines/inkling:free | 1188.0 | 1.04858M | Free | Free | — | |||
| Thinking Machines: Inklingthinkingmachines/inkling | 1188.0 | 1.04858M | $1 | $4.05 | 2026-07-15 | |||
| Thinking Machines: Inkling (batch)thinkingmachines/inkling:batch | 1188.0 | 524.288K | $1 | $4.05 | — | |||
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | 1186.0 | 1M | $3 | $15 | — | |||
| Anthropic: Claude Sonnet 4.5 (batch)anthropic/claude-sonnet-4.5:batch | 1186.0 | 1M | $1.5 | $7.5 | — | |||
| Anthropic: Claude Opus 4.1 (batch)anthropic/claude-opus-4.1:batch | 1182.0 | 200K | $7.5 | $37.5 | — | |||
| Anthropic: Claude Opus 4.1anthropic/claude-opus-4.1 | 1182.0 | 200K | $15 | $75 | — | |||
| DeepSeek: DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp | 1177.0 | 163.84K | $0.27 | $0.41 | — | |||
| NVIDIA: Nemotron 3 Ultra (batch)nvidia/nemotron-3-ultra-550b-a55b:batch | 1173.0 | 512.288K | $0.6 | $3.6 | — | |||
| DeepSeek: DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminus | 1169.0 | 131.072K | $0.27 | $1 | — | |||
| Anthropic: Claude Sonnet 4anthropic/claude-sonnet-4 | 1168.0 | 200K | $3 | $15 | — | |||
| Anthropic: Claude Opus 4anthropic/claude-opus-4 | 1167.0 | 200K | $15 | $75 | — | |||
| NVIDIA: Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b | 1163.0 | 256K | $0.625 | $3.125 | 2026-06-04 | |||
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | 1163.0 | 1M | Free | Free | — | |||
| DeepSeek: DeepSeek V3.2deepseek/deepseek-v3.2 | 1161.0 | 163.84K | $0.269 | $0.4 | 2025-12-01 | |||
| Z.ai: GLM 4.6z-ai/glm-4.6 | 1157.0 | 198K | $0.43 | $1.75 | — | |||
| StepFun: Step 3.7 Flashstepfun/step-3.7-flash | 1156.0 | 256K | $0.2 | $1.15 | 2026-05-29 | |||
| SpaceXAI: Grok 4.3x-ai/grok-4.3 | 1153.0 | 1M | $1.25 | $2.5 | — | |||
| SpaceXAI: Grok 4.3 (batch)x-ai/grok-4.3:batch | 1153.0 | 1M | $1 | $2 | — | |||
| Z.ai: GLM 4.5 Airz-ai/glm-4.5-air | 1151.0 | 131.072K | $0.13 | $0.85 | — | |||
| Z.ai: GLM 4.7 Flashz-ai/glm-4.7-flash | 1149.0 | 131.072K | $0.061 | $0.4 | — |