Coding-optimized GPT model for repository edits, reviews, and agentic software work
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...
GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...
Hybrid-reasoning DeepSeek model with thinking and non-thinking modes, sparse attention, and tool-use
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...
DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...
Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in...
Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...
Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,...
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...
Fast Gemini workhorse for multimodal apps where latency and price matter
Small GPT-5 for responsive agents, coding help, and everyday automation
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Tencent Hy reasoning model for coding, instruction following, and agent tasks
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...
Long-lived GPT workhorse for coding, instruction following, and production apps
GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...
Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It...
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...
Reasoning-optimized 398B MoE agent model with extended thinking for long-horizon and multi-turn tool use
Coding-optimized GPT model for repository edits, reviews, and agentic software work
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...
DeepSeek chat model for instruction following, coding, and analysis
Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the...
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...
GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...
Tiny GPT-5 lane for routing, extraction, classification, and bulk jobs
Low-latency Gemini model for high-volume multimodal and agent workloads
Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances state-of-the-art reasoning and multimodal performance with 8× lower cost...
Affordable GPT-4.1 lane for fast coding help and structured extraction
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...
Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)
Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)
Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for...
Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving...
Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless switching between a "thinking" mode for complex reasoning, math, and...
OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning...
Fast o-series model for compact reasoning, coding, and tool use
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| GPT-5.3 Codexopenai/gpt-5.3-codex | 1180.0 | 400K | $1.75 | $14 | 2026-02-05 | |||
| Z.ai: GLM 4.6z-ai/glm-4.6 | 1177.0 | 198K | $0.43 | $1.75 | — | |||
| Z.ai: GLM 4.5z-ai/glm-4.5 | 1175.0 | 131.072K | $0.6 | $2.2 | — | |||
| DeepSeek V3.2deepseek/deepseek-v3.2 | 1174.0 | 128K | $0.18 | $0.35 | 2025-12-01 | |||
| Anthropic: Claude Sonnet 4anthropic/claude-sonnet-4 | 1173.0 | 200K | $3 | $15 | — | |||
| DeepSeek: DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp | 1168.0 | 163.84K | $0.27 | $0.41 | — | |||
| Anthropic: Claude Opus 4anthropic/claude-opus-4 | 1161.0 | 200K | $15 | $75 | — | |||
| Mistral: Mistral Medium 3.1mistralai/mistral-medium-3.1 | 1160.0 | 131.072K | $0.4 | $2 | — | |||
| Mistral: Mistral Medium 3.1 (batch)mistralai/mistral-medium-3.1:batch | 1160.0 | 131.072K | $0.2 | $1 | — | |||
| NVIDIA: Nemotron 3 Ultra (batch)nvidia/nemotron-3-ultra-550b-a55b:batch | 1158.0 | 512.288K | $0.6 | $3.6 | — | |||
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | 1156.0 | 1M | Free | Free | — | |||
| Nemotron 3 Ultra 550B A55Bnvidia/nemotron-3-ultra-550b-a55b | 1156.0 | 1M | $0.5 | $2.5 | 2026-06-04 | |||
| MiniMax: MiniMax M2minimax/minimax-m2 | 1156.0 | 204.8K | $0.255 | $1.02 | — | |||
| Qwen: Qwen3.5 Plus 2026-02-15qwen/qwen3.5-plus-02-15 | 1152.0 | 1M | $0.26 | $1.56 | — | |||
| Mistral: Mistral Large 3 2512 (batch)mistralai/mistral-large-2512:batch | 1152.0 | 262.144K | $0.25 | $0.75 | — | |||
| Mistral: Mistral Large 3 2512mistralai/mistral-large-2512 | 1152.0 | 262.144K | $0.5 | $1.5 | — | |||
| Google: Gemini 2.5 Flash (batch)google/gemini-2.5-flash:batch | 1149.0 | 1.04858M | $0.15 | $1.25 | — | |||
| Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1149.0 | 1.04858M | $0.3 | $2.5 | 2025-06-17 | |||
| GPT-5 Miniopenai/gpt-5-mini | 1146.0 | 400K | $0.25 | $2 | 2025-08-07 | |||
| OpenAI: GPT-5 Mini (batch)openai/gpt-5-mini:batch | 1146.0 | 400K | $0.125 | $1 | — | |||
| DeepSeek V4 Flashdeepseek/deepseek-v4-flash | 1142.0 | 1M | $0.15 | $0.6 | 2026-04-24 | |||
| Hy3tencent/hy3 | 1141.0 | 256K | $0.066 | $0.26 | 2026-07-06 | |||
| Anthropic: Claude Haiku 4.5 (batch)anthropic/claude-haiku-4.5:batch | 1140.0 | 200K | $0.5 | $2.5 | — | |||
| Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5 | 1140.0 | 200K | $1 | $5 | — | |||
| Z.ai: GLM 4.7 Flashz-ai/glm-4.7-flash | 1139.0 | 131.072K | $0.061 | $0.4 | — | |||
| GPT-4.1openai/gpt-4.1 | 1118.0 | 1.04758M | $2 | $8 | 2025-04-14 | |||
| OpenAI: GPT-4.1 (batch)openai/gpt-4.1:batch | 1118.0 | 1.04758M | $1 | $4 | — | |||
| Qwen: Qwen3 Maxqwen/qwen3-max | 1116.0 | 262.144K | $0.78 | $3.9 | — | |||
| Qwen: Qwen3 Coder 480B A35B (free)qwen/qwen3-coder:free | 1116.0 | 262K | Free | Free | — | |||
| Trinity Large Thinkingarcee-ai/trinity-large-thinking | 1114.0 | 524.288K | $0.25 | $0.9 | 2026-04-01 | |||
| GPT-5.1 Codex miniopenai/gpt-5.1-codex-mini | 1113.0 | 400K | $0.22 | $1.8 | 2025-11-13 | |||
| DeepSeek: DeepSeek V3.1deepseek/deepseek-chat-v3.1 | 1111.0 | 163.84K | $0.25 | $0.95 | — | |||
| DeepSeek Chatdeepseek/deepseek-chat | 1104.0 | 1M | $0.147 | $0.295 | 2025-12-01 | |||
| Qwen: Qwen3 Coder 30B A3B Instructqwen/qwen3-coder-30b-a3b-instruct | 1098.0 | 262.144K | $0.07 | $0.28 | — | |||
| Qwen: Qwen3 Coder 480B A35Bqwen/qwen3-coder | 1097.0 | 262.144K | $0.3 | $1 | — | |||
| Qwen: Qwen3 235B A22B Instruct 2507qwen/qwen3-235b-a22b-2507 | 1085.0 | 262.144K | $0.087 | $0.35 | — | |||
| OpenAI: GPT-5 Nano (batch)openai/gpt-5-nano:batch | 1071.0 | 400K | $0.025 | $0.2 | — | |||
| GPT-5 Nanoopenai/gpt-5-nano | 1071.0 | 400K | $0.05 | $0.4 | 2025-08-07 | |||
| Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 1061.0 | 1.04858M | $0.25 | $1.5 | 2026-03-03 | |||
| Mistral: Mistral Medium 3mistralai/mistral-medium-3 | 1049.0 | 131.072K | $0.4 | $2 | — | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1048.0 | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| OpenAI: GPT-4.1 Mini (batch)openai/gpt-4.1-mini:batch | 1048.0 | 1.04758M | $0.2 | $0.8 | — | |||
| OpenAI: gpt-oss-120b (free)openai/gpt-oss-120b:free | 1040.0 | 131.072K | Free | Free | — | |||
| Mistral: Codestral 2508mistralai/codestral-2508 | 1033.0 | 256K | $0.3 | $0.9 | — | |||
| Mistral: Codestral 2508 (batch)mistralai/codestral-2508:batch | 1033.0 | 256K | $0.15 | $0.45 | — | |||
| MoonshotAI: Kimi K2 0711moonshotai/kimi-k2 | 1032.0 | 131.072K | $0.57 | $2.3 | — | |||
| Inception: Mercury 2inception/mercury-2 | 1011.0 | 128K | $0.25 | $0.75 | — | |||
| Qwen: Qwen3 235B A22Bqwen/qwen3-235b-a22b | 1011.0 | 131.072K | $0.455 | $1.82 | — | |||
| OpenAI: o4 Mini (batch)openai/o4-mini:batch | 1006.0 | 200K | $0.55 | $2.2 | — | |||
| o4-miniopenai/o4-mini | 1006.0 | 200K | $1.1 | $4.4 | 2025-04-16 |