Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
No provider description is available for this model yet.
No provider description is available for this model yet.
Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost combination.
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model supports the `reasoning_effort` parameter, which can be set to...
GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256K context window. Devstral 2 supports exploring...
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...
Relace Apply 3 is a specialized code-patching LLM that merges AI-suggested edits straight into your source files. It can apply updates from GPT-4o, Claude, and others into your files at...
GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It...
Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...
Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)
OpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high. OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining...
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Mistral: Codestral 2508 (batch)mistralai/codestral-2508:batch | 256K | $0.15 | $0.45 | — | |||
| qwen3-next-80b-a3b-thinkingqwen_ai_platform/qwen3-next-80b-a3b-thinking | 262.144K | $0.15 | $1.2 | — | |||
| qwen3-vl-plusqwen_ai_platform/qwen3-vl-plus | 260.096K | — | — | — | |||
| Qwen: Qwen Plus 0728qwen/qwen-plus-2025-07-28 | 1M | $0.26 | $0.78 | — | |||
| OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch | 1.04758M | $0.05 | $0.2 | — | |||
| qwen3.5-plusqwen_ai_platform/qwen3.5-plus | 991.808K | — | — | — | |||
| qwen3.7-maxqwen_ai_platform/qwen3.7-max | 991.808K | $2.5 | $7.5 | — | |||
| qwen3.7-plusqwen_ai_platform/qwen3.7-plus | 991.808K | — | — | — | |||
| qwen3.8-maxqwen_ai_platform/qwen3.8-max | 991.808K | $2 | $6 | — | |||
| databricks-deepseek-v4-flash-0731databricks/databricks-deepseek-v4-flash-0731 | 1M | $0.14 | $0.28 | — | |||
| databricks-deepseek-v4-pro-0813databricks/databricks-deepseek-v4-pro-0813 | 1M | $1.32 | $3.96 | — | |||
| zai-org/GLM-5.3-Flashfriendliai/zai-org/glm-5.3-flash | 1.04858M | $0.15 | $0.5 | — | |||
| zai-org/GLM-5.3friendliai/zai-org/glm-5.3 | 1.04858M | $1.26 | $3.96 | — | |||
| claude-fable-5-1vertex_ai-anthropic_models/claude-fable-5-1 | 1M | $10 | $50 | — | |||
| Google: Gemini 3.1 Pro Preview (batch)google/gemini-3.1-pro-preview:batch | 1.04858M | $1 | $6 | — | |||
| claude-fable-5-1@defaultvertex_ai-anthropic_models/claude-fable-5-1@default | 1M | $10 | $50 | — | |||
| glm-5.2zai/glm-5.2 | 1M | $1.4 | $4.4 | — | |||
| Qwen/Qwen3.8-Flashtogether_ai/qwen/qwen3.8-flash | 1M | $0.15 | $0.47 | — | |||
| OpenAI: GPT-5.6 Terra (batch)openai/gpt-5.6-terra:batch | 1.05M | $1 | $6 | — | |||
| MiniMax: MiniMax M3 (free)minimax/minimax-m3:free | 1.04858M | Free | Free | — | |||
| Nex AGI: Nex-N2.5-Pro (free)nex-agi/nex-n2.5-pro:free | 262.144K | Free | Free | — | |||
| MiniMax: MiniMax M3 (batch)minimax/minimax-m3:batch | 524.288K | $0.3 | $1.2 | — | |||
| OpenAI: o3 Mini (batch)openai/o3-mini:batch | 200K | $0.55 | $2.2 | — | |||
| OpenAI: GPT-5.4 (batch)openai/gpt-5.4:batch | 1.05M | $1.25 | $7.5 | — | |||
| Qwen: Qwen3.6 Plusqwen/qwen3.6-plus | 1M | $0.325 | $1.95 | — | |||
| DeepSeek: DeepSeek V4 Flash 0731 (batch)deepseek/deepseek-v4-flash-0731:batch | 1.04858M | $0.11 | $0.33 | — | |||
| Mistral: Devstral 2 2512mistralai/devstral-2512 | 262.144K | $0.4 | $2 | — | |||
| Anthropic: Claude Opus 4.8 (batch)anthropic/claude-opus-4.8:batch | 1M | $2.5 | $12.5 | — | |||
| Google: Gemini 3.8 Flash (batch)google/gemini-3.8-flash:batch | 1.04858M | $0.375 | $1.875 | — | |||
| OpenAI: GPT-5.4 Nano (batch)openai/gpt-5.4-nano:batch | 400K | $0.1 | $0.625 | — | |||
| OpenAI: GPT-4.1 Mini (batch)openai/gpt-4.1-mini:batch | 1.04758M | $0.2 | $0.8 | — | |||
| Claude Opus 5 (batch)anthropic/claude-opus-5:batch | 1M | $2.5 | $12.5 | — | |||
| Relace: Relace Apply 3relace/relace-apply-3 | 256K | $0.85 | $1.25 | — | |||
| Z.ai: GLM 5V Turboz-ai/glm-5v-turbo | 202.752K | $1.2 | $4 | — | |||
| Mistral: Ministral 3 8B 2512 (batch)mistralai/ministral-8b-2512:batch | 262.144K | $0.075 | $0.075 | — | |||
| Mistral: Mistral Large 3 2512mistralai/mistral-large-2512 | 262.144K | $0.5 | $1.5 | — | |||
| inclusionAI: Ling 3.0 Flashinclusionai/ling-3.0-flash | 262.144K | $0.021 | $0.063 | — | |||
| OpenAI: GPT-5 Pro (batch)openai/gpt-5-pro:batch | 400K | $7.5 | $60 | — | |||
| Anthropic: Claude Sonnet 4.5 (batch)anthropic/claude-sonnet-4.5:batch | 1M | $1.5 | $7.5 | — | |||
| Qwen: Qwen3 Maxqwen/qwen3-max | 262.144K | $0.78 | $3.9 | — | |||
| Anthropic: Claude Opus 4.1 (batch)anthropic/claude-opus-4.1:batch | 200K | $7.5 | $37.5 | — | |||
| Mistral: Codestral 2508mistralai/codestral-2508 | 256K | $0.3 | $0.9 | — | |||
| OpenAI: o4 Mini Highopenai/o4-mini-high | 200K | $1.1 | $4.4 | — | |||
| Anthropic: Claude Opus 4.7 (batch)anthropic/claude-opus-4.7:batch | 1M | $2.5 | $12.5 | — | |||
| gpt-chat-latestazure_ai/gpt-chat-latest | 272K | $5 | $30 | — | |||
| model-routerazure_ai/model-router | 200K | $0.14 | — | — | |||
| xai/grok-4.3vertex_ai/xai/grok-4.3 | 200K | $1.25 | $2.5 | — | |||
| grok-4-20-reasoningazure_ai/grok-4-20-reasoning | 262K | $1.25 | $2.5 | — | |||
| xai/grok-4.6vertex_ai/xai/grok-4.6 | 524.288K | $2 | $6 | — | |||
| grok-4-20-non-reasoningazure_ai/grok-4-20-non-reasoning | 262K | $1.25 | $2.5 | — |