GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2. It achieves state-of-the-art results...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,...
Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks. Launch video: https://youtu.be/Gc82AXLa0Rg?si=4RLn6WBz33qT--B7...
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It...
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...
Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32...
GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex
GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...
GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...
Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across...
The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...
GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...
GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...
o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....
o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....
Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)
Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning...
OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning...
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| OpenAI: GPT-5.3-Codexopenai/gpt-5.3-codex | 1160.0 | 400K | $1.75 | $14 | 2026-02-05 | |||
| Qwen: Qwen3 Coder 480B A35Bqwen/qwen3-coder | 1158.0 | 262.144K | $0.3 | $1 | — | |||
| Mistral: Mistral Large 3 2512mistralai/mistral-large-2512 | 1157.0 | 262.144K | $0.5 | $1.5 | — | |||
| Mistral: Mistral Large 3 2512 (batch)mistralai/mistral-large-2512:batch | 1157.0 | 262.144K | $0.25 | $0.75 | — | |||
| Anthropic: Claude Sonnet 4anthropic/claude-sonnet-4 | 1156.0 | 200K | $3 | $15 | — | |||
| NVIDIA: Nemotron 3 Ultra (batch)nvidia/nemotron-3-ultra-550b-a55b:batch | 1156.0 | 512.288K | $0.6 | $3.6 | — | |||
| NVIDIA: Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b | 1154.0 | 256K | $0.625 | $3.125 | 2026-06-04 | |||
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | 1154.0 | 1M | Free | Free | — | |||
| MiniMax: MiniMax M2minimax/minimax-m2 | 1150.0 | 204.8K | $0.255 | $1.02 | — | |||
| Arcee AI: Trinity Large Thinkingarcee-ai/trinity-large-thinking | 1132.0 | 262.144K | $0.25 | $0.8 | 2026-04-01 | |||
| Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5 | 1130.0 | 200K | $1 | $5 | — | |||
| OpenAI: GPT-5 Miniopenai/gpt-5-mini | 1130.0 | 400K | $0.25 | $2 | 2025-08-07 | |||
| OpenAI: GPT-5 Mini (batch)openai/gpt-5-mini:batch | 1130.0 | 400K | $0.125 | $1 | — | |||
| Anthropic: Claude Haiku 4.5 (batch)anthropic/claude-haiku-4.5:batch | 1130.0 | 200K | $0.5 | $2.5 | — | |||
| Qwen: Qwen3 Maxqwen/qwen3-max | 1125.0 | 262.144K | $0.78 | $3.9 | — | |||
| Google: Gemini 2.5 Flash (batch)google/gemini-2.5-flash:batch | 1119.0 | 1.04858M | $0.15 | $1.25 | — | |||
| Google: Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1119.0 | 1.04858M | $0.3 | $2.5 | 2025-06-17 | |||
| MoonshotAI: Kimi K2 0905moonshotai/kimi-k2-0905 | 1113.0 | 262.144K | $0.6 | $2.5 | — | |||
| OpenAI: GPT-5.1-Codex-Miniopenai/gpt-5.1-codex-mini | 1109.0 | 400K | $0.25 | $2 | 2025-11-13 | |||
| OpenAI: GPT-5 Nanoopenai/gpt-5-nano | 1099.0 | 400K | $0.05 | $0.4 | 2025-08-07 | |||
| OpenAI: GPT-5 Nano (batch)openai/gpt-5-nano:batch | 1099.0 | 400K | $0.025 | $0.2 | — | |||
| Google: Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 1087.0 | 1.04858M | $0.25 | $1.5 | 2026-03-03 | |||
| Mistral: Ministral 3 14B 2512mistralai/ministral-14b-2512 | 1080.0 | 262.144K | $0.2 | $0.2 | — | |||
| Mistral: Ministral 3 8B 2512 (batch)mistralai/ministral-8b-2512:batch | 1068.0 | 262.144K | $0.075 | $0.075 | — | |||
| Mistral: Ministral 3 8B 2512mistralai/ministral-8b-2512 | 1068.0 | 262.144K | $0.15 | $0.15 | — | |||
| Qwen: Qwen3 235B A22B Instruct 2507qwen/qwen3-235b-a22b-2507 | 1053.0 | 262.144K | $0.087 | $0.35 | — | |||
| OpenAI: GPT-4.1 (batch)openai/gpt-4.1:batch | 1040.0 | 1.04758M | $1 | $4 | — | |||
| OpenAI: GPT-4.1openai/gpt-4.1 | 1040.0 | 1.04758M | $2 | $8 | 2025-04-14 | |||
| OpenAI: o3openai/o3 | 1035.0 | 200K | $2 | $8 | 2025-04-16 | |||
| OpenAI: o3 (batch)openai/o3:batch | 1035.0 | 200K | $1 | $4 | — | |||
| Mistral: Codestral 2508 (batch)mistralai/codestral-2508:batch | 1022.0 | 256K | $0.15 | $0.45 | — | |||
| Mistral: Codestral 2508mistralai/codestral-2508 | 1022.0 | 256K | $0.3 | $0.9 | — | |||
| OpenAI: GPT-4.1 Mini (batch)openai/gpt-4.1-mini:batch | 1008.0 | 1.04758M | $0.2 | $0.8 | — | |||
| OpenAI: GPT-4.1 Miniopenai/gpt-4.1-mini | 1008.0 | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| OpenAI: o4 Miniopenai/o4-mini | 991.0 | 200K | $1.1 | $4.4 | 2025-04-16 | |||
| OpenAI: o4 Mini (batch)openai/o4-mini:batch | 991.0 | 200K | $0.55 | $2.2 | — | |||
| OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch | 978.0 | 1.04758M | $0.05 | $0.2 | — | |||
| OpenAI: GPT-4.1 Nanoopenai/gpt-4.1-nano | 978.0 | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| Meta: Llama 4 Scoutmeta-llama/llama-4-scout | 805.0 | 327.68K | $0.1 | $0.3 | — |