May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...
Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly...
GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...
GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks....
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It...
MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| DeepSeek: R1 0528deepseek/deepseek-r1-0528 | 163.84K | $0.5 | $2.15 | — | |||
| Anthropic: Claude Opus 4.5anthropic/claude-opus-4.5 | 200K | $5 | $25 | — | |||
| Thinking Machines: Inkling Small (batch)thinkingmachines/inkling-small:batch | 524.288K | $0.5 | $1.2 | — | |||
| OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch | 1.04758M | $0.05 | $0.2 | — | |||
| OpenAI: gpt-oss-20b (free)openai/gpt-oss-20b:free | 131.072K | Free | Free | — | |||
| OpenAI: GPT-5.5 (batch)openai/gpt-5.5:batch | 1.05M | $2.5 | $15 | — | |||
| Anthropic: Claude Opus 4.7 (batch)anthropic/claude-opus-4.7:batch | 1M | $2.5 | $12.5 | — | |||
| SpaceXAI: Grok Build 0.1x-ai/grok-build-0.1 | 256K | $1 | $2 | — | |||
| Anthropic: Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | 1M | $3 | $15 | — | |||
| OpenAI: GPT-5.6 Sol (batch)openai/gpt-5.6-sol:batch | 1.05M | $1 | $5 | — | |||
| OpenAI: GPT-5 (batch)openai/gpt-5:batch | 400K | $0.625 | $5 | — | |||
| Meta: Llama 4 Maverickmeta-llama/llama-4-maverick | 128K | $0.2 | $0.696 | — | |||
| Z.ai: GLM 4.6z-ai/glm-4.6 | 198K | $0.43 | $1.75 | — | |||
| Qwen: Qwen3.8 2.4T A95B (batch)qwen/qwen3.8-2.4t-a95b:batch | 1.01M | $2 | $6 | — | |||
| Qwen: Qwen3 32Bqwen/qwen3-32b | 40.96K | $0.08 | $0.28 | — | |||
| OpenAI: GPT-5.2 (batch)openai/gpt-5.2:batch | 400K | $0.875 | $7 | — | |||
| Z.ai: GLM 4.7z-ai/glm-4.7 | 202.752K | $0.4 | $1.75 | — | |||
| Mistral: Ministral 3 8B 2512 (batch)mistralai/ministral-8b-2512:batch | 262.144K | $0.075 | $0.075 | — | |||
| Z.ai: GLM 5.3 Flashz-ai/glm-5.3-flash | 1.04858M | $0.075 | $0.25 | — | |||
| Qwen: Qwen3.8 2.4T A95Bqwen/qwen3.8-2.4t-a95b | 1M | $2 | $6 | — | |||
| OpenAI: GPT-5 Codex (batch)openai/gpt-5-codex:batch | 400K | $0.625 | $5 | — | |||
| Google: Gemini 3.5 Flash (batch)google/gemini-3.5-flash:batch | 1.04858M | $0.75 | $4.5 | — | |||
| OpenAI: GPT-4o-mini (batch)openai/gpt-4o-mini:batch | 128K | $0.075 | $0.3 | — | |||
| Qwen: Qwen3 Maxqwen/qwen3-max | 262.144K | $0.78 | $3.9 | — | |||
| MiniMax: MiniMax M2.5minimax/minimax-m2.5 | 200K | $0.27 | $1.08 | — | |||
| Qwen: Qwen3.6 Plusqwen/qwen3.6-plus | 1M | $0.325 | $1.95 | — |