156 models
Ranked by Design Arena: website
1260.0

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading...

z-ai/glm-5 198K context $0.6/M input $1.92/M output
1260.0

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...

minimax/minimax-m2.7:free 196.608K context Free input Free output
1259.0

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...

anthropic/claude-opus-4.5:batch 200K context $2.5/M input $12.5/M output
1259.0

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...

anthropic/claude-opus-4.5 200K context $5/M input $25/M output
1258.0

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...

minimax/minimax-m2.7 204.8K context $0.3/M input $1.2/M output
1253.0

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

qwen/qwen3.6-plus 1M context $0.325/M input $1.95/M output
1251.0

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

deepseek/deepseek-v4-flash-0731:batch 1.04858M context $0.11/M input $0.33/M output
1243.0

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

z-ai/glm-5v-turbo 202.752K context $1.2/M input $4/M output
1242.0

Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...

x-ai/grok-4.20 2M context $1.25/M input $2.5/M output
1238.0

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...

z-ai/glm-4.7 202.752K context $0.4/M input $1.75/M output
1236.0

Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts text and image input and produces...

nex-agi/nex-n2-pro 262.144K context $0.25/M input $1/M output
1235.0

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...

minimax/minimax-m2.5 200K context $0.27/M input $1.08/M output
1232.0

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...

openai/gpt-5.4:batch 1.05M context $1.25/M input $7.5/M output
Reasoning Tools JSON 1232.0

Agent-ready GPT for coding and computer-use workflows at a lower cost

openai/gpt-5.4 2026-03-05 1.05M context $2.5/M input $15/M output
42 providers
1230.0

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...

tencent/hy3:free 262.144K context Free input Free output
1226.0

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

thinkingmachines/inkling:free 1.04858M context Free input Free output
1226.0

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

thinkingmachines/inkling:batch 524.288K context $1/M input $4.05/M output
1214.0

MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world...

minimax/minimax-m2.1 204.8K context $0.3/M input $1.2/M output
Reasoning Tools JSON 1208.0

New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs

google/gemini-3-flash-preview 2025-12-17 1.04858M context $0.5/M input $3/M output
25 providers
1208.0

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...

google/gemini-3-flash-preview:batch 1.04858M context $0.25/M input $1.5/M output
1207.0

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly...

openai/gpt-5.2:batch 400K context $0.875/M input $7/M output
Reasoning Tools 1207.0

Reliable GPT generation for broad coding, writing, and tool-assisted product work

openai/gpt-5.2 2025-12-11 400K context $1.75/M input $14/M output
33 providers
1206.0

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

x-ai/grok-4.3 1M context $1.25/M input $2.5/M output
1206.0

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...

z-ai/glm-4.7-flash 131.072K context $0.061/M input $0.4/M output
1206.0

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

x-ai/grok-4.3:batch 1M context $1/M input $2/M output
1203.0

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...

qwen/qwen3.5-397b-a17b 262.144K context $0.55/M input $3.5/M output
1202.0

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

anthropic/claude-sonnet-4.5:batch 1M context $1.5/M input $7.5/M output
1202.0

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

anthropic/claude-sonnet-4.5 1M context $3/M input $15/M output
1200.0

The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...

qwen/qwen3.5-plus-02-15 1M context $0.26/M input $1.56/M output
Reasoning Tools JSON 1199.0

Sharper GPT-5 generation for coding, product work, and tool-assisted tasks

openai/gpt-5.1 2025-11-13 400K context $1.25/M input $10/M output
28 providers
1199.0

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...

openai/gpt-5.1:batch 400K context $0.625/M input $5/M output
1199.0

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...

deepseek/deepseek-v3.1-terminus 131.072K context $0.27/M input $1/M output
1197.0

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...

openai/gpt-5:batch 400K context $0.625/M input $5/M output
Tools 1197.0

Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows

openai/gpt-5 2025-08-07 400K context $1.25/M input $10/M output
30 providers
1191.0

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

deepseek/deepseek-v3.2-exp 163.84K context $0.27/M input $0.41/M output
1189.0

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...

anthropic/claude-opus-4.1 200K context $15/M input $75/M output
1189.0

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...

anthropic/claude-opus-4.1:batch 200K context $7.5/M input $37.5/M output
1189.0

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

qwen/qwen3-coder:free 262K context Free input Free output
Tools JSON 1188.0

Upstage's flagship model, specialized for agentic use

upstage/solar-pro4 2026-08-06 524.288K context $0.3/M input $1.2/M output
5 providers
1187.0

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...

z-ai/glm-4.6 198K context $0.43/M input $1.75/M output
1182.0

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...

z-ai/glm-4.5 131.072K context $0.6/M input $2.2/M output
Reasoning Tools JSON 1179.0

Google's proven reasoning model for coding, math, and multimodal analysis

google/gemini-2.5-pro 2025-06-17 1.04858M context $1.25/M input $10/M output
29 providers
1179.0

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

google/gemini-2.5-pro:batch 1.04858M context $0.625/M input $5/M output
1177.0

Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in...

anthropic/claude-opus-4 200K context $15/M input $75/M output
Reasoning Tools JSON 1175.0

Coding-optimized GPT model for repository edits, reviews, and agentic software work

openai/gpt-5.3-codex 2026-02-05 400K context $1.75/M input $14/M output
28 providers
1175.0

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

mistralai/mistral-large-2512:batch 262.144K context $0.25/M input $0.75/M output
1175.0

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

mistralai/mistral-large-2512 262.144K context $0.5/M input $1.5/M output
Reasoning 1174.0

Codex GPT for repository edits, code review, and practical software agents

openai/gpt-5.1-codex 2025-11-13 400K context $1.07/M input $8.5/M output
19 providers
1171.0

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

qwen/qwen3-coder 262.144K context $0.3/M input $1/M output
1161.0

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...

deepseek/deepseek-r1-0528 163.84K context $0.5/M input $2.15/M output