116 models
Ranked by Design Arena: svg
Reasoning Tools JSON Open weights 1058.0

Hybrid-reasoning DeepSeek model with thinking and non-thinking modes, sparse attention, and tool-use

deepseek/deepseek-v3.2 2025-12-01 128K context $0.18/M input $0.35/M output
24 providers
1057.0

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

deepseek/deepseek-v3.2-exp 163.84K context $0.27/M input $0.41/M output
1051.0

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

anthropic/claude-haiku-4.5:batch 200K context $0.5/M input $2.5/M output
1051.0

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

anthropic/claude-haiku-4.5 200K context $1/M input $5/M output
Reasoning Tools JSON 1045.0

Fast Gemini workhorse for multimodal apps where latency and price matter

google/gemini-2.5-flash 2025-06-17 1.04858M context $0.3/M input $2.5/M output
30 providers
1045.0

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

google/gemini-2.5-flash:batch 1.04858M context $0.15/M input $1.25/M output
Reasoning Tools Open weights 1041.0

Reasoning-optimized 398B MoE agent model with extended thinking for long-horizon and multi-turn tool use

arcee-ai/trinity-large-thinking 2026-04-01 524.288K context $0.25/M input $0.9/M output
5 providers
1036.0

Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It...

qwen/qwen3-max 262.144K context $0.78/M input $3.9/M output
1018.0

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

mistralai/mistral-large-2512:batch 262.144K context $0.25/M input $0.75/M output
1018.0

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

mistralai/mistral-large-2512 262.144K context $0.5/M input $1.5/M output
1018.0

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...

mistralai/mistral-medium-3.1 131.072K context $0.4/M input $2/M output
1018.0

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...

mistralai/mistral-medium-3.1:batch 131.072K context $0.2/M input $1/M output
1012.0

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving...

inception/mercury-2 128K context $0.25/M input $0.75/M output
Reasoning 1003.0

Coding-optimized GPT model for repository edits, reviews, and agentic software work

openai/gpt-5.1-codex-mini 2025-11-13 400K context $0.22/M input $1.8/M output
15 providers
Tools Open weights 1002.0

DeepSeek chat model for instruction following, coding, and analysis

deepseek/deepseek-chat 2025-12-01 1M context $0.147/M input $0.295/M output
8 providers
992.0

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...

deepseek/deepseek-chat-v3.1 163.84K context $0.25/M input $0.95/M output