22 models
Ranked by Aider Polyglot
Tools 88.0

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...

openai/gpt-5 2025-08-07 400K context $1.25/M input $10/M output
30 providers
Reasoning Tools 84.9

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to think harder and provide consistently...

openai/o3-pro 2025-06-10 200K context $20/M input $80/M output
8 providers
Reasoning Tools JSON 83.1

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

google/gemini-2.5-pro 2025-06-17 1.04858M context $1.25/M input $10/M output
29 providers
Reasoning 81.3

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....

openai/o3 2025-04-16 200K context $2/M input $8/M output
20 providers
Open weights 74.2

DeepSeek reasoning model for multi-step analysis, math, coding, and tools

deepseek/deepseek-reasoner 2025-12-01 1M context $0.147/M input $0.295/M output
6 providers
Reasoning Tools 72.0

Flagship Claude model for deep reasoning, coding, and long-horizon agents

anthropic/claude-opus-4-20250514 2025-05-22 200K context $15/M input $75/M output
8 providers
72.0

Flagship Claude model for deep reasoning, coding, and long-horizon agents

anthropic/claude-opus-4-0 2025-05-22 200K context Input not listed Output not listed
2 providers
Reasoning Tools JSON 72.0

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning...

openai/o4-mini 2025-04-16 200K context $1.1/M input $4.4/M output
20 providers
Reasoning Tools 64.9

Balanced Claude model for coding, analysis, agent workflows, and cost control

anthropic/claude-3-7-sonnet-20250219 2025-02-19 200K context $3/M input $15/M output
4 providers
Reasoning 61.7

The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason...

openai/o1 2024-12-05 200K context $15/M input $60/M output
16 providers
Reasoning Tools JSON 61.3

Balanced Claude model for coding, analysis, agent workflows, and cost control

anthropic/claude-sonnet-4-0 2025-05-22 200K context $2.898/M input $14.493/M output
4 providers
Reasoning Tools 61.3

Balanced Claude model for coding, analysis, agent workflows, and cost control

anthropic/claude-sonnet-4-20250514 2025-05-22 200K context $3/M input $15/M output
11 providers
Reasoning Tools JSON 60.4

OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model supports the `reasoning_effort` parameter, which can be set to...

openai/o3-mini 2024-12-20 200K context $1.1/M input $4.4/M output
19 providers
Reasoning Tools JSON 55.1

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

google/gemini-2.5-flash 2025-06-17 1.04858M context $0.3/M input $2.5/M output
30 providers
Tools JSON 52.4

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...

openai/gpt-4.1 2025-04-14 1.04758M context $2/M input $8/M output
28 providers
51.6

Balanced Claude model for coding, analysis, agent workflows, and cost control

anthropic/claude-3-5-sonnet-20241022 2024-10-22 200K context Input not listed Output not listed
1 provider
32.4

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...

openai/gpt-4.1-mini 2025-04-14 1.04758M context $0.4/M input $1.6/M output
25 providers
Tools 28.0

Fast Claude model for responsive assistance, classification, and lightweight agents

anthropic/claude-3-5-haiku-20241022 2024-10-22 200K context $0.8/M input $4/M output
3 providers
Tools Open weights 15.6

Open multimodal Llama for strong reasoning with efficient everyday serving

meta/llama-4-maverick-17b-instruct 2025-04-05 1M context $0.14/M input $0.59/M output
6 providers
Reasoning Tools Open weights 12.0

Cohere command model for multilingual enterprise agents, tools, and chat

cohere/command-a-03-2025 2025-03-13 256K context $2.5/M input $10/M output
6 providers
Tools Open weights 11.1

Mistral code model for completions, refactors, and developer IDE workflows

mistral/codestral-latest 2024-05-29 256K context $0.3/M input $0.9/M output
5 providers
8.9

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...

openai/gpt-4.1-nano 2025-04-14 1.04758M context $0.1/M input $0.4/M output
20 providers