85 models
Ranked by Intelligence Index
55.0

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

anthropic/claude-opus-4.7:batch 1M context $2.5/M input $12.5/M output
55.0

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

anthropic/claude-opus-4.7 1M context $5/M input $25/M output
53.4

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

anthropic/claude-fable-5.1:batch 1M context $5/M input $25/M output
53.4

Qwen3.8 Max (0803) is the August 3, 2026 checkpoint of Qwen3.8 Max, the flagship model in Alibaba's Qwen3.8 series and the general-availability successor to the Qwen3.8 Max Preview. It is...

qwen/qwen3.8-max 1M context $2/M input $6/M output
53.4

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

anthropic/claude-fable-5.1 1M context $10/M input $50/M output
53.1

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...

openai/gpt-5.4:batch 1.05M context $1.25/M input $7.5/M output
Reasoning Tools JSON 53.1

Agent-ready GPT for coding and computer-use workflows at a lower cost

openai/gpt-5.4 2026-03-05 1.05M context $2.5/M input $15/M output
42 providers
Reasoning Tools JSON 52.8

GPT-6 Astra is OpenAI's most capable model for complex reasoning, coding, computer use, research, and document creation.

openai/gpt-6-astra 2026-09-04 1.05M context $10/M input $50/M output
18 providers
52.8

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

openai/gpt-6-astra:batch 1.05M context $5/M input $25/M output
52.6

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

z-ai/glm-5.2:batch 1.04858M context $0.7/M input $2.2/M output
50.7

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

anthropic/claude-opus-5:batch 1M context $2.5/M input $12.5/M output
Reasoning Tools 50.7

Strongest Claude Opus model for coding, agents, and professional work

anthropic/claude-opus-5 2026-07-24 1M context $5/M input $25/M output
32 providers
Reasoning Tools 49.7

Claude model for creative writing, analysis, and controlled agent workflows

anthropic/claude-fable-5 2026-06-09 1M context $10/M input $50/M output
37 providers
49.7

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

anthropic/claude-fable-5:batch 1M context $5/M input $25/M output
47.1

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

openai/gpt-5.6-sol:batch 1.05M context $1/M input $5/M output
Reasoning Tools JSON 47.1

Frontier GPT-5.6 model for complex professional work, coding, and agentic workflows

openai/gpt-5.6-sol 2026-07-09 1.05M context $4/M input $20/M output
36 providers
44.9

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

z-ai/glm-5.3:batch 1.04858M context $0.7/M input $2.2/M output
44.9

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

z-ai/glm-5.3 1.04858M context $1.4/M input $4.4/M output
Reasoning Tools JSON Open weights 43.8

Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work

moonshotai/kimi-k3 2026-07-16 1.04858M context $3/M input $15/M output
63 providers
43.8

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

moonshotai/kimi-k3:batch 1.04858M context $3/M input $15/M output
42.3

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

openai/gpt-5.6-terra:batch 1.05M context $1/M input $6/M output
Reasoning Tools JSON 42.3

Balanced GPT-5.6 model for capable, cost-efficient everyday work

openai/gpt-5.6-terra 2026-07-09 1.05M context $2/M input $12/M output
35 providers
42.0

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

anthropic/claude-opus-4.8:batch 1M context $2.5/M input $12.5/M output
42.0

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

anthropic/claude-opus-4.8 1M context $5/M input $25/M output
41.9

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

z-ai/glm-5.3-flash 1.04858M context $0.15/M input $0.5/M output
41.9

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

z-ai/glm-5.3-flash:batch 1.04858M context $0.075/M input $0.25/M output
41.2

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

google/gemini-3.8-flash:batch 1.04858M context $0.375/M input $1.875/M output
Reasoning Tools JSON 41.2

Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows

google/gemini-3.8-flash 2026-09-02 1.04858M context $0.75/M input $3.75/M output
17 providers
40.5

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

qwen/qwen3.6-plus 1M context $0.325/M input $1.95/M output
40.3

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...

qwen/qwen3.8-max-0902 1M context $2/M input $6/M output
40.0

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

qwen/qwen3.8-2.4t-a95b 1M context $2/M input $6/M output
40.0

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

qwen/qwen3.8-2.4t-a95b:batch 1.01M context $2/M input $6/M output
Reasoning Tools JSON 39.8

Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1 with improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows.

meta/muse-spark-1.2 2026-08-05 1.04858M context $1.25/M input $4.25/M output
14 providers
Reasoning Tools JSON Open weights 39.5

DeepSeek V4.1 Flash model for reasoning and agentic coding

deepseek/deepseek-v4.1-flash 2026-09-10 1M context $0.15/M input $0.6/M output
21 providers
Reasoning Tools JSON 39.4

High-efficiency Gemini model for agentic workflows, coding, and multimodal reasoning

google/gemini-3.7-flash 2026-08-13 1.04858M context $0.75/M input $3.75/M output
23 providers
39.4

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

google/gemini-3.7-flash:batch 1.04858M context $0.375/M input $1.875/M output
Reasoning Tools JSON 38.6

Default frontier GPT for coding, computer use, research, and knowledge work

openai/gpt-5.5 2026-04-23 1.05M context $5/M input $30/M output
46 providers
38.6

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

openai/gpt-5.5:batch 1.05M context $2.5/M input $15/M output
Tools JSON 38.4

Everyday Claude agent model for coding, planning, browsing, and general work

anthropic/claude-sonnet-5 2026-06-30 1M context $2/M input $10/M output
35 providers
38.4

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

anthropic/claude-sonnet-5:batch 1M context $1/M input $5/M output
37.5

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

openai/gpt-5.6-luna:batch 1.05M context $0.1/M input $0.6/M output
Reasoning Tools JSON 37.5

Cost-efficient GPT-5.6 model for fast, high-volume workloads

openai/gpt-5.6-luna 2026-07-09 1.05M context $0.2/M input $1.2/M output
36 providers
Reasoning Tools JSON Open weights 36.3

DeepSeek V4 Pro snapshot with million-token context and support for thinking and non-thinking modes

deepseek/deepseek-v4-pro-0813 2026-08-12 1M context $0.442/M input $0.884/M output
30 providers
36.3

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

deepseek/deepseek-v4-pro-0813:batch 1.04858M context $0.66/M input $1.98/M output
35.7

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

minimax/minimax-m3:free 1.04858M context Free input Free output
Reasoning Tools JSON Open weights 34.5

Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding

deepseek/deepseek-v4-flash-0731 2026-07-31 1M context $0.05/M input $0.16/M output
43 providers
34.5

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

deepseek/deepseek-v4-flash-0731:batch 1.04858M context $0.11/M input $0.33/M output
Reasoning Tools JSON 34.3

Fast Gemini model balancing multimodal reasoning, tool use, and cost

google/gemini-3.6-flash 2026-07-21 1.04858M context $0.75/M input $3.75/M output
24 providers
Reasoning 34.3

Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.

meta/muse-spark-1.1 2026-04-08 1.04858M context $1.25/M input $4.25/M output
12 providers
34.3

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

google/gemini-3.6-flash:batch 1.04858M context $0.375/M input $1.875/M output