76 models
Ranked by Design Arena: uicomponent
Reasoning Tools JSON 1394.0

Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It improves long-horizon agent collaboration, instruction following, and coding efficiency relative to Muse Spark 1.2.

meta/muse-spark-1.3 2026-09-02 1.04858M context $1.25/M input $4.25/M output
9 providers
Reasoning Tools 1360.0

Strongest Claude Opus model for coding, agents, and professional work

anthropic/claude-opus-5 2026-07-24 1M context $5/M input $25/M output
32 providers
1360.0

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

anthropic/claude-opus-5:batch 1M context $2.5/M input $12.5/M output
1340.0

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

google/gemini-3.8-flash:batch 1.04858M context $0.375/M input $1.875/M output
Reasoning Tools JSON 1340.0

Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows

google/gemini-3.8-flash 2026-09-02 1.04858M context $0.75/M input $3.75/M output
17 providers
1335.0

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

anthropic/claude-fable-5.1 1M context $10/M input $50/M output
1335.0

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

anthropic/claude-fable-5.1:batch 1M context $5/M input $25/M output
Reasoning Tools JSON 1332.0

Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1 with improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows.

meta/muse-spark-1.2 2026-08-05 1.04858M context $1.25/M input $4.25/M output
14 providers
Reasoning Tools 1331.0

Claude model for creative writing, analysis, and controlled agent workflows

anthropic/claude-fable-5 2026-06-09 1M context $10/M input $50/M output
37 providers
1331.0

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

anthropic/claude-fable-5:batch 1M context $5/M input $25/M output
1329.0

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

anthropic/claude-opus-4.7 1M context $5/M input $25/M output
1329.0

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

anthropic/claude-opus-4.7:batch 1M context $2.5/M input $12.5/M output
Reasoning Tools JSON 1317.0

Fast Gemini model balancing multimodal reasoning, tool use, and cost

google/gemini-3.6-flash 2026-07-21 1.04858M context $0.75/M input $3.75/M output
24 providers
1317.0

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

google/gemini-3.6-flash:batch 1.04858M context $0.375/M input $1.875/M output
1317.0

Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

x-ai/grok-4.6 500K context $2/M input $6/M output
Reasoning Tools JSON 1309.0

High-efficiency Gemini model for agentic workflows, coding, and multimodal reasoning

google/gemini-3.7-flash 2026-08-13 1.04858M context $0.75/M input $3.75/M output
23 providers
1309.0

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

google/gemini-3.7-flash:batch 1.04858M context $0.375/M input $1.875/M output
Reasoning 1306.0

Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.

meta/muse-spark-1.1 2026-04-08 1.04858M context $1.25/M input $4.25/M output
12 providers
1299.0

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

anthropic/claude-opus-4.6:batch 1M context $2.5/M input $12.5/M output
1299.0

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

anthropic/claude-opus-4.6 1M context $5/M input $25/M output
1296.0

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

anthropic/claude-sonnet-5:batch 1M context $1/M input $5/M output
Tools JSON 1296.0

Everyday Claude agent model for coding, planning, browsing, and general work

anthropic/claude-sonnet-5 2026-06-30 1M context $2/M input $10/M output
35 providers
Reasoning Tools JSON 1295.0

Reasoning-first Gemini preview for agentic coding and complex problem solving

google/gemini-3.1-pro-preview 2026-02-19 1.04858M context $2/M input $12/M output
31 providers
1295.0

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

google/gemini-3.1-pro-preview:batch 1.04858M context $1/M input $6/M output
1293.0

Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.

x-ai/grok-4.5 500K context $2/M input $6/M output
Reasoning Tools JSON 1287.0

Fast Gemini model balancing multimodal reasoning, tool use, and cost

google/gemini-3.5-flash 2026-05-19 1.04858M context $1.5/M input $9/M output
31 providers
1287.0

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

google/gemini-3.5-flash:batch 1.04858M context $0.75/M input $4.5/M output
1286.0

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...

anthropic/claude-sonnet-4.6:batch 1M context $1.5/M input $7.5/M output
1286.0

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...

anthropic/claude-sonnet-4.6 1M context $3/M input $15/M output
Reasoning Tools JSON 1271.0

Default frontier GPT for coding, computer use, research, and knowledge work

openai/gpt-5.5 2026-04-23 1.05M context $5/M input $30/M output
46 providers
1271.0

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

openai/gpt-5.5:batch 1.05M context $2.5/M input $15/M output
1271.0

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

anthropic/claude-opus-4.8 1M context $5/M input $25/M output
1271.0

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

anthropic/claude-opus-4.8:batch 1M context $2.5/M input $12.5/M output
1255.0

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...

openai/gpt-5.4:batch 1.05M context $1.25/M input $7.5/M output
Reasoning Tools JSON 1255.0

Agent-ready GPT for coding and computer-use workflows at a lower cost

openai/gpt-5.4 2026-03-05 1.05M context $2.5/M input $15/M output
42 providers
1253.0

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...

anthropic/claude-opus-4.5:batch 200K context $2.5/M input $12.5/M output
1253.0

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...

anthropic/claude-opus-4.5 200K context $5/M input $25/M output
1213.0

Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...

x-ai/grok-4.20 2M context $1.25/M input $2.5/M output
1206.0

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

x-ai/grok-4.3 1M context $1.25/M input $2.5/M output
1206.0

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

x-ai/grok-4.3:batch 1M context $1/M input $2/M output
1205.0

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly...

openai/gpt-5.2:batch 400K context $0.875/M input $7/M output
1194.0

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...

openai/gpt-5:batch 400K context $0.625/M input $5/M output
1187.0

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

anthropic/claude-sonnet-4.5:batch 1M context $1.5/M input $7.5/M output
1187.0

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

anthropic/claude-sonnet-4.5 1M context $3/M input $15/M output
1181.0

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...

openai/gpt-5.1:batch 400K context $0.625/M input $5/M output
1179.0

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...

anthropic/claude-opus-4.1 200K context $15/M input $75/M output
1179.0

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...

anthropic/claude-opus-4.1:batch 200K context $7.5/M input $37.5/M output
1168.0

Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in...

anthropic/claude-opus-4 200K context $15/M input $75/M output
Reasoning Tools JSON 1154.0

Google's proven reasoning model for coding, math, and multimodal analysis

google/gemini-2.5-pro 2025-06-17 1.04858M context $1.25/M input $10/M output
29 providers
1154.0

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

google/gemini-2.5-pro:batch 1.04858M context $0.625/M input $5/M output