Claude Opus 5 vs GPT-5.6 Terra vs Gemini 3.6 Flash
Compare Claude Opus 5 and GPT-5.6 Terra and Gemini 3.6 Flash on key metrics including pricing, context length, capabilities, providers, and published benchmarks.
Overview
Author
Anthropic
Context length
1Mtokens
Reasoning
Supported
Input modalities
Output modalities
Providers
34 providers
Author
OpenAI
Context length
1.05MtokensBest
Reasoning
Supported
Input modalities
Output modalities
Providers
36 providers
Author
Google
Context length
1.04858Mtokens
Reasoning
Supported
Input modalities
Output modalities
Providers
25 providers
Pricing per million tokens
Input
$5
Output
$25
Cached input
$0.5
Cache write
$6.25
Input
$2
Output
$12
Cached input
$0.2
Cache write
$2.5
Input
$0.75Best
Output
$3.75Best
Cached input
$0.075
Cache write
Not listed
Limits
Context length
1Mtokens
Max output
128KtokensBest
Released
2026-07-24
Knowledge cutoff
2026-05
Context length
1.05MtokensBest
Max output
128KtokensBest
Released
2026-07-09
Knowledge cutoff
2026-02-16
Context length
1.04858Mtokens
Max output
65.536Ktokens
Released
2026-07-21
Knowledge cutoff
2026-03
Capabilities
Tool calling
Supported
Structured output
Not reported
Vision input
Supported
Attachments
Supported
Tool calling
Supported
Structured output
Supported
Vision input
Supported
Attachments
Supported
Tool calling
Supported
Structured output
Supported
Vision input
Supported
Attachments
Supported
Availability
Inference providers
34providers
Provider list
302aiAnthropicabacusagentrouteraihubmix+27 more
Inference providers
36providersBest
Provider list
302aiOpenAIabacusai-routeraihubmix+30 more
Inference providers
25providers
Provider list
302aiGoogleabacuscortecscrossmodel+18 more
Benchmarks
- Claude Opus 5
- GPT-5.6 Terra
- Gemini 3.6 Flash
- Claude Opus 5
- GPT-5.6 Terra
- Gemini 3.6 Flash
- Claude Opus 5
- GPT-5.6 Terra
- Gemini 3.6 Flash
All published benchmark results 50 variants
AA-Briefcase
Elo1720.0Best
ARC-AGI-1
accuracy97.5Best
ARC-AGI-2
accuracy90.4Best
ARC-AGI-3
RHAE30.2Best
Agentic Index
index56.2Best
Agents' Last Exam
scoreNot reported
Artificial Analysis Coding Agent Index
index score · 1.1 · CodexNot reported
Artificial Analysis Intelligence Index
index score · 4.1Not reported
AutomationBench
success rate26.0Best
BrowseComp
accuracy90.8Best
CharXiv Reasoning
accuracyNot reported
Coding Index
index78.0Best
DeepSWE
resolve rate · 1.168.8
DeepSearchQA
F195.0Best
Design Arena: 3d
elo1364.0Best
Design Arena: agenticgamedev
elo1265.0Best
Design Arena: androidnative
elo1268.0Best
Design Arena: asciiart
elo1386.0Best
Design Arena: codecategories
elo1339.0Best
Design Arena: dataviz
elo1355.0Best
Design Arena: fullstack
elo1331.0Best
Design Arena: gamedev
elo1366.0Best
Design Arena: htmlslides
eloNot reported
Design Arena: mobileapps
elo1348.0Best
Design Arena: python-pptxslides
elo1265.0Best
Design Arena: svg
elo1351.0Best
Design Arena: uicomponent
elo1360.0Best
Design Arena: webapps
elo1277.0Best
Design Arena: website
elo1320.0Best
Frontier-Bench
mean reward · v0.1 · mini-SWE-agent43.3Best
FrontierCode
mean@5 · 1.153.4Best
FrontierMath
accuracy · v2Not reported
GDM-MRCR
accuracy · v2Not reported
GDPval-AA
Elo · v21861.0Best
GPQA Diamond
accuracyNot reported
HealthBench Professional
score59.8Best
Humanity's Last Exam
accuracy64.7Best
Intelligence Index
index50.7Best
MLE-Bench
average position scoreNot reported
MMMU Pro
accuracyNot reported
OSWorld
success rate · 2.070.6Best
OSWorld-Verified
success rateNot reported
SWE-Bench Multilingual
resolve rate89.5Best
SWE-Bench Multimodal
resolve rate59.4Best
SWE-Bench Pro
resolve rate79.2Best
SWE-Bench Pro
resolve rate · AntigravityNot reported
SWE-Bench Verified
resolved96.0Best
Terminal-Bench
success rate · 2.1Not reported
Terminal-Bench
accuracy · 2.1 · Terminus 2Not reported
Toolathlon
success rateNot reported
AA-Briefcase
EloNot reported
ARC-AGI-1
accuracyNot reported
ARC-AGI-2
accuracyNot reported
ARC-AGI-3
RHAENot reported
Agentic Index
index43.7
Agents' Last Exam
score50.4Best
Artificial Analysis Coding Agent Index
index score · 1.1 · Codex77.4Best
Artificial Analysis Intelligence Index
index score · 4.155.0Best
AutomationBench
success rateNot reported
BrowseComp
accuracy87.5
CharXiv Reasoning
accuracyNot reported
Coding Index
index76.7
DeepSWE
resolve rate · 1.169.6Best
DeepSearchQA
F1Not reported
Design Arena: 3d
eloNot reported
Design Arena: agenticgamedev
eloNot reported
Design Arena: androidnative
eloNot reported
Design Arena: asciiart
eloNot reported
Design Arena: codecategories
eloNot reported
Design Arena: dataviz
eloNot reported
Design Arena: fullstack
eloNot reported
Design Arena: gamedev
eloNot reported
Design Arena: htmlslides
eloNot reported
Design Arena: mobileapps
eloNot reported
Design Arena: python-pptxslides
eloNot reported
Design Arena: svg
eloNot reported
Design Arena: uicomponent
eloNot reported
Design Arena: webapps
eloNot reported
Design Arena: website
eloNot reported
Frontier-Bench
mean reward · v0.1 · mini-SWE-agentNot reported
FrontierCode
mean@5 · 1.1Not reported
FrontierMath
accuracy · v284.9Best
GDM-MRCR
accuracy · v2Not reported
GDPval-AA
Elo · v2Not reported
GPQA Diamond
accuracy92.9Best
HealthBench Professional
scoreNot reported
Humanity's Last Exam
accuracyNot reported
Intelligence Index
index42.3
MLE-Bench
average position scoreNot reported
MMMU Pro
accuracy80.7Best
OSWorld
success rate · 2.050.2
OSWorld-Verified
success rateNot reported
SWE-Bench Multilingual
resolve rateNot reported
SWE-Bench Multimodal
resolve rateNot reported
SWE-Bench Pro
resolve rate63.4
SWE-Bench Pro
resolve rate · AntigravityNot reported
SWE-Bench Verified
resolvedNot reported
Terminal-Bench
success rate · 2.187.4Best
Terminal-Bench
accuracy · 2.1 · Terminus 2Not reported
Toolathlon
success rate53.1Best
AA-Briefcase
EloNot reported
ARC-AGI-1
accuracyNot reported
ARC-AGI-2
accuracyNot reported
ARC-AGI-3
RHAENot reported
Agentic Index
index30.2
Agents' Last Exam
scoreNot reported
Artificial Analysis Coding Agent Index
index score · 1.1 · CodexNot reported
Artificial Analysis Intelligence Index
index score · 4.1Not reported
AutomationBench
success rateNot reported
BrowseComp
accuracyNot reported
CharXiv Reasoning
accuracy89.4Best
Coding Index
index69.2
DeepSWE
resolve rate · 1.149.0
DeepSearchQA
F1Not reported
Design Arena: 3d
elo1302.0
Design Arena: agenticgamedev
elo1181.0
Design Arena: androidnative
elo1215.0
Design Arena: asciiart
elo1291.0
Design Arena: codecategories
elo1305.0
Design Arena: dataviz
elo1312.0
Design Arena: fullstack
elo1195.0
Design Arena: gamedev
elo1284.0
Design Arena: htmlslides
elo1150.0Best
Design Arena: mobileapps
elo1232.0
Design Arena: python-pptxslides
elo1146.0
Design Arena: svg
eloNot reported
Design Arena: uicomponent
elo1317.0
Design Arena: webapps
elo1217.0
Design Arena: website
elo1312.0
Frontier-Bench
mean reward · v0.1 · mini-SWE-agentNot reported
FrontierCode
mean@5 · 1.1Not reported
FrontierMath
accuracy · v2Not reported
GDM-MRCR
accuracy · v254.0Best
GDPval-AA
Elo · v21421.0
GPQA Diamond
accuracyNot reported
HealthBench Professional
scoreNot reported
Humanity's Last Exam
accuracyNot reported
Intelligence Index
index34.3
MLE-Bench
average position score63.9Best
MMMU Pro
accuracyNot reported
OSWorld
success rate · 2.0Not reported
OSWorld-Verified
success rate83.0Best
SWE-Bench Multilingual
resolve rateNot reported
SWE-Bench Multimodal
resolve rateNot reported
SWE-Bench Pro
resolve rateNot reported
SWE-Bench Pro
resolve rate · Antigravity58.7Best
SWE-Bench Verified
resolvedNot reported
Terminal-Bench
success rate · 2.1Not reported
Terminal-Bench
accuracy · 2.1 · Terminus 278.0Best
Toolathlon
success rateNot reported
Provenance
“Not reported” means the source catalog is silent on that field. Benchmark rows keep each published version and harness separate.