Claude Opus 5 vs GPT-5.6 Terra vs Gemini 3.6 Flash

Compare Claude Opus 5 and GPT-5.6 Terra and Gemini 3.6 Flash on key metrics including pricing, context length, capabilities, providers, and published benchmarks.

3 of 6 selected
Clear all

Overview

Author Anthropic
Context length 1Mtokens
Reasoning Supported
Input modalities T
Output modalities T
Providers 34 providers
Author OpenAI
Context length 1.05MtokensBest
Reasoning Supported
Input modalities T
Output modalities T
Providers 36 providers
Author Google
Context length 1.04858Mtokens
Reasoning Supported
Input modalities T
Output modalities T
Providers 25 providers

Pricing per million tokens

Input $5
Output $25
Cached input $0.5
Cache write $6.25
Input $2
Output $12
Cached input $0.2
Cache write $2.5
Input $0.75Best
Output $3.75Best
Cached input $0.075
Cache write Not listed

Limits

Context length 1Mtokens
Max output 128KtokensBest
Released 2026-07-24
Knowledge cutoff 2026-05
Context length 1.05MtokensBest
Max output 128KtokensBest
Released 2026-07-09
Knowledge cutoff 2026-02-16
Context length 1.04858Mtokens
Max output 65.536Ktokens
Released 2026-07-21
Knowledge cutoff 2026-03

Capabilities

Tool calling Supported
Structured output Not reported
Vision input Supported
Attachments Supported
Tool calling Supported
Structured output Supported
Vision input Supported
Attachments Supported
Tool calling Supported
Structured output Supported
Vision input Supported
Attachments Supported

Availability

Inference providers 34providers
Provider list 302aiAnthropicabacusagentrouteraihubmix+27 more
Inference providers 36providersBest
Provider list 302aiOpenAIabacusai-routeraihubmix+30 more
Inference providers 25providers
Provider list 302aiGoogleabacuscortecscrossmodel+18 more

Benchmarks

Intelligence index
0 27.5 55 claude-opus-4… — 55 index 55 claude-opus-4… claude-opus-4… — 55 index 55 claude-opus-4… claude-fable-… — 53.4 index 53.4 claude-fable-… claude-fable-… — 53.4 index 53.4 claude-fable-… qwen3.8-max — 53.4 index 53.4 qwen3.8-max gpt-5.4:batch — 53.1 index 53.1 gpt-5.4:batch gpt-5.4 — 53.1 index 53.1 gpt-5.4 gpt-6-astra:b… — 52.8 index 52.8 gpt-6-astra:b… Claude Opus 5 — 50.7 index 50.7 Claude Opus 5 GPT-5.6 Terra — 42.3 index 42.3 GPT-5.6 Terra Gemini 3.6 Fl… — 34.3 index 34.3 Gemini 3.6 Fl…
  • Claude Opus 5
  • GPT-5.6 Terra
  • Gemini 3.6 Flash
Coding index
0 40.8 81.6 claude-fable-… — 81.6 index 81.6 claude-fable-… claude-fable-… — 81.6 index 81.6 claude-fable-… claude-opus-5… — 78 index 78 claude-opus-5… gpt-5.6-sol:b… — 77.4 index 77.4 gpt-5.6-sol:b… gpt-5.6-sol — 77.4 index 77.4 gpt-5.6-sol gpt-6-astra:b… — 76.9 index 76.9 gpt-6-astra:b… gpt-6-astra — 76.9 index 76.9 gpt-6-astra grok-4.6 — 76.8 index 76.8 grok-4.6 Claude Opus 5 — 78 index 78 Claude Opus 5 GPT-5.6 Terra — 76.7 index 76.7 GPT-5.6 Terra Gemini 3.6 Fl… — 69.2 index 69.2 Gemini 3.6 Fl…
  • Claude Opus 5
  • GPT-5.6 Terra
  • Gemini 3.6 Flash
Agentic index
0 29 58 claude-fable-… — 58 index 58 claude-fable-… claude-fable-… — 58 index 58 claude-fable-… claude-opus-5… — 56.2 index 56.2 claude-opus-5… glm-5.3:batch — 53.4 index 53.4 glm-5.3:batch glm-5.3 — 53.4 index 53.4 glm-5.3 grok-4.6 — 53.4 index 53.4 grok-4.6 gpt-6-astra:b… — 51.5 index 51.5 gpt-6-astra:b… gpt-6-astra — 51.5 index 51.5 gpt-6-astra Claude Opus 5 — 56.2 index 56.2 Claude Opus 5 GPT-5.6 Terra — 43.7 index 43.7 GPT-5.6 Terra Gemini 3.6 Fl… — 30.2 index 30.2 Gemini 3.6 Fl…
  • Claude Opus 5
  • GPT-5.6 Terra
  • Gemini 3.6 Flash
All published benchmark results 50 variants
AA-Briefcase Elo1720.0Best
ARC-AGI-1 accuracy97.5Best
ARC-AGI-2 accuracy90.4Best
ARC-AGI-3 RHAE30.2Best
Agentic Index index56.2Best
Agents' Last Exam scoreNot reported
Artificial Analysis Coding Agent Index index score · 1.1 · CodexNot reported
Artificial Analysis Intelligence Index index score · 4.1Not reported
AutomationBench success rate26.0Best
BrowseComp accuracy90.8Best
CharXiv Reasoning accuracyNot reported
Coding Index index78.0Best
DeepSWE resolve rate · 1.168.8
DeepSearchQA F195.0Best
Design Arena: 3d elo1364.0Best
Design Arena: agenticgamedev elo1265.0Best
Design Arena: androidnative elo1268.0Best
Design Arena: asciiart elo1386.0Best
Design Arena: codecategories elo1339.0Best
Design Arena: dataviz elo1355.0Best
Design Arena: fullstack elo1331.0Best
Design Arena: gamedev elo1366.0Best
Design Arena: htmlslides eloNot reported
Design Arena: mobileapps elo1348.0Best
Design Arena: python-pptxslides elo1265.0Best
Design Arena: svg elo1351.0Best
Design Arena: uicomponent elo1360.0Best
Design Arena: webapps elo1277.0Best
Design Arena: website elo1320.0Best
Frontier-Bench mean reward · v0.1 · mini-SWE-agent43.3Best
FrontierCode mean@5 · 1.153.4Best
FrontierMath accuracy · v2Not reported
GDM-MRCR accuracy · v2Not reported
GDPval-AA Elo · v21861.0Best
GPQA Diamond accuracyNot reported
HealthBench Professional score59.8Best
Humanity's Last Exam accuracy64.7Best
Intelligence Index index50.7Best
MLE-Bench average position scoreNot reported
MMMU Pro accuracyNot reported
OSWorld success rate · 2.070.6Best
OSWorld-Verified success rateNot reported
SWE-Bench Multilingual resolve rate89.5Best
SWE-Bench Multimodal resolve rate59.4Best
SWE-Bench Pro resolve rate79.2Best
SWE-Bench Pro resolve rate · AntigravityNot reported
SWE-Bench Verified resolved96.0Best
Terminal-Bench success rate · 2.1Not reported
Terminal-Bench accuracy · 2.1 · Terminus 2Not reported
Toolathlon success rateNot reported
AA-Briefcase EloNot reported
ARC-AGI-1 accuracyNot reported
ARC-AGI-2 accuracyNot reported
ARC-AGI-3 RHAENot reported
Agentic Index index43.7
Agents' Last Exam score50.4Best
Artificial Analysis Coding Agent Index index score · 1.1 · Codex77.4Best
Artificial Analysis Intelligence Index index score · 4.155.0Best
AutomationBench success rateNot reported
BrowseComp accuracy87.5
CharXiv Reasoning accuracyNot reported
Coding Index index76.7
DeepSWE resolve rate · 1.169.6Best
DeepSearchQA F1Not reported
Design Arena: 3d eloNot reported
Design Arena: agenticgamedev eloNot reported
Design Arena: androidnative eloNot reported
Design Arena: asciiart eloNot reported
Design Arena: codecategories eloNot reported
Design Arena: dataviz eloNot reported
Design Arena: fullstack eloNot reported
Design Arena: gamedev eloNot reported
Design Arena: htmlslides eloNot reported
Design Arena: mobileapps eloNot reported
Design Arena: python-pptxslides eloNot reported
Design Arena: svg eloNot reported
Design Arena: uicomponent eloNot reported
Design Arena: webapps eloNot reported
Design Arena: website eloNot reported
Frontier-Bench mean reward · v0.1 · mini-SWE-agentNot reported
FrontierCode mean@5 · 1.1Not reported
FrontierMath accuracy · v284.9Best
GDM-MRCR accuracy · v2Not reported
GDPval-AA Elo · v2Not reported
GPQA Diamond accuracy92.9Best
HealthBench Professional scoreNot reported
Humanity's Last Exam accuracyNot reported
Intelligence Index index42.3
MLE-Bench average position scoreNot reported
MMMU Pro accuracy80.7Best
OSWorld success rate · 2.050.2
OSWorld-Verified success rateNot reported
SWE-Bench Multilingual resolve rateNot reported
SWE-Bench Multimodal resolve rateNot reported
SWE-Bench Pro resolve rate63.4
SWE-Bench Pro resolve rate · AntigravityNot reported
SWE-Bench Verified resolvedNot reported
Terminal-Bench success rate · 2.187.4Best
Terminal-Bench accuracy · 2.1 · Terminus 2Not reported
Toolathlon success rate53.1Best
AA-Briefcase EloNot reported
ARC-AGI-1 accuracyNot reported
ARC-AGI-2 accuracyNot reported
ARC-AGI-3 RHAENot reported
Agentic Index index30.2
Agents' Last Exam scoreNot reported
Artificial Analysis Coding Agent Index index score · 1.1 · CodexNot reported
Artificial Analysis Intelligence Index index score · 4.1Not reported
AutomationBench success rateNot reported
BrowseComp accuracyNot reported
CharXiv Reasoning accuracy89.4Best
Coding Index index69.2
DeepSWE resolve rate · 1.149.0
DeepSearchQA F1Not reported
Design Arena: 3d elo1302.0
Design Arena: agenticgamedev elo1181.0
Design Arena: androidnative elo1215.0
Design Arena: asciiart elo1291.0
Design Arena: codecategories elo1305.0
Design Arena: dataviz elo1312.0
Design Arena: fullstack elo1195.0
Design Arena: gamedev elo1284.0
Design Arena: htmlslides elo1150.0Best
Design Arena: mobileapps elo1232.0
Design Arena: python-pptxslides elo1146.0
Design Arena: svg eloNot reported
Design Arena: uicomponent elo1317.0
Design Arena: webapps elo1217.0
Design Arena: website elo1312.0
Frontier-Bench mean reward · v0.1 · mini-SWE-agentNot reported
FrontierCode mean@5 · 1.1Not reported
FrontierMath accuracy · v2Not reported
GDM-MRCR accuracy · v254.0Best
GDPval-AA Elo · v21421.0
GPQA Diamond accuracyNot reported
HealthBench Professional scoreNot reported
Humanity's Last Exam accuracyNot reported
Intelligence Index index34.3
MLE-Bench average position score63.9Best
MMMU Pro accuracyNot reported
OSWorld success rate · 2.0Not reported
OSWorld-Verified success rate83.0Best
SWE-Bench Multilingual resolve rateNot reported
SWE-Bench Multimodal resolve rateNot reported
SWE-Bench Pro resolve rateNot reported
SWE-Bench Pro resolve rate · Antigravity58.7Best
SWE-Bench Verified resolvedNot reported
Terminal-Bench success rate · 2.1Not reported
Terminal-Bench accuracy · 2.1 · Terminus 278.0Best
Toolathlon success rateNot reported

Provenance

Verification Source-linked
Confidence Medium
Last updated Sep 11, 2026
Source Models.dev
Verification Source-linked
Confidence Medium
Last updated Sep 11, 2026
Source Models.dev
Verification Source-linked
Confidence Medium
Last updated Sep 11, 2026
Source Models.dev

“Not reported” means the source catalog is silent on that field. Benchmark rows keep each published version and harness separate.