7 models
Ranked by Terminal-Bench 2.0
Reasoning Tools JSON 82.7

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

openai/gpt-5.5 2026-04-23 1.05M context $5/M input $30/M output
46 providers
Reasoning Tools JSON 75.1

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...

openai/gpt-5.4 2026-03-05 1.05M context $2.5/M input $15/M output
42 providers
Reasoning Tools 69.7

Qwen frontier model tuned for agent frameworks, coding assistants, and long tasks

alibaba/qwen3.7-max 2026-05-21 1M context $2.5/M input $7.5/M output
31 providers
Reasoning Tools JSON 60.0

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...

openai/gpt-5.4-mini 2026-03-17 400K context $0.75/M input $4.5/M output
31 providers
Reasoning Tools JSON 54.8

Microsoft coding model built for fast, efficient assistance in everyday developer workflows

microsoft/mai-code-1-flash 2026-06-02 256K context $0.75/M input $4.5/M output
1 provider
Reasoning Tools JSON 46.3

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...

openai/gpt-5.4-nano 2026-03-17 400K context $0.2/M input $1.25/M output
27 providers
Reasoning Tools Open weights 37.5

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...

poolside/laguna-xs-2.1 2026-07-02 262.144K context $0.06/M input $0.12/M output
6 providers