162 models
Ranked by Intelligence Index
21.2

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...

openai/gpt-5.4-nano:batch 400K context $0.1/M input $0.625/M output
21.2

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

anthropic/claude-sonnet-4.5 1M context $3/M input $15/M output
Reasoning Tools JSON 21.2

Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation

openai/gpt-5.4-nano 2026-03-17 400K context $0.2/M input $1.25/M output
27 providers
21.2

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

anthropic/claude-sonnet-4.5:batch 1M context $1.5/M input $7.5/M output
20.2

North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...

cohere/north-mini-code:free 256K context Free input Free output
Reasoning Tools 19.7

Meituan LongCat-2.0, a reasoning model with tool calling and a 1M-token context window

meituan/longcat-2.0 2026-06-30 1M context $0.3/M input $1.2/M output
5 providers
19.1

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...

qwen/qwen3.5-397b-a17b 262.144K context $0.55/M input $3.5/M output
18.8

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...

qwen/qwen3.6-35b-a3b 262.144K context $0.1/M input $0.9/M output
18.4

Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...

amazon/nova-2-lite-v1 1M context $0.3/M input $2.5/M output
17.6

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

anthropic/claude-haiku-4.5 200K context $1/M input $5/M output
17.6

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

anthropic/claude-haiku-4.5:batch 200K context $0.5/M input $2.5/M output
Reasoning 17.4

Small GPT-5 for responsive agents, coding help, and everyday automation

openai/gpt-5-mini 2025-08-07 400K context $0.25/M input $2/M output
29 providers
17.4

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....

openai/gpt-5-mini:batch 400K context $0.125/M input $1/M output
16.9

Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...

qwen/qwen3-next-80b-a3b-thinking 262.144K context $0.15/M input $1.2/M output
16.7

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

google/gemini-2.5-pro:batch 1.04858M context $0.625/M input $5/M output
Reasoning Tools JSON 16.7

Google's proven reasoning model for coding, math, and multimodal analysis

google/gemini-2.5-pro 2025-06-17 1.04858M context $1.25/M input $10/M output
29 providers
16.2

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...

qwen/qwen3.5-122b-a10b 262.144K context $0.26/M input $2.08/M output
Reasoning Tools JSON 16.0

Low-latency Gemini model for high-volume multimodal and agent workloads

google/gemini-3.1-flash-lite-preview 2026-03-03 1.04858M context $0.25/M input $1.5/M output
13 providers
15.7

OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...

openai/o3-mini-high:batch 200K context $0.55/M input $2.2/M output
15.4

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...

deepseek/deepseek-v3.1-terminus 131.072K context $0.27/M input $1/M output
15.4

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

google/gemma-4-31b-it:batch 262.144K context $0.39/M input $0.97/M output
15.4

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

google/gemma-4-31b-it:free 262.144K context Free input Free output
15.2

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...

openai/gpt-oss-20b:free 131.072K context Free input Free output
14.9

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

mistralai/mistral-medium-3-5:batch 262.144K context $0.75/M input $3.75/M output
14.9

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

mistralai/mistral-medium-3-5 262.144K context $1.5/M input $7.5/M output
14.8

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...

openai/gpt-4.1-mini:batch 1.04758M context $0.2/M input $0.8/M output
14.8

Affordable GPT-4.1 lane for fast coding help and structured extraction

openai/gpt-4.1-mini 2025-04-14 1.04758M context $0.4/M input $1.6/M output
25 providers
14.7

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...

mistralai/mistral-medium-3.1 131.072K context $0.4/M input $2/M output
14.7

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...

mistralai/mistral-medium-3.1:batch 131.072K context $0.2/M input $1/M output
14.5

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...

nvidia/nemotron-3-nano-30b-a3b:free 256K context Free input Free output
14.2

Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....

inclusionai/ling-2.6-flash 262.144K context $0.01/M input $0.03/M output
13.9

Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary...

cohere/command-a 256K context $2.5/M input $10/M output
13.6

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

nvidia/nemotron-3.5-lightning:free 1M context Free input Free output
13.6

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

nvidia/nemotron-3-super-120b-a12b:free 262.144K context Free input Free output
12.7

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...

qwen/qwen3-235b-a22b-thinking-2507 131.072K context $0.23/M input $2.3/M output
12.3

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...

openai/gpt-oss-120b:batch 131.072K context $0.15/M input $0.6/M output
11.5

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving...

inception/mercury-2 128K context $0.25/M input $0.75/M output
11.5

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...

mistralai/mistral-small-2603:batch 262.144K context $0.075/M input $0.3/M output
11.5

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...

mistralai/mistral-small-2603 262.144K context $0.15/M input $0.6/M output
11.0

OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...

openai/o3-mini-high 200K context $1.1/M input $4.4/M output
10.1

Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total parameters and only 3B activated per...

qwen/qwen3-coder-next 262.144K context $0.12/M input $0.8/M output
9.8

Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model is designed specifically for “thinking mode,” where internal reasoning traces are separated...

qwen/qwen3-30b-a3b-thinking-2507 81.92K context $0.2/M input $2.4/M output
9.7

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...

deepseek/deepseek-chat-v3-0324 163.84K context $0.25/M input $1/M output
9.7

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

mistralai/mistral-large-2512:batch 262.144K context $0.25/M input $0.75/M output
9.7

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

mistralai/mistral-large-2512 262.144K context $0.5/M input $1.5/M output
9.6

Tiny GPT-4.1 option for classification, routing, and very high-volume tasks

openai/gpt-4.1-nano 2025-04-14 1.04758M context $0.1/M input $0.4/M output
20 providers
9.6

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...

openai/gpt-4.1-nano:batch 1.04758M context $0.05/M input $0.2/M output
9.4

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

meta-llama/llama-3.3-70b-instruct:free 65.536K context Free input Free output
9.4

Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256K context window. Devstral 2 supports exploring...

mistralai/devstral-2512 262.144K context $0.4/M input $2/M output
9.3

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

meta-llama/llama-3.3-70b-instruct 131.072K context $0.1/M input $0.32/M output