164 models
Ranked by Coding Index
17.4

Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...

qwen/qwen3-next-80b-a3b-thinking 262.144K context $0.15/M input $1.2/M output
16.3

OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...

openai/o3-mini-high 200K context $1.1/M input $4.4/M output
16.3

OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...

openai/o3-mini-high:batch 200K context $0.55/M input $2.2/M output
Reasoning 15.6

Small GPT-5 for responsive agents, coding help, and everyday automation

openai/gpt-5-mini 2025-08-07 400K context $0.25/M input $2/M output
29 providers
15.6

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....

openai/gpt-5-mini:batch 400K context $0.125/M input $1/M output
14.4

The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...

mistralai/ministral-14b-2512 262.144K context $0.2/M input $0.2/M output
Open weights 14.4

Small Nemotron 3 MoE for efficient coding, math, and long-context agents

nvidia/nemotron-3-nano-30b-a3b 2025-12-15 262.144K context $0.05/M input $0.2/M output
11 providers
14.4

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...

nvidia/nemotron-3-nano-30b-a3b:free 256K context Free input Free output
13.8

NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free 256K context Free input Free output
11.1

Tiny GPT-4.1 option for classification, routing, and very high-volume tasks

openai/gpt-4.1-nano 2025-04-14 1.04758M context $0.1/M input $0.4/M output
20 providers
11.1

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...

openai/gpt-4.1-nano:batch 1.04758M context $0.05/M input $0.2/M output
9.7

A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

mistralai/ministral-8b-2512 262.144K context $0.15/M input $0.15/M output
9.7

A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

mistralai/ministral-8b-2512:batch 262.144K context $0.075/M input $0.075/M output
8.2

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...

meta-llama/llama-4-scout 327.68K context $0.1/M input $0.3/M output