Reasoning Tools

Flagship Claude model for deep reasoning, coding, and long-horizon agents

anthropic/claude-opus-4-1 2025-08-05 200K context $15/M input $75/M output
9 providers
Reasoning Tools Open weights

Hybrid-reasoning GLM release that made the 4.5 line broadly useful

zhipuai/glm-4.5 2025-07-28 131.072K context $0.6/M input $2.2/M output
14 providers

Qwen coding model for software agents, repository edits, and code reasoning

alibaba/qwen3-coder-flash 2025-07-28 1M context $0.3/M input $1.5/M output
12 providers
Tools

Efficient Qwen model for fast chat, extraction, and high-volume workloads

alibaba/qwen-flash 2025-07-28 1M context $0.05/M input $0.4/M output
7 providers
Reasoning Tools Open weights

Nemotron model for efficient reasoning, coding, and specialized AI agents

nvidia/llama-3.3-nemotron-super-49b-v1.5 2025-07-25 131.072K context $0.4/M input $0.4/M output
4 providers

Hosted Qwen coder for software agents, repo edits, and long-context code

alibaba/qwen3-coder-plus 2025-07-23 1.04858M context $1/M input $5/M output
14 providers
Tools Open weights

Updated large open Qwen3 MoE instruct model for multilingual chat, coding, and tool use

alibaba/qwen3-235b-a22b-instruct-2507 2025-07-21 262.144K context $0.069/M input $0.455/M output
7 providers
Tools Open weights

Instruct model with native audio input for speech understanding and tool use

mistral/voxtral-small-latest 2025-07-15 32K context $0.1/M input $0.3/M output
3 providers
Reasoning Tools JSON

Fast Gemini workhorse for multimodal apps where latency and price matter

google/gemini-2.5-flash 2025-06-17 1.04858M context $0.3/M input $2.5/M output
30 providers
Reasoning Tools JSON

Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents

google/gemini-2.5-flash-lite 2025-06-17 1.04858M context $0.1/M input $0.4/M output
18 providers
Reasoning Tools JSON

Google's proven reasoning model for coding, math, and multimodal analysis

google/gemini-2.5-pro 2025-06-17 1.04858M context $1.25/M input $10/M output
29 providers
Reasoning Tools

High-effort o3 tier for difficult technical reasoning and careful answers

openai/o3-pro 2025-06-10 200K context $20/M input $80/M output
8 providers
Reasoning Tools

Flagship Claude model for deep reasoning, coding, and long-horizon agents

anthropic/claude-opus-4-20250514 2025-05-22 200K context $15/M input $75/M output
8 providers
Reasoning Tools JSON

Balanced Claude model for coding, analysis, agent workflows, and cost control

anthropic/claude-sonnet-4-0 2025-05-22 200K context $2.898/M input $14.493/M output
4 providers

Mistral model for multilingual chat, reasoning, and tool-assisted workflows

mistral/mistral-medium-2505 2025-05-07 131.072K context $0.4/M input $2/M output
10 providers
Open weights

Sparse MoE Qwen model with 3B active parameters for efficient chat and reasoning

alibaba/qwen3-30b-a3b 2025-04-28 131.072K context $0.08/M input $0.29/M output
8 providers
Reasoning Tools JSON

Fast o-series model for compact reasoning, coding, and tool use

openai/o4-mini 2025-04-16 200K context $1.1/M input $4.4/M output
20 providers
Reasoning

Deliberate o-series reasoner for hard math, coding, and multi-step analysis

openai/o3 2025-04-16 200K context $2/M input $8/M output
20 providers

Tiny GPT-4.1 option for classification, routing, and very high-volume tasks

openai/gpt-4.1-nano 2025-04-14 1.04758M context $0.1/M input $0.4/M output
20 providers

Affordable GPT-4.1 lane for fast coding help and structured extraction

openai/gpt-4.1-mini 2025-04-14 1.04758M context $0.4/M input $1.6/M output
25 providers
Tools JSON

Long-lived GPT workhorse for coding, instruction following, and production apps

openai/gpt-4.1 2025-04-14 1.04758M context $2/M input $8/M output
28 providers
Reasoning Tools JSON

Mistral vision-language model for image understanding and multimodal chat

mistral/pixtral-large-2502 2025-04-08 128K context $1.993/M input $5.978/M output
1 provider
Reasoning Tools Open weights

Flagship Nemotron model for high-throughput reasoning and complex agents

nvidia/llama-3.1-nemotron-ultra-253b 2025-04-07 128K context Input not listed Output not listed
2 providers
Tools Open weights

Open multimodal Llama for strong reasoning with efficient everyday serving

meta/llama-4-maverick-17b-instruct 2025-04-05 1M context $0.14/M input $0.59/M output
6 providers
Tools Open weights

Smaller Qwen coder for efficient local agents and repo-level fixes

alibaba/qwen3-coder-30b-a3b-instruct 2025-04 262.144K context $0.45/M input $2.25/M output
13 providers
Reasoning Tools Open weights

Large open Qwen MoE for multilingual reasoning, coding, and tool use

alibaba/qwen3-235b-a22b 2025-04 131.072K context $0.7/M input $2.8/M output
5 providers
Reasoning Tools JSON Open weights

Dense open Qwen model for self-hosted chat, reasoning, and coding

alibaba/qwen3-32b 2025-04 131.072K context $0.7/M input $2.8/M output
15 providers
Tools Open weights

Open Qwen coding heavyweight for repository reasoning and agentic engineering

alibaba/qwen3-coder-480b-a35b-instruct 2025-04 262.144K context $1.5/M input $7.5/M output
8 providers
Tools JSON Open weights

March 2025 checkpoint of DeepSeek-V3 with improved reasoning and coding

deepseek/deepseek-v3-0324 2025-03-24 163.84K context $0.2/M input $0.8/M output
9 providers

O-series reasoning model for hard analysis, math, coding, and planning

openai/o1-pro 2025-03-19 200K context $150/M input $600/M output
7 providers
Reasoning Tools

Mistral reasoning model for transparent analysis, math, and complex decisions

mistral/magistral-medium-latest 2025-03-17 128K context $2/M input $5/M output
6 providers
Tools JSON Open weights

Efficient multimodal model for instruction following, coding, reasoning, and function calling

mistral/mistral-small-3-1-24b-instruct-2503 2025-03-17 128K context $0.106/M input $0.318/M output
2 providers
Reasoning Tools Open weights

Cohere command model for multilingual enterprise agents, tools, and chat

cohere/command-a-03-2025 2025-03-13 256K context $2.5/M input $10/M output
6 providers
Open weights

Open multimodal Gemma instruction model for multilingual text generation and image understanding

google/gemma-3-12b-it 2025-03-12 131.072K context $0.05/M input $0.1/M output
8 providers
Open weights

Largest open Gemma 3 instruction model for multilingual text generation and visual understanding

google/gemma-3-27b-it 2025-03-12 131.072K context $0.08/M input $0.16/M output
10 providers
Tools JSON Open weights

Open multimodal Gemma instruction model for efficient text generation and image understanding

google/gemma-3-4b-it 2025-03-12 131.072K context $0.04/M input $0.08/M output
7 providers
Reasoning Tools

Balanced Claude model for coding, analysis, agent workflows, and cost control

anthropic/claude-3-7-sonnet-20250219 2025-02-19 200K context $3/M input $15/M output
4 providers
Reasoning Tools Open weights

Classic open reasoning model for transparent math, coding, and deliberate problem solving

deepseek/deepseek-r1 2025-01-20 128K context $0.7/M input $2.5/M output
14 providers
Tools Open weights

Open DeepSeek MoE chat model for coding, math, and general reasoning

deepseek/deepseek-v3 2024-12-26 131.072K context $0.27/M input $1.12/M output
8 providers
Reasoning Tools JSON

Smaller o-series reasoner for economical coding, math, and planning tasks

openai/o3-mini 2024-12-20 200K context $1.1/M input $4.4/M output
19 providers
Tools

Low-latency Gemini model for high-volume multimodal and agent workloads

google/gemini-2.0-flash-lite 2024-12-11 1.04858M context $0.052/M input $0.21/M output
2 providers
Tools

Earlier Gemini Flash workhorse for responsive multimodal apps and tool use

google/gemini-2.0-flash 2024-12-11 1.04858M context $0.1/M input $0.42/M output
2 providers
Tools Open weights

Popular open Llama workhorse for multilingual chat, coding, and self-hosting

meta/llama-3.3-70b-instruct 2024-12-06 128K context $0.1/M input $0.32/M output
25 providers
Reasoning

O-series reasoning model for hard analysis, math, coding, and planning

openai/o1 2024-12-05 200K context $15/M input $60/M output
16 providers
Tools

Efficient model for low-latency assistance, extraction, and routine automation

amazon/nova-micro 2024-12-03 128K context $0.035/M input $0.14/M output
3 providers
Tools

Efficient model for low-latency assistance, extraction, and routine automation

amazon/nova-lite 2024-12-03 300K context $0.06/M input $0.24/M output
3 providers
Tools

Flagship model for demanding analysis, coding, and production agent workflows

amazon/nova-pro 2024-12-03 300K context $0.8/M input $3.2/M output
3 providers
JSON Open weights

Cohere retrieval model for long-context chat and enterprise RAG workflows

cohere/command-r7b-12-2024 2024-12-02 128K context $0.037/M input $0.15/M output
5 providers
Tools

GPT model for general reasoning, writing, coding, and tool-assisted tasks

openai/gpt-4o-2024-11-20 2024-11-20 128K context $2.5/M input $10/M output
8 providers
Tools JSON Open weights

Open coding-focused Qwen model for code generation, repair, and repository reasoning

alibaba/qwen2.5-coder-32b-instruct 2024-11-12 131.072K context $0.06/M input $0.2/M output
2 providers