2,824 models
Open weights

R1 reasoning distilled into Qwen 2.5 32B for efficient open-weight step-by-step problem solving

deepseek/deepseek-r1-distill-qwen-32b 2025-01-20 131.072K context $0.3/M input $0.3/M output
1 provider
Reasoning Tools Open weights

Classic open reasoning model for transparent math, coding, and deliberate problem solving

deepseek/deepseek-r1 2025-01-20 128K context $0.7/M input $2.5/M output
14 providers
Tools Open weights

Open DeepSeek MoE chat model for coding, math, and general reasoning

deepseek/deepseek-v3 2024-12-26 131.072K context $0.27/M input $1.12/M output
8 providers
Reasoning Tools JSON

Smaller o-series reasoner for economical coding, math, and planning tasks

openai/o3-mini 2024-12-20 200K context $1.1/M input $4.4/M output
19 providers
Tools

Earlier Gemini Flash workhorse for responsive multimodal apps and tool use

google/gemini-2.0-flash 2024-12-11 1.04858M context $0.1/M input $0.42/M output
2 providers
Tools Open weights

Compact Microsoft instruction model tuned for efficient coding assistance, reasoning, and low-latency agent tasks

microsoft/phi-4-mini 2024-12-11 128K context $0.075/M input $0.3/M output
2 providers
Tools

Low-latency Gemini model for high-volume multimodal and agent workloads

google/gemini-2.0-flash-lite 2024-12-11 1.04858M context $0.052/M input $0.21/M output
2 providers
Tools Open weights

Popular open Llama workhorse for multilingual chat, coding, and self-hosting

meta/llama-3.3-70b-instruct 2024-12-06 128K context $0.1/M input $0.32/M output
25 providers
Reasoning

O-series reasoning model for hard analysis, math, coding, and planning

openai/o1 2024-12-05 200K context $15/M input $60/M output
16 providers
Tools

Efficient model for low-latency assistance, extraction, and routine automation

amazon/nova-lite 2024-12-03 300K context $0.06/M input $0.24/M output
3 providers
Tools

Efficient model for low-latency assistance, extraction, and routine automation

amazon/nova-micro 2024-12-03 128K context $0.035/M input $0.14/M output
3 providers
Tools

Flagship model for demanding analysis, coding, and production agent workflows

amazon/nova-pro 2024-12-03 300K context $0.8/M input $3.2/M output
3 providers
JSON Open weights

Cohere retrieval model for long-context chat and enterprise RAG workflows

cohere/command-r7b-12-2024 2024-12-02 128K context $0.037/M input $0.15/M output
5 providers
Tools

GPT model for general reasoning, writing, coding, and tool-assisted tasks

openai/gpt-4o-2024-11-20 2024-11-20 128K context $2.5/M input $10/M output
8 providers
Tools Open weights

Flagship Mistral model for advanced reasoning, coding, and multilingual work

mistral/mistral-large-2411 2024-11-18 131.072K context $2/M input $6/M output
2 providers
Tools JSON Open weights

Open coding-focused Qwen model for code generation, repair, and repository reasoning

alibaba/qwen2.5-coder-32b-instruct 2024-11-12 131.072K context $0.06/M input $0.2/M output
2 providers
Tools Open weights

Mistral's larger vision model for document-heavy image understanding and chat

mistral/pixtral-large-latest 2024-11-01 128K context $2/M input $6/M output
5 providers
Open weights

Flagship Mistral model for advanced reasoning, coding, and multilingual work

mistral/mistral-large-latest 2024-11-01 262.144K context $0.5/M input $1.5/M output
7 providers

Efficient Qwen model for fast chat, extraction, and high-volume workloads

alibaba/qwen-turbo 2024-11-01 1M context $0.05/M input $0.2/M output
5 providers
Open weights

Open multilingual model optimized for generation across 23 languages

cohere/c4ai-aya-expanse-32b 2024-10-24 128K context Input not listed Output not listed
1 provider
Tools

Fast Claude model for responsive assistance, classification, and lightweight agents

anthropic/claude-3-5-haiku-20241022 2024-10-22 200K context $0.8/M input $4/M output
3 providers

Balanced Claude model for coding, analysis, agent workflows, and cost control

anthropic/claude-3-5-sonnet-20241022 2024-10-22 200K context Input not listed Output not listed
1 provider
Tools Open weights

Efficient open Mistral edge model for on-device chat and function calling

mistral/ministral-8b-instruct-2410 2024-10-16 131.072K context $0.15/M input $0.15/M output
1 provider
Open weights

Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads

mistral/ministral-3b 2024-10-16 128K context $0.04/M input $0.04/M output
5 providers
Reasoning Tools

Enterprise language model for workflow automation, coding, data analysis, and tool use

writer/palmyra-x4 2024-10-09 128K context $2.5/M input $10/M output
2 providers
Open weights

Small open Llama base model for lightweight text generation and self-hosting

meta/llama-3.2-3b 2024-09-25 131.072K context $0.1/M input $0.1/M output
1 provider
Open weights

Compact open Llama base model for lightweight and on-device use

meta/llama-3.2-1b 2024-09-25 131.072K context $0.1/M input $0.1/M output
1 provider
Tools JSON Open weights

Open multimodal Llama model for image understanding, captioning, and visual QA

meta/llama-3.2-11b-vision-instruct 2024-09-25 128K context $0.055/M input $0.055/M output
3 providers
Tools Open weights

Mistral vision-language model for image understanding and multimodal chat

mistral/pixtral-12b 2024-09-01 128K context $0.15/M input $0.15/M output
4 providers
Tools Open weights

Qwen vision-language model for visual reasoning, documents, and agent tasks

alibaba/qwen2-5-vl-72b-instruct 2024-09 131.072K context $2.8/M input $8.4/M output
3 providers
Open weights

Cohere's RAG workhorse for long-context enterprise search and tool use

cohere/command-r-plus-08-2024 2024-08-30 128K context $2.5/M input $10/M output
9 providers
Tools JSON Open weights

Cohere retrieval model for long-context chat and enterprise RAG workflows

cohere/command-r-08-2024 2024-08-30 128K context $0.15/M input $0.6/M output
7 providers
Tools Open weights

Compact Nemotron model for efficient reasoning and deployable AI agents

nvidia/nemotron-mini-4b-instruct 2024-08-21 128K context Input not listed Output not listed
1 provider
Tools

GPT model for general reasoning, writing, coding, and tool-assisted tasks

openai/gpt-4o-2024-08-06 2024-08-06 128K context $2.5/M input $10/M output
7 providers
Open weights

Llama 3.1-based safety classifier for moderating prompts and model responses

meta/llama-guard-3-8b 2024-07-23 128K context Input not listed Output not listed
1 provider
Tools JSON Open weights

Compact open Llama model for lightweight chat, drafting, and self-hosting

meta/llama-3.1-8b-instruct 2024-07-23 128K context $0.02/M input $0.04/M output
7 providers
Tools JSON Open weights

Open Llama instruction model for multilingual chat, reasoning, and coding

meta/llama-3.1-70b-instruct 2024-07-23 128K context $0.4/M input $0.4/M output
4 providers

Small omni GPT for cheap multimodal assistance and production-scale traffic

openai/gpt-4o-mini 2024-07-18 128K context $0.15/M input $0.6/M output
23 providers
Open weights

Efficient Mistral-NVIDIA open model for multilingual chat and local deployment

mistral/mistral-nemo 2024-07-01 128K context $0.15/M input $0.15/M output
11 providers
Reasoning Tools

Research model for long-horizon investigation, synthesis, and analytical reports

openai/o3-deep-research 2024-06-26 200K context $9/M input $36/M output
6 providers
Reasoning Tools

Research model for long-horizon investigation, synthesis, and analytical reports

openai/o4-mini-deep-research 2024-06-26 200K context $1.8/M input $7.2/M output
5 providers
Tools Open weights

Mistral code model for completions, refactors, and developer IDE workflows

mistral/codestral-latest 2024-05-29 256K context $0.3/M input $0.9/M output
5 providers
Tools

Omni-era GPT for multimodal chat, practical coding, and general assistants

openai/gpt-4o 2024-05-13 128K context $2.5/M input $10/M output
23 providers
Tools JSON

GPT model for general reasoning, writing, coding, and tool-assisted tasks

openai/gpt-4o-2024-05-13 2024-05-13 128K context $5/M input $15/M output
5 providers
Tools

Qwen vision-language model for visual reasoning, documents, and agent tasks

alibaba/qwen-vl-max 2024-04-08 131.072K context $0.8/M input $3.2/M output
5 providers
Tools

Legacy model retained for compatibility with older integrations

anthropic/claude-3-haiku-20240307 2024-03-13 200K context $0.25/M input $1.25/M output
2 providers
Reasoning

Qwen instruction model for multilingual chat, reasoning, and tool use

alibaba/qwen-plus 2024-01-25 1M context $0.4/M input $1.2/M output
10 providers
Tools

Qwen vision-language model for visual reasoning, documents, and agent tasks

alibaba/qwen-vl-plus 2024-01-25 131.072K context $0.21/M input $0.63/M output
4 providers
JSON

Web-grounded Sonar for multi-step research questions that need cited reasoning

perplexity/sonar-reasoning-pro 2024-01-01 128K context $2/M input $8/M output
9 providers
JSON

Fast web-grounded Sonar for current answers, citations, and lightweight retrieval

perplexity/sonar 2024-01-01 128K context $1/M input $1/M output
9 providers