4,213 models

No provider description is available for this model yet.

azure/eu/gpt-5.6-terra 1.05M context $2.75/M input $16.5/M output

No provider description is available for this model yet.

azure/eu/gpt-5.6-luna 1.05M context $1.1/M input $6.6/M output

No provider description is available for this model yet.

perplexity/pplx-7b-online 4.096K context Input not listed $0.28/M output

No provider description is available for this model yet.

perplexity/sonar-medium-chat 16.384K context $0.6/M input $1.8/M output

No provider description is available for this model yet.

perplexity/sonar-medium-online 12K context Input not listed $1.8/M output
Open weights

No provider description is available for this model yet.

meta-llama/Llama-2-13b-hf Not documented context Input not listed Output not listed

No provider description is available for this model yet.

cerebras/zai-glm-4.6 128K context $2.25/M input $2.75/M output

No provider description is available for this model yet.

azure/gpt-5.5 1.05M context $5/M input $30/M output

No provider description is available for this model yet.

azure/us/gpt-5.5 1.05M context $5.5/M input $33/M output

No provider description is available for this model yet.

azure/eu/gpt-5.5 1.05M context $5.5/M input $33/M output

Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...

inclusionai/ling-3.0-flash-fin:free 262.144K context Free input Free output

No provider description is available for this model yet.

bedrock/eu.twelvelabs.pegasus-1-2-v1:0 Not documented context Input not listed $7.5/M output

No provider description is available for this model yet.

nvidia/Kimi-K2.6-DFlash Not documented context Input not listed Output not listed

No provider description is available for this model yet.

azure/gpt-5.5-2026-04-23 1.05M context $5/M input $30/M output

No provider description is available for this model yet.

cerebras/llama-3.3-70b 128K context $0.85/M input $1.2/M output

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...

nex-agi/nex-n2.5-mini:free 262.144K context Free input Free output

No provider description is available for this model yet.

bedrock/us.twelvelabs.pegasus-1-2-v1:0 Not documented context Input not listed $7.5/M output

No provider description is available for this model yet.

azure/eu/gpt-5.5-2026-04-23 1.05M context $5.5/M input $33/M output

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

mistralai/mistral-medium-3-5 262.144K context $1.5/M input $7.5/M output

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...

inception/mercury-2.5-preview 260K context $0.2/M input $0.75/M output

No provider description is available for this model yet.

azure/gpt-5.4-mini 272K context $0.75/M input $4.5/M output

No provider description is available for this model yet.

cerebras/llama3.1-70b 128K context $0.6/M input $0.6/M output
Open weights

No provider description is available for this model yet.

ibm-granite/granite-4.2-30b Not documented context Input not listed Output not listed
Open weights

No provider description is available for this model yet.

ibm-granite/granite-4.2-3b Not documented context Input not listed Output not listed

No provider description is available for this model yet.

azure/gpt-5.4-mini-2026-03-17 272K context $0.75/M input $4.5/M output

A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

mistralai/ministral-8b-2512:batch 262.144K context $0.075/M input $0.075/M output

The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.

mistralai/ministral-3b-2512 131.072K context $0.1/M input $0.1/M output

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

deepseek/deepseek-v4-pro-0813:batch 1.04858M context $0.66/M input $1.98/M output

Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...

qwen/qwen3.5-9b:batch 262.144K context $0.17/M input $0.25/M output

Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fetch enabled, then a...

openrouter/fusion 1M context Input not listed Output not listed

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...

qwen/qwen3.8-max-0902 1M context $2/M input $6/M output

No provider description is available for this model yet.

azure/gpt-5.4-nano-2026-03-17 272K context $0.2/M input $1.25/M output

No provider description is available for this model yet.

azure/mistral-large-2402 32K context $8/M input $24/M output

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...

qwen/qwen3.7-flash 1M context $0.03/M input $0.13/M output

No provider description is available for this model yet.

azure/mistral-large-latest 32K context $8/M input $24/M output

No provider description is available for this model yet.

azure/o1 200K context $15/M input $60/M output

NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and...

nvidia/nemotron-nano-9b-v2:free 128K context Free input Free output