3,455 models

Ministral 8B is an 8B parameter model featuring a unique interleaved sliding-window attention pattern for faster, memory-efficient inference. Designed for edge use cases, it supports up to 128k context length...

mistralai/ministral-8b 128K context $0.11/M input $0.11/M output

Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...

sakana/fugu-max 1M context $2/M input $6/M output

No provider description is available for this model yet.

openrouter/openai/gpt-5.6-sol 1.05M context $2/M input $10/M output

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

anthropic/claude-opus-4.6:batch 1M context $2.5/M input $12.5/M output

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,...

z-ai/glm-4.5v 65.536K context $0.6/M input $1.8/M output

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...

openai/gpt-4o-mini:batch 128K context $0.075/M input $0.3/M output
Open weights

Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort,...

ibm-granite/granite-4.2-8b 131.072K context $0.06/M input $0.25/M output

No provider description is available for this model yet.

zai/glm-4.5v 128K context $0.6/M input $1.8/M output

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

z-ai/glm-5.3-flash 1.04858M context $0.15/M input $0.5/M output

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K...

openai/gpt-5.4-pro:batch 1.05M context $15/M input $90/M output

Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

x-ai/grok-4.6 500K context $2/M input $6/M output

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...

inclusionai/ling-3.0-flash 262.144K context $0.021/M input $0.063/M output

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...

anthropic/claude-opus-4.1 200K context $15/M input $75/M output

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

openai/gpt-5.6-terra:batch 1.05M context $1/M input $6/M output

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

nvidia/nemotron-3-ultra-550b-a55b:free 1M context Free input Free output

No provider description is available for this model yet.

azure_ai/gpt-chat-latest 272K context $5/M input $30/M output

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...

openai/gpt-5.1:batch 400K context $0.625/M input $5/M output

No provider description is available for this model yet.

azure_ai/model-router 200K context $0.14/M input Output not listed

No provider description is available for this model yet.

azure_ai/cohere-command-a 131.072K context $2.5/M input $10/M output

No provider description is available for this model yet.

vertex_ai/xai/grok-4.3 200K context $1.25/M input $2.5/M output

No provider description is available for this model yet.

azure_ai/grok-4-20-reasoning 262K context $1.25/M input $2.5/M output

No provider description is available for this model yet.

vertex_ai/xai/grok-4.6 524.288K context $2/M input $6/M output

No provider description is available for this model yet.

azure_ai/grok-4-20-non-reasoning 262K context $1.25/M input $2.5/M output

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous programming via tool calling and...

qwen/qwen3-coder-plus 1M context $0.65/M input $3.25/M output

No provider description is available for this model yet.

gemini/lyria-3.5 1.04858M context Input not listed Output not listed

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

google/gemini-3.7-flash:batch 1.04858M context $0.375/M input $1.875/M output

No provider description is available for this model yet.

azure/gpt-6-astra 922K context $10/M input $50/M output

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)

mistralai/codestral-2508 256K context $0.3/M input $0.9/M output

No provider description is available for this model yet.

azure/us/gpt-6-astra 922K context $11/M input $55/M output

No provider description is available for this model yet.

azure_ai/codestral-2501 256K context $0.3/M input $0.9/M output

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,...

minimax/minimax-m2 204.8K context $0.255/M input $1.02/M output

No provider description is available for this model yet.

azure_ai/mai-thinking-1 256K context $2/M input $8/M output

No provider description is available for this model yet.

azure_ai/grok-4.6 200K context $2/M input $6/M output

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...

qwen/qwen3.7-plus 1M context $0.32/M input $1.28/M output

No provider description is available for this model yet.

databricks/databricks-glm-5-3 1.04858M context $1.4/M input $4.4/M output

No provider description is available for this model yet.

databricks/databricks-grok-4-6 500K context $2.5/M input $7.5/M output