3,520 models

OpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high. OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining...

openai/o4-mini-high 200K context $1.1/M input $4.4/M output

No provider description is available for this model yet.

dashscope/deepseek-v4-flash 1M context $0.2/M input $0.4/M output

No provider description is available for this model yet.

dashscope/deepseek-v4-flash-0731 1M context $0.2/M input $0.4/M output

No provider description is available for this model yet.

dashscope/deepseek-v4-pro 1M context $2.4/M input $4.8/M output

No provider description is available for this model yet.

dashscope/glm-5.1 202.745K context $1.4/M input $4.4/M output

No provider description is available for this model yet.

dashscope/glm-5.2 1.04858M context $1.4/M input $4.4/M output

No provider description is available for this model yet.

dashscope/kimi-k2.7-code 229.376K context $0.95/M input $4/M output

Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...

qwen/qwen3.5-9b:batch 262.144K context $0.17/M input $0.25/M output

No provider description is available for this model yet.

dashscope/qwen3.8-max 991.808K context $2/M input $6/M output

No provider description is available for this model yet.

vertex_ai/gemini-3.7-flash 1.04858M context $0.75/M input $3.75/M output

No provider description is available for this model yet.

gemini/gemini-3.7-flash 1.04858M context $0.75/M input $3.75/M output

GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

openai/gpt-5.6-sol-pro:batch 1.05M context $1/M input $5/M output

A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

mistralai/ministral-8b-2512 262.144K context $0.15/M input $0.15/M output

No provider description is available for this model yet.

anthropic/claude-3-opus-20240229 200K context $15/M input $75/M output

No provider description is available for this model yet.

groq/qwen/qwen3.6-27b 131.072K context $0.6/M input $3/M output

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost combination.

qwen/qwen-plus-2025-07-28 1M context $0.26/M input $0.78/M output

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,...

z-ai/glm-4.5v 65.536K context $0.6/M input $1.8/M output

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

meta-llama/llama-3.3-70b-instruct:free 65.536K context Free input Free output

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

thinkingmachines/inkling:free 1.04858M context Free input Free output

Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, graphics, and layouts within images.

qwen/qwen2.5-vl-72b-instruct 128K context $0.8/M input $1/M output