GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

z-ai/glm-5.2:free 256K context Free input Free output

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost combination.

qwen/qwen-plus-2025-07-28 1M context $0.26/M input $0.78/M output

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

meta-llama/llama-3.3-70b-instruct:free 65.536K context Free input Free output

No provider description is available for this model yet.

azure/us/gpt-5-mini-2025-08-07 272K context $0.275/M input $2.2/M output

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

minimax/minimax-m3 524.288K context $0.3/M input $1.2/M output

No provider description is available for this model yet.

azure/us/gpt-5-nano-2025-08-07 272K context $0.055/M input $0.44/M output

No provider description is available for this model yet.

azure/us/gpt-5.1 272K context $1.38/M input $11/M output

Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing in autonomous programming via tool calling...

qwen/qwen3-coder-flash 1M context $0.195/M input $0.975/M output

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...

qwen/qwen3-vl-30b-a3b-instruct 262.144K context $0.15/M input $0.6/M output

No provider description is available for this model yet.

azure/us/gpt-5.1-chat 128K context $1.38/M input $11/M output

No provider description is available for this model yet.

azure/us/o1-2024-12-17 200K context $16.5/M input $66/M output

No provider description is available for this model yet.

azure/us/o1-mini-2024-09-12 128K context $1.21/M input $4.84/M output

Fast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

anthropic/claude-opus-4.7-fast 1M context $30/M input $150/M output

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

qwen/qwen3.6-plus 1M context $0.325/M input $1.95/M output

GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

openai/gpt-5.6-terra-pro 1.05M context $2/M input $12/M output

No provider description is available for this model yet.

azure/us/o1-preview-2024-09-12 128K context $16.5/M input $66/M output

No provider description is available for this model yet.

cerebras/qwen-3-32b 128K context $0.4/M input $0.8/M output

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

thinkingmachines/inkling-small:batch 524.288K context $0.5/M input $1.2/M output

Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a 256K context window.

bytedance-seed/seed-1.6 262.144K context $0.25/M input $2/M output

No provider description is available for this model yet.

azure/us/o3-2025-04-16 200K context $2.2/M input $8.8/M output

No provider description is available for this model yet.

azure/us/o3-mini-2025-01-31 200K context $1.21/M input $4.84/M output

No provider description is available for this model yet.

azure/us/o4-mini-2025-04-16 200K context $1.21/M input $4.84/M output

No provider description is available for this model yet.

azure_ai/llama-3.3-70b-instruct 128K context $0.71/M input $0.71/M output

No provider description is available for this model yet.

azure_ai/meta-llama-3-70b-instruct 8.192K context $1.1/M input $0.37/M output

No provider description is available for this model yet.

azure_ai/phi-3-mini-128k-instruct 128K context $0.13/M input $0.52/M output

No provider description is available for this model yet.

azure_ai/phi-3-small-8k-instruct 8.192K context $0.15/M input $0.6/M output

No provider description is available for this model yet.

azure_ai/phi-3.5-moe-instruct 128K context $0.16/M input $0.64/M output

No provider description is available for this model yet.

azure_ai/phi-3.5-mini-instruct 128K context $0.13/M input $0.52/M output

No provider description is available for this model yet.

groq/qwen/qwen3.6-27b 131.072K context $0.6/M input $3/M output

No provider description is available for this model yet.

azure_ai/phi-4 16.384K context $0.125/M input $0.5/M output

No provider description is available for this model yet.

azure_ai/phi-4-mini-instruct 131.072K context $0.075/M input $0.3/M output

No provider description is available for this model yet.

azure_ai/phi-4-multimodal-instruct 131.072K context $0.08/M input $0.32/M output

No provider description is available for this model yet.

openai/daybreak-blue-latest 1.05M context $5/M input $30/M output

The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...

mistralai/ministral-14b-2512 262.144K context $0.2/M input $0.2/M output