No provider description is available for this model yet.

azure/o1-2024-12-17 200K context $15/M input $60/M output

No provider description is available for this model yet.

azure/o1-mini 128K context $1.21/M input $4.84/M output

No provider description is available for this model yet.

azure/o1-preview 128K context $15/M input $60/M output

No provider description is available for this model yet.

azure/o1-preview-2024-09-12 128K context $15/M input $60/M output

No provider description is available for this model yet.

cerebras/llama3.1-8b 128K context $0.1/M input $0.1/M output

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

qwen/qwen3.8-27b 262.144K context $0.214/M input $2.55/M output

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous programming via tool calling and...

qwen/qwen3-coder-plus 1M context $0.65/M input $3.25/M output

No provider description is available for this model yet.

azure/o3 200K context $2/M input $8/M output

The preview GPT-4 model with improved instruction following, JSON mode, reproducible outputs, parallel function calling, and more. Training data: up to Dec 2023. **Note:** heavily rate limited by OpenAI while...

openai/gpt-4-turbo-preview 128K context $10/M input $30/M output

No provider description is available for this model yet.

azure/o3-2025-04-16 200K context $2/M input $8/M output

No provider description is available for this model yet.

azure/o3-mini 200K context $1.1/M input $4.4/M output

No provider description is available for this model yet.

azure/o3-mini-2025-01-31 200K context $1.1/M input $4.4/M output

No provider description is available for this model yet.

cerebras/gpt-oss-120b 131.072K context $0.35/M input $0.75/M output

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

z-ai/glm-5.2:free 256K context Free input Free output

No provider description is available for this model yet.

azure/o4-mini 200K context $1.1/M input $4.4/M output

No provider description is available for this model yet.

azure/o4-mini-2025-04-16 200K context $1.1/M input $4.4/M output

No provider description is available for this model yet.

azure/us/gpt-4.1-2025-04-14 1.04758M context $2.2/M input $8.8/M output

No provider description is available for this model yet.

azure/us/gpt-4.1-mini-2025-04-14 1.04758M context $0.44/M input $1.76/M output

No provider description is available for this model yet.

azure/us/gpt-4.1-nano-2025-04-14 1.04758M context $0.11/M input $0.44/M output

No provider description is available for this model yet.

azure/us/gpt-4o-2024-08-06 128K context $2.75/M input $11/M output

No provider description is available for this model yet.

azure/us/gpt-4o-2024-11-20 128K context $2.75/M input $11/M output

No provider description is available for this model yet.

mistral/glm-5-2 1.04858M context $1.4/M input $4.4/M output

No provider description is available for this model yet.

azure/us/gpt-5-2025-08-07 272K context $1.375/M input $11/M output

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost combination.

qwen/qwen-plus-2025-07-28 1M context $0.26/M input $0.78/M output

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

meta-llama/llama-3.3-70b-instruct:free 65.536K context Free input Free output

No provider description is available for this model yet.

azure/us/gpt-5-nano-2025-08-07 272K context $0.055/M input $0.44/M output

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

minimax/minimax-m3 524.288K context $0.3/M input $1.2/M output

No provider description is available for this model yet.

azure/us/gpt-5.1-chat 128K context $1.38/M input $11/M output

No provider description is available for this model yet.

azure/us/o1-2024-12-17 200K context $16.5/M input $66/M output

No provider description is available for this model yet.

azure/us/o1-mini-2024-09-12 128K context $1.21/M input $4.84/M output

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

qwen/qwen3.6-plus 1M context $0.325/M input $1.95/M output

No provider description is available for this model yet.

azure/us/o1-preview-2024-09-12 128K context $16.5/M input $66/M output

No provider description is available for this model yet.

cerebras/qwen-3-32b 128K context $0.4/M input $0.8/M output

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

google/gemini-2.5-pro:batch 1.04858M context $0.625/M input $5/M output

No provider description is available for this model yet.

mistral/zai-glm-5-2 1.04858M context $1.4/M input $4.4/M output

Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing in autonomous programming via tool calling...

qwen/qwen3-coder-flash 1M context $0.195/M input $0.975/M output

No provider description is available for this model yet.

azure/us/o3-mini-2025-01-31 200K context $1.21/M input $4.84/M output

No provider description is available for this model yet.

azure/us/o4-mini-2025-04-16 200K context $1.21/M input $4.84/M output

No provider description is available for this model yet.

azure_ai/llama-3.3-70b-instruct 128K context $0.71/M input $0.71/M output

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...

qwen/qwen3-vl-30b-a3b-instruct 262.144K context $0.15/M input $0.6/M output

Fast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

anthropic/claude-opus-4.7-fast 1M context $30/M input $150/M output