4,213 models

No provider description is available for this model yet.

baseten/zai-org/glm-4.6 Not documented context $0.6/M input $2.2/M output

ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with 47B active per token. It is trained jointly on text and image data...

baidu/ernie-4.5-vl-424b-a47b 123K context $0.42/M input $1.25/M output
Open weights

No provider description is available for this model yet.

google/gemma-4-E2B Not documented context Input not listed Output not listed

No provider description is available for this model yet.

baseten/zai-org/glm-4.7 Not documented context $0.6/M input $2.2/M output

No provider description is available for this model yet.

baseten/zai-org/glm-5 Not documented context $0.95/M input $3.15/M output
Open weights

No provider description is available for this model yet.

google/gemma-4-E4B Not documented context Input not listed Output not listed
Open weights

No provider description is available for this model yet.

google/gemma-4-26B-A4B Not documented context Input not listed Output not listed

No provider description is available for this model yet.

baseten/nvidia/nemotron-120b-a12b Not documented context $0.3/M input $0.75/M output

No provider description is available for this model yet.

baseten/minimaxai/minimax-m2.5 Not documented context $0.3/M input $1.2/M output

Ox Alpha is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon software engineering, complex reasoning, and workflows that combine text with...

stealth/ox-alpha 1.04858M context Input not listed Output not listed

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...

qwen/qwen3-235b-a22b-2507 262.144K context $0.087/M input $0.35/M output

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

google/gemini-2.5-pro-preview 1.04858M context $1.25/M input $10/M output

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...

qwen/qwen3.6-27b 262.144K context $0.3/M input $2/M output

Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...

qwen/qwen-2.5-72b-instruct 32.768K context $0.36/M input $0.4/M output
Open weights

No provider description is available for this model yet.

google/gemma-4-31B Not documented context Input not listed Output not listed
Open weights

No provider description is available for this model yet.

google/gemma-4-31B-it Not documented context Input not listed Output not listed

No provider description is available for this model yet.

gmi/minimaxai/minimax-m2.1 196.608K context $0.3/M input $1.2/M output

No provider description is available for this model yet.

gmi/moonshotai/kimi-k2-thinking 262.144K context $0.8/M input $1.2/M output

No provider description is available for this model yet.

friendliai/zai-org/glm-5.2 1.04858M context $1.4/M input $4.4/M output

No provider description is available for this model yet.

friendliai/google/gemma-4-31b-it 262.144K context $0.14/M input $0.4/M output

GPT-4o Search Previewis a specialized model for web search in Chat Completions. It is trained to understand and execute web search queries.

openai/gpt-4o-search-preview 128K context $2.5/M input $10/M output

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the...

qwen/qwen3-coder-30b-a3b-instruct 262.144K context $0.07/M input $0.28/M output

No provider description is available for this model yet.

cerebras/qwen-3.8-27b 65.536K context $0.99/M input $1.49/M output

No provider description is available for this model yet.

gmi/google/gemini-3-flash-preview 1.04858M context $0.5/M input $3/M output

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for...

openai/gpt-5.5-pro:batch 1.05M context $15/M input $90/M output

No provider description is available for this model yet.

openai/computer-use-preview 8.192K context $3/M input $12/M output

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous programming via tool calling and...

qwen/qwen3-coder-plus 1M context $0.65/M input $3.25/M output

No provider description is available for this model yet.

bedrock/anthropic.claude-v1 100K context $8/M input $24/M output

No provider description is available for this model yet.

bedrock/anthropic.claude-v2:1 100K context $8/M input $24/M output

No provider description is available for this model yet.

gmi/google/gemini-3-pro-preview 1.04858M context $2/M input $12/M output