Compact GPT model for low-latency assistance and high-volume workloads

openai/gpt-4-turbo 2023-11-06 128K context $10/M input $30/M output
15 providers
Tools

GPT model for general reasoning, writing, coding, and tool-assisted tasks

openai/gpt-4 2023-11-06 8.192K context $30/M input $60/M output
11 providers

Compact GPT model for low-latency assistance and high-volume workloads

openai/gpt-3.5-turbo 2023-03-01 16.385K context $0.5/M input $1.5/M output
13 providers

No provider description is available for this model yet.

zai/glm-5.3-flash 1.04858M context $0.15/M input $0.5/M output

No provider description is available for this model yet.

novita/zai-org/glm-4.7-h 204.8K context $0.6/M input $2.2/M output

No provider description is available for this model yet.

together_ai/zai-org/glm-5.3 1.04858M context $1.4/M input $4.4/M output

No provider description is available for this model yet.

novita/moonshotai/kimi-k2.5 262.144K context $0.6/M input $3/M output

No provider description is available for this model yet.

novita/qwen/qwen3.7-max 1M context $1.25/M input $3.75/M output

No provider description is available for this model yet.

moonshot/kimi-k2.7-code 262.144K context $0.95/M input $4/M output

No provider description is available for this model yet.

gemini/gemma-4-31b-it 262.144K context Input not listed Output not listed

No provider description is available for this model yet.

gemini/gemma-4-26b-a4b-it 262.144K context Input not listed Output not listed

No provider description is available for this model yet.

databricks/databricks-glm-5-3-flash 1.04858M context Input not listed Output not listed

No provider description is available for this model yet.

novita/xiaomimimo/mimo-v2.5 1.04858M context $0.168/M input $0.336/M output

Aion-RP-Llama-3.1-8B ranks the highest in the character evaluation portion of the RPBench-Auto benchmark, a roleplaying-specific variant of Arena-Hard-Auto, where LLMs evaluate each other’s responses. It is a fine-tuned base model...

aion-labs/aion-rp-llama-3.1-8b 32.768K context $0.8/M input $1.6/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3-max 256K context $1.2/M input $6/M output

No provider description is available for this model yet.

novita/deepseek/deepseek-ocr-2 8.192K context $0.03/M input $0.03/M output

No provider description is available for this model yet.

deepinfra/xiaomimimo/mimo-v2.5 262.144K context $0.4/M input $2/M output

No provider description is available for this model yet.

qwencloud/qwen-coder 1M context $0.3/M input $1.5/M output

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless switching between a "thinking" mode for complex reasoning, math, and...

qwen/qwen3-235b-a22b 131.072K context $0.455/M input $1.82/M output

No provider description is available for this model yet.

novita/qwen/qwen3-coder-next 262.144K context $0.2/M input $1.5/M output

No provider description is available for this model yet.

qwencloud/kimi-k2.7-code 229.376K context $0.95/M input $4/M output

GPT-5.1 Chat (AKA Instant is the fast, lightweight member of the 5.1 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on...

openai/gpt-5.1-chat 128K context $1.25/M input $10/M output

No provider description is available for this model yet.

qwencloud/glm-5.2 1.04858M context $1.4/M input $4.4/M output

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

meta/muse-spark-1.3-contributor 1.04858M context $0.1/M input $0.2/M output

No provider description is available for this model yet.

novita/zai-org/glm-5 202.8K context $1/M input $3.2/M output

No provider description is available for this model yet.

qwencloud/qwen-flash 997.952K context Input not listed Output not listed

This model always redirects to the latest GLM model from Z.ai.

~z-ai/glm-latest 1.04858M context $0.873/M input $3.36/M output

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...

z-ai/glm-4.7-flash 131.072K context $0.061/M input $0.4/M output

No provider description is available for this model yet.

qwencloud/qwen-flash-2025-07-28 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-max 30.72K context $1.6/M input $6.4/M output

No provider description is available for this model yet.

qwencloud/qwen-plus 129.024K context $0.4/M input $1.2/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-01-25 129.024K context $0.4/M input $1.2/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-04-28 129.024K context $0.4/M input $1.2/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-07-14 129.024K context $0.4/M input $1.2/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-07-28 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-09-11 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-plus-latest 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-turbo 129.024K context $0.05/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen-turbo-2024-11-01 1M context $0.05/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen-turbo-2025-04-28 1M context $0.05/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen-turbo-latest 1M context $0.05/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen3-30b-a3b 129.024K context Input not listed Output not listed

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

google/gemini-3.6-flash:batch 1.04858M context $0.375/M input $1.875/M output

No provider description is available for this model yet.

qwencloud/glm-5.1 202.745K context $1.4/M input $4.4/M output