1,829 models

No provider description is available for this model yet.

deepinfra/qwen/qwen3.5-122b-a10b 262.144K context $0.29/M input $2.4/M output

No provider description is available for this model yet.

deepinfra/zai-org/glm-5.1 202.752K context $1.05/M input $3.5/M output

No provider description is available for this model yet.

deepinfra/zai-org/glm-5.2 1.04858M context $0.75/M input $2.4/M output

No provider description is available for this model yet.

deepinfra/moonshotai/kimi-k3 1.04858M context $2.85/M input $14.25/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.6-27b 262.144K context $0.32/M input $3.2/M output

No provider description is available for this model yet.

deepinfra/google/gemma-4-26b-a4b-it 262.144K context $0.07/M input $0.34/M output

No provider description is available for this model yet.

deepinfra/google/gemini-3.1-pro 1M context $2/M input $12/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.5-9b 262.144K context $0.1/M input $0.15/M output

No provider description is available for this model yet.

deepinfra/zai-org/glm-4.7 202.752K context $0.4/M input $1.75/M output

No provider description is available for this model yet.

deepinfra/google/gemma-4-31b-it 262.144K context $0.13/M input $0.38/M output

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

nvidia/nemotron-3-ultra-550b-a55b:batch 512.288K context $0.6/M input $3.6/M output

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

google/gemini-2.5-flash:batch 1.04858M context $0.15/M input $1.25/M output

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)

mistralai/codestral-2508 256K context $0.3/M input $0.9/M output

No provider description is available for this model yet.

azure_ai/claude-fable-5-1 1M context $10/M input $50/M output

No provider description is available for this model yet.

azure_ai/deepseek-v4-flash-0731 1M context $0.19/M input $0.51/M output

No provider description is available for this model yet.

qwencloud/deepseek-v4-flash 1M context $0.2/M input $0.4/M output

No provider description is available for this model yet.

qwencloud/deepseek-v4-flash-0731 1M context $0.2/M input $0.4/M output

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...

anthropic/claude-opus-4.1:batch 200K context $7.5/M input $37.5/M output

No provider description is available for this model yet.

moonshot/kimi-k3 1.04858M context $3/M input $15/M output

No provider description is available for this model yet.

qwencloud/deepseek-v4-pro 1M context $2.4/M input $4.8/M output

OpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high. OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining...

openai/o4-mini-high 200K context $1.1/M input $4.4/M output

This model always redirects to the latest model in the GPT Terra family.

~openai/gpt-terra-latest 1.05M context $2/M input $12/M output

No provider description is available for this model yet.

qwencloud/glm-5.1 202.745K context $1.4/M input $4.4/M output

No provider description is available for this model yet.

qwencloud/glm-5.2 1.04858M context $1.4/M input $4.4/M output

No provider description is available for this model yet.

qwencloud/kimi-k2.7-code 229.376K context $0.95/M input $4/M output

No provider description is available for this model yet.

qwencloud/qwen-coder 1M context $0.3/M input $1.5/M output

No provider description is available for this model yet.

qwencloud/qwen-flash 997.952K context Input not listed Output not listed

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

meta/muse-spark-1.3-contributor 1.04858M context $0.1/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen-flash-2025-07-28 997.952K context Input not listed Output not listed

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....

openai/o3:batch 200K context $1/M input $4/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-07-28 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-09-11 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-plus-latest 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-turbo-2024-11-01 1M context $0.05/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen-turbo-2025-04-28 1M context $0.05/M input $0.2/M output

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

moonshotai/kimi-k3:batch 1.04858M context $3/M input $15/M output

No provider description is available for this model yet.

qwencloud/qwen-turbo-latest 1M context $0.05/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen3-coder-flash 997.952K context Input not listed Output not listed