No provider description is available for this model yet.

deepinfra/google/gemma-4-26b-a4b-it 262.144K context $0.07/M input $0.34/M output

No provider description is available for this model yet.

deepinfra/google/gemini-3.1-pro 1M context $2/M input $12/M output

No provider description is available for this model yet.

deepinfra/openai/gpt-oss-120b-ultra 131.072K context $0.2/M input $0.95/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.5-9b 262.144K context $0.1/M input $0.15/M output

No provider description is available for this model yet.

deepinfra/zai-org/glm-4.7 202.752K context $0.4/M input $1.75/M output

No provider description is available for this model yet.

deepinfra/google/gemma-4-31b-it 262.144K context $0.13/M input $0.38/M output

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

nvidia/nemotron-3-ultra-550b-a55b:batch 512.288K context $0.6/M input $3.6/M output

Virtuoso‑Large is Arcee's top‑tier general‑purpose LLM at 72 B parameters, tuned to tackle cross‑domain reasoning, creative writing and enterprise QA. Unlike many 70 B peers, it retains the 128 k...

arcee-ai/virtuoso-large 131.072K context $0.75/M input $1.2/M output

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

nvidia/nemotron-3.5-lightning:free 1M context Free input Free output

OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model supports the `reasoning_effort` parameter, which can be set to...

openai/o3-mini:batch 200K context $0.55/M input $2.2/M output

No provider description is available for this model yet.

azure_ai/claude-fable-5-1 1M context $10/M input $50/M output

No provider description is available for this model yet.

azure_ai/deepseek-v4-flash-0731 1M context $0.19/M input $0.51/M output

No provider description is available for this model yet.

qwencloud/deepseek-v4-flash 1M context $0.2/M input $0.4/M output

No provider description is available for this model yet.

qwencloud/deepseek-v4-flash-0731 1M context $0.2/M input $0.4/M output

No provider description is available for this model yet.

moonshot/kimi-k3 1.04858M context $3/M input $15/M output

No provider description is available for this model yet.

qwencloud/deepseek-v4-pro 1M context $2.4/M input $4.8/M output

No provider description is available for this model yet.

qwencloud/glm-5.1 202.745K context $1.4/M input $4.4/M output

No provider description is available for this model yet.

qwencloud/glm-5.2 1.04858M context $1.4/M input $4.4/M output

This model always redirects to the latest model in the GPT Terra family.

~openai/gpt-terra-latest 1.05M context $2/M input $12/M output

No provider description is available for this model yet.

qwencloud/kimi-k2.7-code 229.376K context $0.95/M input $4/M output

No provider description is available for this model yet.

qwencloud/qwen-coder 1M context $0.3/M input $1.5/M output

No provider description is available for this model yet.

qwencloud/qwen-flash 997.952K context Input not listed Output not listed

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

meta/muse-spark-1.3-contributor 1.04858M context $0.1/M input $0.2/M output

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

anthropic/claude-opus-4.8:batch 1M context $2.5/M input $12.5/M output

No provider description is available for this model yet.

qwencloud/qwen-flash-2025-07-28 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-max 30.72K context $1.6/M input $6.4/M output

No provider description is available for this model yet.

qwencloud/qwen-plus 129.024K context $0.4/M input $1.2/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-01-25 129.024K context $0.4/M input $1.2/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-04-28 129.024K context $0.4/M input $1.2/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-07-14 129.024K context $0.4/M input $1.2/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-07-28 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-09-11 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-plus-latest 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-turbo 129.024K context $0.05/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen-turbo-2024-11-01 1M context $0.05/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen-turbo-2025-04-28 1M context $0.05/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen-turbo-latest 1M context $0.05/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen3-30b-a3b 129.024K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen3-coder-flash 997.952K context Input not listed Output not listed

This model always redirects to the latest GLM model from Z.ai.

~z-ai/glm-latest 1.04858M context $0.9/M input $3/M output

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

qwen/qwen3.8-2.4t-a95b 1M context $2/M input $6/M output