Virtuoso‑Large is Arcee's top‑tier general‑purpose LLM at 72 B parameters, tuned to tackle cross‑domain reasoning, creative writing and enterprise QA. Unlike many 70 B peers, it retains the 128 k...

arcee-ai/virtuoso-large 131.072K context $0.75/M input $1.2/M output

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

nvidia/nemotron-3.5-lightning:free 1M context Free input Free output

No provider description is available for this model yet.

azure_ai/claude-fable-5-1 1M context $10/M input $50/M output

No provider description is available for this model yet.

azure_ai/deepseek-v4-flash-0731 1M context $0.19/M input $0.51/M output

No provider description is available for this model yet.

qwencloud/deepseek-v4-flash 1M context $0.2/M input $0.4/M output

No provider description is available for this model yet.

qwencloud/deepseek-v4-flash-0731 1M context $0.2/M input $0.4/M output

No provider description is available for this model yet.

moonshot/kimi-k3 1.04858M context $3/M input $15/M output

No provider description is available for this model yet.

qwencloud/deepseek-v4-pro 1M context $2.4/M input $4.8/M output

This model always redirects to the latest model in the GPT Terra family.

~openai/gpt-terra-latest 1.05M context $2/M input $12/M output

No provider description is available for this model yet.

qwencloud/glm-5.1 202.745K context $1.4/M input $4.4/M output

No provider description is available for this model yet.

qwencloud/glm-5.2 1.04858M context $1.4/M input $4.4/M output

No provider description is available for this model yet.

qwencloud/kimi-k2.7-code 229.376K context $0.95/M input $4/M output

No provider description is available for this model yet.

qwencloud/qwen-coder 1M context $0.3/M input $1.5/M output

No provider description is available for this model yet.

qwencloud/qwen-flash 997.952K context Input not listed Output not listed

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

meta/muse-spark-1.3-contributor 1.04858M context $0.1/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen-flash-2025-07-28 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-max 30.72K context $1.6/M input $6.4/M output

No provider description is available for this model yet.

qwencloud/qwen-plus 129.024K context $0.4/M input $1.2/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-01-25 129.024K context $0.4/M input $1.2/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-04-28 129.024K context $0.4/M input $1.2/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-07-14 129.024K context $0.4/M input $1.2/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-07-28 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-09-11 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-plus-latest 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-turbo 129.024K context $0.05/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen-turbo-2024-11-01 1M context $0.05/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen-turbo-2025-04-28 1M context $0.05/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen-turbo-latest 1M context $0.05/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen3-30b-a3b 129.024K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen3-coder-flash 997.952K context Input not listed Output not listed

This model always redirects to the latest GLM model from Z.ai.

~z-ai/glm-latest 262.144K context $0.877/M input $2.97/M output

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

qwen/qwen3.8-2.4t-a95b 1M context $2/M input $6/M output

No provider description is available for this model yet.

qwencloud/qwen3-coder-plus 997.952K context Input not listed Output not listed

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to think harder and provide consistently...

openai/o3-pro:batch 200K context $10/M input $40/M output

No provider description is available for this model yet.

qwencloud/qwen3-coder-plus-2025-07-22 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen3-max-preview 258.048K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen3-max 258.048K context Input not listed Output not listed

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...

openai/gpt-oss-120b:batch 131.072K context $0.15/M input $0.6/M output

No provider description is available for this model yet.

qwencloud/qwen3-max-2026-01-23 258.048K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen3-vl-32b-instruct 131.072K context $0.16/M input $0.64/M output

No provider description is available for this model yet.

qwencloud/qwen3-vl-32b-thinking 131.072K context $0.16/M input $2.87/M output