3,311 models

No provider description is available for this model yet.

deepinfra/qwen/qwen3.5-9b 262.144K context $0.1/M input $0.15/M output

No provider description is available for this model yet.

deepinfra/zai-org/glm-4.7 202.752K context $0.4/M input $1.75/M output

No provider description is available for this model yet.

deepinfra/google/gemma-4-31b-it 262.144K context $0.13/M input $0.38/M output

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...

openai/gpt-5.4-mini:batch 400K context $0.375/M input $2.25/M output

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

nvidia/nemotron-3-ultra-550b-a55b:batch 512.288K context $0.6/M input $3.6/M output

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...

z-ai/glm-4.5-air 131.072K context $0.13/M input $0.85/M output

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

google/gemini-2.5-flash:batch 1.04858M context $0.15/M input $1.25/M output

Virtuoso‑Large is Arcee's top‑tier general‑purpose LLM at 72 B parameters, tuned to tackle cross‑domain reasoning, creative writing and enterprise QA. Unlike many 70 B peers, it retains the 128 k...

arcee-ai/virtuoso-large 131.072K context $0.75/M input $1.2/M output

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

qwen/qwen3.8-2.4t-a95b 1M context $2/M input $6/M output

No provider description is available for this model yet.

dashscope/qwen3-max-2026-01-23 258.048K context Input not listed Output not listed

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

google/gemini-3.1-pro-preview:batch 1.04858M context $1/M input $6/M output

No provider description is available for this model yet.

dashscope/qwen3-max 258.048K context Input not listed Output not listed

No provider description is available for this model yet.

dashscope/qwen3-max-preview 258.048K context Input not listed Output not listed

Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information...

x-ai/grok-4.20-multi-agent 2M context $1.25/M input $2.5/M output

No provider description is available for this model yet.

azure_ai/claude-fable-5-1 1M context $10/M input $50/M output

No provider description is available for this model yet.

azure_ai/deepseek-v4-flash-0731 1M context $0.19/M input $0.51/M output

No provider description is available for this model yet.

qwencloud/deepseek-v4-flash 1M context $0.2/M input $0.4/M output

No provider description is available for this model yet.

qwencloud/deepseek-v4-flash-0731 1M context $0.2/M input $0.4/M output

No provider description is available for this model yet.

moonshot/kimi-k3 1.04858M context $3/M input $15/M output

No provider description is available for this model yet.

qwencloud/deepseek-v4-pro 1M context $2.4/M input $4.8/M output

No provider description is available for this model yet.

friendliai/google/gemma-4-31b-it 262.144K context $0.14/M input $0.4/M output

This model always redirects to the latest model in the GPT Terra family.

~openai/gpt-terra-latest 1.05M context $2/M input $12/M output

No provider description is available for this model yet.

qwencloud/glm-5.1 202.745K context $1.4/M input $4.4/M output

No provider description is available for this model yet.

qwencloud/glm-5.2 1.04858M context $1.4/M input $4.4/M output

No provider description is available for this model yet.

qwencloud/kimi-k2.7-code 229.376K context $0.95/M input $4/M output

No provider description is available for this model yet.

qwencloud/qwen-coder 1M context $0.3/M input $1.5/M output

No provider description is available for this model yet.

qwencloud/qwen-flash 997.952K context Input not listed Output not listed

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

meta/muse-spark-1.3-contributor 1.04858M context $0.1/M input $0.2/M output

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,...

z-ai/glm-4.5v 65.536K context $0.6/M input $1.8/M output

No provider description is available for this model yet.

qwencloud/qwen-flash-2025-07-28 997.952K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-plus 129.024K context $0.4/M input $1.2/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-01-25 129.024K context $0.4/M input $1.2/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-04-28 129.024K context $0.4/M input $1.2/M output

No provider description is available for this model yet.

qwencloud/qwen-plus-2025-07-14 129.024K context $0.4/M input $1.2/M output