3,677 models

No provider description is available for this model yet.

fireworks_ai/deepseek-v4-flash-0731 1.04858M context $0.14/M input $0.28/M output

No provider description is available for this model yet.

gmi/moonshotai/kimi-k2-thinking 262.144K context $0.8/M input $1.2/M output

No provider description is available for this model yet.

gmi/minimaxai/minimax-m2.1 196.608K context $0.3/M input $1.2/M output

No provider description is available for this model yet.

baseten/minimaxai/minimax-m2.5 Not documented context $0.3/M input $1.2/M output

No provider description is available for this model yet.

baseten/nvidia/nemotron-120b-a12b Not documented context $0.3/M input $0.75/M output

No provider description is available for this model yet.

baseten/zai-org/glm-5 Not documented context $0.95/M input $3.15/M output

No provider description is available for this model yet.

baseten/zai-org/glm-4.7 Not documented context $0.6/M input $2.2/M output

No provider description is available for this model yet.

baseten/zai-org/glm-4.6 Not documented context $0.6/M input $2.2/M output

No provider description is available for this model yet.

baseten/moonshotai/kimi-k2.5 Not documented context $0.6/M input $3/M output

No provider description is available for this model yet.

baseten/moonshotai/kimi-k2-thinking Not documented context $0.6/M input $2.5/M output

Coder‑Large is a 32 B‑parameter offspring of Qwen 2.5‑Instruct that has been further trained on permissively‑licensed GitHub, CodeSearchNet and synthetic bug‑fix corpora. It supports a 32k context window, enabling multi‑file...

arcee-ai/coder-large 32.768K context $0.5/M input $0.8/M output

No provider description is available for this model yet.

baseten/openai/gpt-oss-120b Not documented context $0.1/M input $0.5/M output

No provider description is available for this model yet.

azure_ai/fw-inkling 1.04858M context $1/M input $4.05/M output

No provider description is available for this model yet.

openai/chatgpt-4o-latest 128K context $5/M input $15/M output

This model always redirects to the latest model in the GLM Flash family.

~z-ai/glm-flash-latest 1.04858M context $0.075/M input $0.25/M output

LFM2.5-2.6B is a compact reasoning model from Liquid AI. It is suited for agent workflows, data extraction, RAG, and long-context processing. Liquid advises against using it for agentic coding or...

liquid/lfm-2.5-2.6b:free 65.536K context Free input Free output

No provider description is available for this model yet.

baseten/deepseek-ai/deepseek-v3.1 Not documented context $0.5/M input $1.5/M output

No provider description is available for this model yet.

azure_ai/fw-glm-5.2-fast 1.04858M context $2.1/M input $6.6/M output

A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million...

openai/gpt-audio-mini 128K context $0.6/M input $2.4/M output

No provider description is available for this model yet.

gmi/zai-org/glm-4.7-fp8 202.752K context $0.4/M input $2/M output

No provider description is available for this model yet.

azure/gpt-audio-mini 128K context $0.6/M input $2.4/M output

No provider description is available for this model yet.

scx-ai/glm-5.2 1.04858M context $0.61/M input $1.98/M output

No provider description is available for this model yet.

scx-ai/qwen3.8-max 1M context $1.65/M input $4.99/M output

No provider description is available for this model yet.

together_ai/together-ai-8.1b-21b 1K context $0.3/M input $0.3/M output

No provider description is available for this model yet.

openrouter/x-ai/grok-4.20 1M context $1.25/M input $2.5/M output

No provider description is available for this model yet.

together_ai/together-ai-up-to-4b Not documented context $0.1/M input $0.1/M output

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

qwen/qwen3-coder:free 262K context Free input Free output