4,182 models

No provider description is available for this model yet.

fireworks_ai/glm-5p2-fast 1.04858M context $2.1/M input $6.6/M output
Open weights

No provider description is available for this model yet.

meta-llama/Llama-2-13b-chat-hf Not documented context Input not listed Output not listed
Open weights

No provider description is available for this model yet.

meta-llama/Llama-2-7b-chat-hf Not documented context Input not listed Output not listed

No provider description is available for this model yet.

fireworks_ai/deepseek-v4-flash-0731 1.04858M context $0.14/M input $0.28/M output

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...

mistralai/mistral-medium-3.1:batch 131.072K context $0.2/M input $1/M output

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....

openai/gpt-5-mini:batch 400K context $0.125/M input $1/M output

Fast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

anthropic/claude-opus-4.7-fast 1M context $30/M input $150/M output

No provider description is available for this model yet.

azure_ai/fw-inkling 1.04858M context $1/M input $4.05/M output

No provider description is available for this model yet.

azure_ai/fw-glm-5.2-fast 1.04858M context $2.1/M input $6.6/M output

No provider description is available for this model yet.

bedrock/twelvelabs.pegasus-1-2-v1:0 Not documented context Input not listed $7.5/M output

Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired with natural language queries, and produces detailed visual understanding...

perceptron/perceptron-mk1 32.768K context $0.15/M input $1.5/M output

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)

mistralai/codestral-2508:batch 256K context $0.15/M input $0.45/M output

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

qwen/qwen3-coder 262.144K context $0.3/M input $1/M output

No provider description is available for this model yet.

bedrock/eu.twelvelabs.pegasus-1-2-v1:0 Not documented context Input not listed $7.5/M output
Open weights

No provider description is available for this model yet.

meta-llama/Llama-2-13b-hf Not documented context Input not listed Output not listed

Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in...

anthropic/claude-opus-4 200K context $15/M input $75/M output
Open weights

No provider description is available for this model yet.

meta-llama/Llama-2-7b-hf Not documented context Input not listed Output not listed

No provider description is available for this model yet.

bedrock/amazon.titan-text-lite-v1 42K context $0.3/M input $0.4/M output

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.

openai/gpt-3.5-turbo:batch 16.385K context $0.25/M input $0.75/M output

No provider description is available for this model yet.

gmi/deepseek-ai/deepseek-v3.2 163.84K context $0.28/M input $0.4/M output

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

google/gemini-3.1-flash-lite:batch 1.04858M context $0.125/M input $0.75/M output
Open weights

No provider description is available for this model yet.

meta-llama/Llama-2-70b-chat Not documented context Input not listed Output not listed

The o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o1-pro model uses more compute to think harder and provide...

openai/o1-pro:batch 200K context $75/M input $300/M output
Open weights

No provider description is available for this model yet.

meta-llama/Llama-2-13b-chat Not documented context Input not listed Output not listed
Open weights

No provider description is available for this model yet.

meta-llama/Llama-2-7b-chat Not documented context Input not listed Output not listed

Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances state-of-the-art reasoning and multimodal performance with 8× lower cost...

mistralai/mistral-medium-3 131.072K context $0.4/M input $2/M output

No provider description is available for this model yet.

azure_ai/fw-glm-5.2 1.04858M context $1.54/M input $4.84/M output

Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fetch enabled, then a...

openrouter/fusion 1M context Input not listed Output not listed

No provider description is available for this model yet.

openrouter/openai/gpt-5.4-pro 1.05M context $30/M input $180/M output
Open weights

No provider description is available for this model yet.

meta-llama/Llama-2-70b Not documented context Input not listed Output not listed

Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided translation.

tencent/hy-mt2-7b 8.192K context $0.074/M input $0.295/M output

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...

openai/gpt-5-nano:batch 400K context $0.025/M input $0.2/M output

No provider description is available for this model yet.

azure_ai/fw-glm-5.1 202.8K context $1.54/M input $4.84/M output