1,808 models

Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination.

qwen/qwen-plus 1M context $0.26/M input $0.78/M output

No provider description is available for this model yet.

azure_ai/claude-fable-5 1M context $10/M input $50/M output

No provider description is available for this model yet.

azure_ai/claude-opus-4-7 1M context $5/M input $25/M output

OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...

openai/o3-mini-high:batch 200K context $0.55/M input $2.2/M output

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

z-ai/glm-5v-turbo 202.752K context $1.2/M input $4/M output

The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason...

openai/o1:batch 200K context $7.5/M input $30/M output

This model always redirects to the latest model in the GPT Astra family.

~openai/gpt-astra-latest 1.05M context $10/M input $50/M output

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

nvidia/nemotron-3-ultra-550b-a55b:free 1M context Free input Free output

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

google/gemini-2.5-flash:batch 1.04858M context $0.15/M input $1.25/M output

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

google/gemma-4-31b-it:free 262.144K context Free input Free output

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

google/gemma-4-31b-it:batch 262.144K context $0.39/M input $0.97/M output

This model always redirects to the latest model in the GPT Sol family.

~openai/gpt-sol-latest 1.05M context $2/M input $10/M output

No provider description is available for this model yet.

azure_ai/claude-haiku-4-5 200K context $1/M input $5/M output

No provider description is available for this model yet.

azure_ai/claude-opus-4-5 200K context $5/M input $25/M output

No provider description is available for this model yet.

azure_ai/claude-opus-4-6 1M context $5/M input $25/M output

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,...

minimax/minimax-m2 204.8K context $0.255/M input $1.02/M output

No provider description is available for this model yet.

mistral/labs-leanstral-1-5 262.144K context Input not listed Output not listed

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

anthropic/claude-fable-5.1 1M context $10/M input $50/M output

No provider description is available for this model yet.

azure_ai/claude-opus-5 1M context $5/M input $25/M output

No provider description is available for this model yet.

azure_ai/claude-opus-4-8 1M context $5/M input $25/M output

No provider description is available for this model yet.

azure_ai/claude-opus-4-1 200K context $15/M input $75/M output

No provider description is available for this model yet.

anthropic/claude-mythos-preview 1M context $10/M input $50/M output

No provider description is available for this model yet.

azure_ai/claude-sonnet-4-5 200K context $3/M input $15/M output

No provider description is available for this model yet.

xai/grok-4.20-multi-agent-0309 1M context $1.25/M input $2.5/M output

No provider description is available for this model yet.

azure_ai/claude-sonnet-4-6 1M context $3/M input $15/M output

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

thinkingmachines/inkling-small:free 1.04858M context Free input Free output

No provider description is available for this model yet.

vertex_ai/gemini-3.8-flash 1.04858M context $0.75/M input $3.75/M output

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...

moonshotai/kimi-k2.7-code:batch 262.144K context $0.95/M input $4/M output

No provider description is available for this model yet.

azure_ai/gpt-5.5-2026-04-23 1.05M context $5/M input $30/M output

UnslopNemo v4.1 is the latest addition from the creator of Rocinante, designed for adventure writing and role-play scenarios.

thedrummer/unslopnemo-12b 1.024M context $0.4/M input $0.4/M output