1,128 models

No provider description is available for this model yet.

mistral/mistral-vibe-cli-fast 262.144K context $0.15/M input $0.6/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.5-397b-a17b 262.144K context $0.45/M input $3/M output

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

anthropic/claude-fable-5.1:batch 1M context $5/M input $25/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.5-122b-a10b 262.144K context $0.29/M input $2.4/M output

No provider description is available for this model yet.

deepinfra/moonshotai/kimi-k3 1.04858M context $2.85/M input $14.25/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.6-27b 262.144K context $0.32/M input $3.2/M output

No provider description is available for this model yet.

deepinfra/google/gemma-4-26b-a4b-it 262.144K context $0.07/M input $0.34/M output

No provider description is available for this model yet.

deepinfra/google/gemini-3.1-pro 1M context $2/M input $12/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.5-9b 262.144K context $0.1/M input $0.15/M output

No provider description is available for this model yet.

deepinfra/google/gemma-4-31b-it 262.144K context $0.13/M input $0.38/M output

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...

google/gemini-3-flash-preview:batch 1.04858M context $0.25/M input $1.5/M output

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

google/gemini-2.5-flash:batch 1.04858M context $0.15/M input $1.25/M output

No provider description is available for this model yet.

azure_ai/claude-fable-5-1 1M context $10/M input $50/M output

No provider description is available for this model yet.

moonshot/kimi-k3 1.04858M context $3/M input $15/M output

This model always redirects to the latest model in the GPT Terra family.

~openai/gpt-terra-latest 1.05M context $2/M input $12/M output

No provider description is available for this model yet.

qwencloud/kimi-k2.7-code 229.376K context $0.95/M input $4/M output

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

meta/muse-spark-1.3-contributor 1.04858M context $0.1/M input $0.2/M output

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

moonshotai/kimi-k3:batch 1.04858M context $3/M input $15/M output

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to think harder and provide consistently...

openai/o3-pro:batch 200K context $10/M input $40/M output

No provider description is available for this model yet.

qwencloud/qwen3-vl-plus 260.096K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen3.5-plus 991.808K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen3.7-plus 991.808K context Input not listed Output not listed

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

qwen/qwen3.8-flash 1M context $0.15/M input $0.47/M output

No provider description is available for this model yet.

qwencloud/qwen3.8-max 991.808K context $2/M input $6/M output

No provider description is available for this model yet.

qwen_ai_platform/kimi-k2.7-code 229.376K context $0.95/M input $4/M output

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

minimax/minimax-m3:batch 524.288K context $0.3/M input $1.2/M output

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

openai/gpt-6-astra:batch 1.05M context $5/M input $25/M output

No provider description is available for this model yet.

qwen_ai_platform/qwen3-vl-plus 260.096K context Input not listed Output not listed

No provider description is available for this model yet.

qwen_ai_platform/qwen3.5-plus 991.808K context Input not listed Output not listed

No provider description is available for this model yet.

qwen_ai_platform/qwen3.7-plus 991.808K context Input not listed Output not listed

No provider description is available for this model yet.

qwen_ai_platform/qwen3.8-max 991.808K context $2/M input $6/M output

No provider description is available for this model yet.

friendliai/zai-org/glm-5.3-flash 1.04858M context $0.15/M input $0.5/M output

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

openai/gpt-5.6-terra:batch 1.05M context $1/M input $6/M output

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

minimax/minimax-m3:free 1.04858M context Free input Free output

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

z-ai/glm-5.3-flash 1.04858M context $0.15/M input $0.5/M output

Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

x-ai/grok-4.6 500K context $2/M input $6/M output

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

qwen/qwen3.6-plus 1M context $0.325/M input $1.95/M output

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

google/gemini-2.5-pro-preview-05-06 1.04858M context $1.25/M input $10/M output

Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...

anthropic/claude-sonnet-4 200K context $3/M input $15/M output

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly...

openai/gpt-5.2:batch 400K context $0.875/M input $7/M output