OpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high. OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining...

openai/o4-mini-high:batch 200K context $0.55/M input $2.2/M output

No provider description is available for this model yet.

deepinfra/bytedance/seed-2.0-pro 256K context $0.5/M input $3/M output

No provider description is available for this model yet.

deepinfra/moonshotai/kimi-k2.7-code 262.144K context $0.68/M input $3.4/M output

No provider description is available for this model yet.

mistral/mistral-vibe-cli-fast 262.144K context $0.15/M input $0.6/M output

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

anthropic/claude-opus-4.6:batch 1M context $2.5/M input $12.5/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.5-397b-a17b 262.144K context $0.45/M input $3/M output

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

anthropic/claude-fable-5.1:batch 1M context $5/M input $25/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.5-122b-a10b 262.144K context $0.29/M input $2.4/M output

Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video. The Instruct model targets general vision-language use (VQA, document parsing, chart/table...

qwen/qwen3-vl-235b-a22b-instruct 131.072K context $0.21/M input $1.9/M output

No provider description is available for this model yet.

deepinfra/moonshotai/kimi-k3 1.04858M context $2.85/M input $14.25/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.6-27b 262.144K context $0.32/M input $3.2/M output

No provider description is available for this model yet.

deepinfra/google/gemma-4-26b-a4b-it 262.144K context $0.07/M input $0.34/M output

No provider description is available for this model yet.

deepinfra/google/gemini-3.1-pro 1M context $2/M input $12/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.5-9b 262.144K context $0.1/M input $0.15/M output

No provider description is available for this model yet.

deepinfra/google/gemma-4-31b-it 262.144K context $0.13/M input $0.38/M output

Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math....

qwen/qwen3-vl-235b-a22b-thinking 131.072K context $0.4/M input $4/M output

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

google/gemini-2.5-flash:batch 1.04858M context $0.15/M input $1.25/M output

No provider description is available for this model yet.

azure_ai/claude-fable-5-1 1M context $10/M input $50/M output

No provider description is available for this model yet.

moonshot/kimi-k3 1.04858M context $3/M input $15/M output

This model always redirects to the latest model in the GPT Terra family.

~openai/gpt-terra-latest 1.05M context $2/M input $12/M output

No provider description is available for this model yet.

qwencloud/kimi-k2.7-code 229.376K context $0.95/M input $4/M output

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

meta/muse-spark-1.3-contributor 1.04858M context $0.1/M input $0.2/M output

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....

openai/o3:batch 200K context $1/M input $4/M output

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

moonshotai/kimi-k3:batch 1.04858M context $3/M input $15/M output

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to think harder and provide consistently...

openai/o3-pro:batch 200K context $10/M input $40/M output

No provider description is available for this model yet.

qwencloud/qwen3-vl-32b-instruct 131.072K context $0.16/M input $0.64/M output

No provider description is available for this model yet.

qwencloud/qwen3-vl-32b-thinking 131.072K context $0.16/M input $2.87/M output

No provider description is available for this model yet.

qwencloud/qwen3-vl-plus 260.096K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen3.5-plus 991.808K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen3.7-plus 991.808K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen3.8-max 991.808K context $2/M input $6/M output

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

qwen/qwen3.8-flash 1M context $0.15/M input $0.47/M output

Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to...

sakana/fugu-ultra-v2 1M context $5/M input $30/M output

No provider description is available for this model yet.

qwen_ai_platform/kimi-k2.7-code 229.376K context $0.95/M input $4/M output

Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information...

x-ai/grok-4.20-multi-agent 2M context $1.25/M input $2.5/M output

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

openai/gpt-6-astra:batch 1.05M context $5/M input $25/M output