No provider description is available for this model yet.

deepinfra/qwen/qwen3.6-27b 262.144K context $0.32/M input $3.2/M output

No provider description is available for this model yet.

deepinfra/google/gemma-4-26b-a4b-it 262.144K context $0.07/M input $0.34/M output

No provider description is available for this model yet.

deepinfra/google/gemini-3.1-pro 1M context $2/M input $12/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.5-9b 262.144K context $0.1/M input $0.15/M output

No provider description is available for this model yet.

deepinfra/google/gemma-4-31b-it 262.144K context $0.13/M input $0.38/M output

Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math....

qwen/qwen3-vl-235b-a22b-thinking 131.072K context $0.4/M input $4/M output

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...

google/gemini-2.5-flash-lite:batch 1.04858M context $0.05/M input $0.2/M output

Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...

anthropic/claude-sonnet-4 200K context $3/M input $15/M output

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

google/gemini-2.5-flash:batch 1.04858M context $0.15/M input $1.25/M output

No provider description is available for this model yet.

azure_ai/claude-fable-5-1 1M context $10/M input $50/M output

Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and improved function calling. Compared to the 3.1 release, version 3.2 significantly improves accuracy on...

mistralai/mistral-small-3.2-24b-instruct 128K context $0.075/M input $0.2/M output

No provider description is available for this model yet.

moonshot/kimi-k3 1.04858M context $3/M input $15/M output

This model always redirects to the latest model in the GPT Terra family.

~openai/gpt-terra-latest 1.05M context $2/M input $12/M output

No provider description is available for this model yet.

qwencloud/kimi-k2.7-code 229.376K context $0.95/M input $4/M output

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

meta/muse-spark-1.3-contributor 1.04858M context $0.1/M input $0.2/M output

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

moonshotai/kimi-k3:batch 1.04858M context $3/M input $15/M output

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to think harder and provide consistently...

openai/o3-pro:batch 200K context $10/M input $40/M output

No provider description is available for this model yet.

qwencloud/qwen3-vl-32b-instruct 131.072K context $0.16/M input $0.64/M output

No provider description is available for this model yet.

qwencloud/qwen3-vl-32b-thinking 131.072K context $0.16/M input $2.87/M output

No provider description is available for this model yet.

qwencloud/qwen3-vl-plus 260.096K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen3.5-plus 991.808K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen3.7-plus 991.808K context Input not listed Output not listed

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

qwen/qwen3.8-flash 1M context $0.15/M input $0.47/M output

No provider description is available for this model yet.

qwencloud/qwen3.8-max 991.808K context $2/M input $6/M output

Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to...

sakana/fugu-ultra-v2 1M context $5/M input $30/M output

No provider description is available for this model yet.

qwen_ai_platform/kimi-k2.7-code 229.376K context $0.95/M input $4/M output

Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information...

x-ai/grok-4.20-multi-agent 2M context $1.25/M input $2.5/M output

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

minimax/minimax-m3:batch 524.288K context $0.3/M input $1.2/M output

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

openai/gpt-6-astra:batch 1.05M context $5/M input $25/M output

GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and...

openai/gpt-5-pro:batch 400K context $7.5/M input $60/M output

No provider description is available for this model yet.

qwen_ai_platform/qwen3-vl-plus 260.096K context Input not listed Output not listed

No provider description is available for this model yet.

qwen_ai_platform/qwen3.5-plus 991.808K context Input not listed Output not listed

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon...

meta/muse-glimmer-30b:batch 131.072K context $0.175/M input $0.75/M output

No provider description is available for this model yet.

qwen_ai_platform/qwen3.7-plus 991.808K context Input not listed Output not listed

No provider description is available for this model yet.

qwen_ai_platform/qwen3.8-max 991.808K context $2/M input $6/M output

No provider description is available for this model yet.

friendliai/zai-org/glm-5.3-flash 1.04858M context $0.15/M input $0.5/M output