No provider description is available for this model yet.

deepinfra/zai-org/glm-5.1 202.752K context $1.05/M input $3.5/M output

No provider description is available for this model yet.

wandb/moonshotai/kimi-k2.7-code 262.144K context $0.71/M input $3.5/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.5-122b-a10b 262.144K context $0.29/M input $2.4/M output

No provider description is available for this model yet.

wandb/minimaxai/minimax-m3 262.144K context $0.23/M input $0.96/M output

No provider description is available for this model yet.

novita/google/gemma-4-26b-a4b-it 262.144K context $0.13/M input $0.4/M output

No provider description is available for this model yet.

novita/qwen/qwen3.8-max 1M context $2/M input $6/M output

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

google/gemini-3.1-pro-preview:batch 1.04858M context $1/M input $6/M output

No provider description is available for this model yet.

zai/glm-5.3 1M context $1.4/M input $4.4/M output

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

google/gemini-3.5-flash:batch 1.04858M context $0.75/M input $4.5/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.8-max 256K context $1.65/M input $4.951/M output

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.

openai/gpt-4-turbo:batch 128K context $5/M input $15/M output

No provider description is available for this model yet.

deepinfra/deepseek-ai/deepseek-v3.2 163.84K context $0.26/M input $0.38/M output

No provider description is available for this model yet.

deepinfra/google/gemma-4-e4b-it 131.072K context $0.02/M input $0.1/M output

No provider description is available for this model yet.

wandb/ibm-granite/granite-4.1-8b 131.072K context $0.05/M input $0.1/M output

No provider description is available for this model yet.

mistral/labs-leanstral-1-5-1 262.144K context Input not listed Output not listed

No provider description is available for this model yet.

wandb/google/gemma-4-31b-it 262.144K context $0.1/M input $0.34/M output

Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It...

qwen/qwen3-max 262.144K context $0.78/M input $3.9/M output

No provider description is available for this model yet.

mistral/mistral-code-fim-latest 128K context $0.3/M input $0.9/M output

No provider description is available for this model yet.

wandb/deepseek-ai/deepseek-v4-pro 1.04858M context $1.15/M input $2.55/M output

No provider description is available for this model yet.

mistral/mistral-code-latest 128K context $0.3/M input $0.9/M output

No provider description is available for this model yet.

novita/zai-org/glm-5v-turbo 204.8K context $1.2/M input $4/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.5-397b-a17b 262.144K context $0.45/M input $3/M output

No provider description is available for this model yet.

mistral/mistral-vibe-cli-fast 262.144K context $0.15/M input $0.6/M output

No provider description is available for this model yet.

wandb/deepseek-ai/deepseek-v4-flash 1.04858M context $0.14/M input $0.28/M output

No provider description is available for this model yet.

deepinfra/moonshotai/kimi-k2.7-code 262.144K context $0.68/M input $3.4/M output

Relace Apply 3 is a specialized code-patching LLM that merges AI-suggested edits straight into your source files. It can apply updates from GPT-4o, Claude, and others into your files at...

relace/relace-apply-3 256K context $0.85/M input $1.25/M output

No provider description is available for this model yet.

novita/deepseek/deepseek-v4-pro 1.04858M context $1.6/M input $3.2/M output

No provider description is available for this model yet.

novita/moonshotai/kimi-k2.7-code 262.144K context $0.95/M input $4/M output

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

mistralai/mistral-large-2512:batch 262.144K context $0.25/M input $0.75/M output

No provider description is available for this model yet.

deepinfra/zai-org/glm-5 202.752K context $0.6/M input $2.08/M output

No provider description is available for this model yet.

novita/thudm/glm-4-32b-0414 32K context $0.55/M input $1.66/M output

Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances state-of-the-art reasoning and multimodal performance with 8× lower cost...

mistralai/mistral-medium-3 131.072K context $0.4/M input $2/M output

No provider description is available for this model yet.

deepinfra/bytedance/seed-2.0-pro 256K context $0.5/M input $3/M output

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

minimax/minimax-m3:free 1.04858M context Free input Free output

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

anthropic/claude-opus-4.8:batch 1M context $2.5/M input $12.5/M output