1,564 models

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

thinkingmachines/inkling:free 1.04858M context Free input Free output

No provider description is available for this model yet.

novita/qwen/qwen3.5-397b-a17b 262.144K context $0.6/M input $3.6/M output

No provider description is available for this model yet.

novita/deepseek/deepseek-ocr-2 8.192K context $0.03/M input $0.03/M output

No provider description is available for this model yet.

novita/moonshotai/kimi-k2.5 262.144K context $0.6/M input $3/M output

No provider description is available for this model yet.

novita/qwen/qwen3.6-35b-a3b 262.144K context $0.248/M input $1.485/M output

No provider description is available for this model yet.

wandb/google/gemma-4-31b-it 262.144K context $0.1/M input $0.34/M output

No provider description is available for this model yet.

wandb/minimaxai/minimax-m3 262.144K context $0.23/M input $0.96/M output

No provider description is available for this model yet.

wandb/moonshotai/kimi-k2.7-code 262.144K context $0.71/M input $3.5/M output

No provider description is available for this model yet.

wandb/moonshotai/kimi-k2.6 262.144K context $0.65/M input $3.41/M output

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...

openai/gpt-5.4-mini:batch 400K context $0.375/M input $2.25/M output

No provider description is available for this model yet.

wandb/qwen/qwen3.8-27b 262.144K context $0.4/M input $3/M output

No provider description is available for this model yet.

wandb/qwen/qwen3.6-35b-a3b 262.144K context $0.25/M input $1.25/M output

No provider description is available for this model yet.

wandb/qwen/qwen3.6-27b 262.144K context $0.6/M input $3.6/M output

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K...

openai/gpt-5.4-pro:batch 1.05M context $15/M input $90/M output

No provider description is available for this model yet.

wandb/qwen/qwen3.5-35b-a3b 262.144K context $0.25/M input $1.25/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.8-27b 262.144K context $0.4/M input $3/M output

GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks....

openai/gpt-5-codex:batch 400K context $0.625/M input $5/M output

No provider description is available for this model yet.

deepinfra/moonshotai/kimi-k2.5 262.144K context $0.45/M input $2.25/M output

GPT-5.1 Chat (AKA Instant is the fast, lightweight member of the 5.1 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on...

openai/gpt-5.1-chat 128K context $1.25/M input $10/M output

No provider description is available for this model yet.

deepinfra/xiaomimimo/mimo-v2.5 262.144K context $0.4/M input $2/M output

No provider description is available for this model yet.

databricks/databricks-glm-5-3-flash 1.04858M context Input not listed Output not listed

No provider description is available for this model yet.

gemini/gemma-4-26b-a4b-it 262.144K context Input not listed Output not listed

No provider description is available for this model yet.

gemini/gemma-4-31b-it 262.144K context Input not listed Output not listed

No provider description is available for this model yet.

moonshot/kimi-k2.7-code 262.144K context $0.95/M input $4/M output

No provider description is available for this model yet.

zai/glm-5.3-flash 1.04858M context $0.15/M input $0.5/M output

No provider description is available for this model yet.

gemini/gemini-omni-1.1-flash 131.072K context $1.5/M input $9/M output

No provider description is available for this model yet.

xai/grok-4.20 1M context $1.25/M input $2.5/M output

No provider description is available for this model yet.

xai/grok-4.20-reasoning 1M context $1.25/M input $2.5/M output

No provider description is available for this model yet.

xai/grok-4.20-reasoning-latest 1M context $1.25/M input $2.5/M output

No provider description is available for this model yet.

xai/grok-4.20-non-reasoning 1M context $1.25/M input $2.5/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.5-27b 262.144K context $0.26/M input $2.6/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.6-35b-a3b 262.144K context $0.1/M input $0.95/M output

No provider description is available for this model yet.

groq/qwen/qwen3.8-27b 131.042K context $0.8/M input $4/M output

No provider description is available for this model yet.

mistral/mistral-medium-3.5 262.144K context $1.5/M input $7.5/M output

No provider description is available for this model yet.

mistral/mistral-vibe-cli-latest 262.144K context $1.5/M input $7.5/M output

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

x-ai/grok-4.3 1M context $1.25/M input $2.5/M output

No provider description is available for this model yet.

deepinfra/thinkingmachines/inkling 524.288K context $0.95/M input $4.05/M output

No provider description is available for this model yet.

deepinfra/moonshotai/kimi-k2.6 262.144K context $0.75/M input $3.5/M output