3,223 models

No provider description is available for this model yet.

azure/us/gpt-5.6-sol 1.05M context $5.5/M input $33/M output

No provider description is available for this model yet.

azure/us/gpt-5.6-terra 1.05M context $2.75/M input $16.5/M output

Aion-RP-Llama-3.1-8B ranks the highest in the character evaluation portion of the RPBench-Auto benchmark, a roleplaying-specific variant of Arena-Hard-Auto, where LLMs evaluate each other’s responses. It is a fine-tuned base model...

aion-labs/aion-rp-llama-3.1-8b 32.768K context $0.8/M input $1.6/M output

No provider description is available for this model yet.

ollama/mixtral-8x22b-instruct-v0.1 65.536K context Input not listed Output not listed

No provider description is available for this model yet.

azure/us/gpt-5.6-luna 1.05M context $1.1/M input $6.6/M output

No provider description is available for this model yet.

azure/eu/gpt-5.6 1.05M context $5.5/M input $33/M output

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)

mistralai/codestral-2508:batch 256K context $0.15/M input $0.45/M output

No provider description is available for this model yet.

azure/eu/gpt-5.6-sol 1.05M context $5.5/M input $33/M output

No provider description is available for this model yet.

azure/eu/gpt-5.6-terra 1.05M context $2.75/M input $16.5/M output

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...

qwen/qwen3.6-35b-a3b 262.144K context $0.1/M input $0.9/M output

No provider description is available for this model yet.

azure/gpt-5.5 1.05M context $5/M input $30/M output

No provider description is available for this model yet.

azure/us/gpt-5.5 1.05M context $5.5/M input $33/M output

No provider description is available for this model yet.

azure/eu/gpt-5.5 1.05M context $5.5/M input $33/M output

Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances state-of-the-art reasoning and multimodal performance with 8× lower cost...

mistralai/mistral-medium-3 131.072K context $0.4/M input $2/M output

No provider description is available for this model yet.

openai/chatgpt-4o-latest 128K context $5/M input $15/M output

No provider description is available for this model yet.

azure/gpt-5.5-2026-04-23 1.05M context $5/M input $30/M output

No provider description is available for this model yet.

cerebras/llama-3.3-70b 128K context $0.85/M input $1.2/M output

No provider description is available for this model yet.

azure/us/gpt-5.5-2026-04-23 1.05M context $5.5/M input $33/M output

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...

tencent/hy3:free 262.144K context Free input Free output

Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast execution and high efficiency at scale. It uses a “fast...

inclusionai/ling-2.6-1t 262.144K context $0.075/M input $0.625/M output

No provider description is available for this model yet.

openrouter/openai/gpt-4-turbo 128K context $10/M input $30/M output

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

openai/gpt-5.6-luna:batch 1.05M context $0.1/M input $0.6/M output

Coder‑Large is a 32 B‑parameter offspring of Qwen 2.5‑Instruct that has been further trained on permissively‑licensed GitHub, CodeSearchNet and synthetic bug‑fix corpora. It supports a 32k context window, enabling multi‑file...

arcee-ai/coder-large 32.768K context $0.5/M input $0.8/M output

Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...

x-ai/grok-build-0.1 256K context $1/M input $2/M output

GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimized for complex tasks that require step-by-step reasoning,...

openai/gpt-5.2-pro:batch 400K context $10.5/M input $84/M output

KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

kwaipilot/kat-coder-air-v2.5:free 256K context Free input Free output

No provider description is available for this model yet.

gmi/minimaxai/minimax-m2.1 196.608K context $0.3/M input $1.2/M output

No provider description is available for this model yet.

gmi/moonshotai/kimi-k2-thinking 262.144K context $0.8/M input $1.2/M output

KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

kwaipilot/kat-coder-pro-v2.5:free 256K context Free input Free output

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...

qwen/qwen3.7-max 1M context $1.475/M input $4.425/M output

No provider description is available for this model yet.

openrouter/openai/o3-pro 200K context $20/M input $80/M output

No provider description is available for this model yet.

gmi/google/gemini-3-flash-preview 1.04858M context $0.5/M input $3/M output

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for...

openai/gpt-5.5-pro:batch 1.05M context $15/M input $90/M output

Fast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

anthropic/claude-opus-4.7-fast 1M context $30/M input $150/M output

No provider description is available for this model yet.

gmi/google/gemini-3-pro-preview 1.04858M context $2/M input $12/M output

No provider description is available for this model yet.

gmi/deepseek-ai/deepseek-v3-0324 163.84K context $0.28/M input $0.88/M output

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...

nousresearch/hermes-3-llama-3.1-405b:free 131.072K context Free input Free output

Uncensored and creative writing model based on Mistral Small 3.2 24B with good recall, prompt adherence, and intelligence.

thedrummer/cydonia-24b-v4.1 131.072K context $0.3/M input $0.5/M output