3,646 models

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

thinkingmachines/inkling-small:free 1.04858M context Free input Free output

No provider description is available for this model yet.

google/medgemma-4b-it Not documented context Input not listed Output not listed

No provider description is available for this model yet.

google/medgemma-4b-pt Not documented context Input not listed Output not listed

No provider description is available for this model yet.

google/medgemma-27b-text-it Not documented context Input not listed Output not listed

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

anthropic/claude-fable-5.1 1M context $10/M input $50/M output

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

google/gemma-4-31b-it:batch 262.144K context $0.39/M input $0.97/M output

No provider description is available for this model yet.

azure/command-r-plus 128K context $3/M input $15/M output

GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

openai/gpt-6-astra-pro 1.05M context $10/M input $50/M output

No provider description is available for this model yet.

azure_ai/claude-haiku-4-5 200K context $1/M input $5/M output

No provider description is available for this model yet.

azure_ai/claude-opus-4-5 200K context $5/M input $25/M output

No provider description is available for this model yet.

azure_ai/claude-opus-4-6 1M context $5/M input $25/M output

No provider description is available for this model yet.

azure_ai/claude-opus-4-7 1M context $5/M input $25/M output

No provider description is available for this model yet.

azure_ai/claude-fable-5 1M context $10/M input $50/M output

No provider description is available for this model yet.

azure_ai/claude-opus-5 1M context $5/M input $25/M output

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K...

openai/gpt-5.4-pro:batch 1.05M context $15/M input $90/M output

No provider description is available for this model yet.

azure_ai/claude-opus-4-8 1M context $5/M input $25/M output

Inflection 3 Productivity is optimized for following instructions. It is better for tasks requiring JSON output or precise adherence to provided guidelines. It has access to recent news. For emotional...

inflection/inflection-3-productivity 8K context $2.5/M input $10/M output

No provider description is available for this model yet.

azure_ai/claude-opus-4-1 200K context $15/M input $75/M output

Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts. It is suited for Japanese instruction following,...

sakana/namazu 262.144K context $0.95/M input $4/M output

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...

nvidia/nemotron-3-nano-30b-a3b:free 256K context Free input Free output

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

qwen/qwen3.8-2.4t-a95b:batch 1.01M context $2/M input $6/M output

Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed for deeper reasoning and analysis. Pricing is based...

perplexity/sonar-pro-search 200K context $3/M input $15/M output

No provider description is available for this model yet.

azure_ai/claude-sonnet-4-6 1M context $3/M input $15/M output

No provider description is available for this model yet.

azure/computer-use-preview 8.192K context $3/M input $12/M output

No provider description is available for this model yet.

azure/container Not documented context Input not listed Output not listed

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

z-ai/glm-5.2 202.752K context $0.6/M input $2/M output

No provider description is available for this model yet.

azure_ai/gpt-oss-120b 131.072K context $0.15/M input $0.6/M output

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.

openai/gpt-3.5-turbo:batch 16.385K context $0.25/M input $0.75/M output

No provider description is available for this model yet.

azure_ai/gpt-5.5 1.05M context $5/M input $30/M output

No provider description is available for this model yet.

gemini/gemini-3.8-flash 1.04858M context $0.75/M input $3.75/M output