4,182 models

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...

tencent/hy3:free 262.144K context Free input Free output
Open weights

No provider description is available for this model yet.

meta-llama/Meta-Llama-3-70B-Instruct Not documented context Input not listed Output not listed

No provider description is available for this model yet.

xai/grok-4.20-multi-agent-0309 1M context $1.25/M input $2.5/M output

Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it...

qwen/qwen3-max-thinking 262.144K context $0.78/M input $3.9/M output

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...

qwen/qwen3.5-35b-a3b 256K context $0.312/M input $1.25/M output

No provider description is available for this model yet.

mistral/labs-leanstral-1-5 262.144K context Input not listed Output not listed

No provider description is available for this model yet.

meta-llama/Llama-Guard-4-12B Not documented context Input not listed Output not listed

KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

kwaipilot/kat-coder-air-v2.5 256K context $0.15/M input $0.6/M output

GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

openai/gpt-5.6-luna-pro 1.05M context $0.2/M input $1.2/M output

No provider description is available for this model yet.

meta-llama/Llama-4-Scout-17B-16E Not documented context Input not listed Output not listed

This model always redirects to the latest model in the DeepSeek V4 Flash family.

~deepseek/deepseek-v4-flash-latest 1.04858M context $0.04/M input $0.08/M output
Open weights

No provider description is available for this model yet.

meta-llama/Llama-3.3-70B-Instruct Not documented context Input not listed Output not listed

No provider description is available for this model yet.

github_copilot/gpt-4o 64K context Input not listed Output not listed

No provider description is available for this model yet.

bedrock/us-east-2/deepseek.v3.2 163.84K context $0.62/M input $1.85/M output

No provider description is available for this model yet.

github_copilot/gpt-4o-2024-08-06 64K context Input not listed Output not listed
Open weights

No provider description is available for this model yet.

meta-llama/Llama-3.1-70B-Instruct Not documented context Input not listed Output not listed
Open weights

No provider description is available for this model yet.

microsoft/Fara1.5-27B Not documented context Input not listed Output not listed

This model always redirects to the latest model in the OpenAI GPT Astra family.

~openai/gpt-astra-latest 1.05M context $10/M input $50/M output

This model always redirects to the latest model in the OpenAI GPT Sol family.

~openai/gpt-sol-latest 1.05M context $2/M input $10/M output

The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason...

openai/o1:batch 200K context $7.5/M input $30/M output

No provider description is available for this model yet.

snowflake/llama3.3-70b 128K context $0.72/M input $0.72/M output

This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement [here](https://mistral.ai/news/mistral-large-2407/)....

mistralai/mistral-large 128K context $2/M input $6/M output

No provider description is available for this model yet.

snowflake/mistral-large 32K context Input not listed Output not listed

Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather...

inference-net/schematron-v2-turbo 128K context $0.03/M input $0.15/M output