Euryale L3.1 70B v2.2 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.1](/models/sao10k/l3-euryale-70b).
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Hermes 3 is a generalist language model with many improvements over [Hermes 2](/models/nousresearch/nous-hermes-2-mistral-7b-dpo), including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...
Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases. It has demonstrated strong...
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese,...
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See the launch announcement and benchmark results [here](https://www.anthropic.com/news/claude-3-haiku) #multimodal
This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement [here](https://mistral.ai/news/mistral-large-2407/)....
Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it...
UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile systems, and games. Built by ByteDance, it builds upon the UI-TARS framework with reinforcement...
Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...
This model always redirects to the latest model in the Claude Haiku family.
This model always redirects to the latest model in the GPT Mini family.
This model always redirects to the latest model in the Gemini Pro family.
This model always redirects to the latest model in the Kimi family.
This model always redirects to the latest model in the Gemini Flash family.
This model always redirects to the latest model in the Claude Sonnet family.
Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...
This model always redirects to the latest model in the OpenAI GPT family.
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...
Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts text and image input and produces...
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...
GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rolls out new Instant model updates...
Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and operational efficiency. It is optimized for coding agents, tool...
Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal capabilities. It provides state-of-the-art performance in text-based reasoning and...
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows...
GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text...
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the...
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...
MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...
Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...
Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total parameters and only 3B activated per...
The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenRouter. The router smartly filters for models that...
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...
Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...
Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing in autonomous programming via tool calling...
The preview GPT-4 model with improved instruction following, JSON mode, reproducible outputs, parallel function calling, and more. Training data: up to Dec 2023. **Note:** heavily rate limited by OpenAI while...
NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a hybrid Transformer-Mamba architecture, combining transformer-level accuracy with Mamba’s...
Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models. These models are the latest in a series of models released by IBM. They are fine-tuned for long...
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model, designed for advanced visual and textual reasoning across complex scenes, documents, and temporal sequences. It integrates enhanced multimodal alignment and...
Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual...
Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the...
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...
May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Sao10K: Llama 3.1 Euryale 70B v2.2sao10k/l3.1-euryale-70b | 131.072K | $0.85 | $0.85 | — | |||
| Nous: Hermes 3 70B Instructnousresearch/hermes-3-llama-3.1-70b | 131.072K | $0.7 | $0.7 | — | |||
| Nous: Hermes 3 405B Instructnousresearch/hermes-3-llama-3.1-405b | 131.072K | $1 | $1 | — | |||
| Meta: Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instruct | 131.072K | $0.4 | $0.4 | — | |||
| Meta: Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | 131.072K | $0.05 | $0.08 | — | |||
| Mistral: Mistral Nemomistralai/mistral-nemo | 131.072K | $0.019 | $0.03 | — | |||
| OpenAI: GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18 | 128K | $0.15 | $0.6 | — | |||
| Anthropic: Claude 3 Haikuanthropic/claude-3-haiku | 200K | $0.25 | $1.25 | — | |||
| Mistral Largemistralai/mistral-large | 128K | $2 | $6 | — | |||
| Qwen: Qwen3 Max Thinkingqwen/qwen3-max-thinking | 262.144K | $0.78 | $3.9 | — | |||
| ByteDance: UI-TARS 7B bytedance/ui-tars-1.5-7b | 128K | $0.1 | $0.2 | — | |||
| Meta: Llama Guard 4 12Bmeta-llama/llama-guard-4-12b | 163.84K | $0.18 | $0.18 | — | |||
| NVIDIA: Nemotron 3 Nano Omni (free)nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | 256K | Free | Free | — | |||
| Anthropic: Claude Haiku Latest~anthropic/claude-haiku-latest | 200K | $1 | $5 | — | |||
| OpenAI: GPT Mini Latest~openai/gpt-mini-latest | 400K | $0.75 | $4.5 | — | |||
| Google: Gemini Pro Latest~google/gemini-pro-latest | 1.04858M | $2 | $12 | — | |||
| MoonshotAI: Kimi Latest~moonshotai/kimi-latest | 1.04858M | $2.025 | $11.34 | — | |||
| Google: Gemini Flash Latest~google/gemini-flash-latest | 1.04858M | $0.75 | $3.75 | — | |||
| Anthropic: Claude Sonnet Latest~anthropic/claude-sonnet-latest | 1M | $2 | $10 | — | |||
| Poolside: Laguna XS 2.1 (free)poolside/laguna-xs-2.1:free | 262.144K | Free | Free | — | |||
| OpenAI GPT Latest~openai/gpt-latest | 1.05M | $2 | $10 | — | |||
| NVIDIA: Nemotron 3.5 Content Safety (free)nvidia/nemotron-3.5-content-safety:free | 128K | Free | Free | — | |||
| Nex AGI: Nex-N2-Pronex-agi/nex-n2-pro | 262.144K | $0.25 | $1 | — | |||
| Mistral: Mistral Medium 3.5mistralai/mistral-medium-3-5 | 262.144K | $1.5 | $7.5 | — | |||
| OpenAI: GPT Chat Latestopenai/gpt-chat-latest | 400K | $5 | $30 | — | |||
| inclusionAI: Ring-2.6-1Tinclusionai/ring-2.6-1t | 262.144K | $0.075 | $0.625 | — | |||
| Mistral: Mistral Small 3.1 24Bmistralai/mistral-small-3.1-24b-instruct | 128K | $0.351 | $0.555 | — | |||
| Google: Gemma 4 26B A4B (free)google/gemma-4-26b-a4b-it:free | 262.144K | Free | Free | — | |||
| Google: Gemma 4 31B (free)google/gemma-4-31b-it:free | 262.144K | Free | Free | — | |||
| Z.ai: GLM 5 Turboz-ai/glm-5-turbo | 202.752K | $1.2 | $4 | — | |||
| OpenAI: GPT-5 Image Miniopenai/gpt-5-image-mini | 400K | $2.5 | $2 | — | |||
| Qwen: Qwen3.5-Flashqwen/qwen3.5-flash-02-23 | 1M | $0.065 | $0.26 | — | |||
| Qwen: Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | 262.144K | $0.55 | $3.5 | — | |||
| MiniMax: MiniMax M2.5minimax/minimax-m2.5 | 200K | $0.27 | $1.08 | — | |||
| Anthropic: Claude Opus 4.6anthropic/claude-opus-4.6 | 1M | $5 | $25 | — | |||
| Qwen: Qwen3 Coder Nextqwen/qwen3-coder-next | 262.144K | $0.12 | $0.8 | — | |||
| Free Models Routeropenrouter/free | 200K | Free | Free | — | |||
| NVIDIA: Nemotron 3 Nano 30B A3B (free)nvidia/nemotron-3-nano-30b-a3b:free | 256K | Free | Free | — | |||
| Mistral: Ministral 3 14B 2512mistralai/ministral-14b-2512 | 262.144K | $0.2 | $0.2 | — | |||
| Anthropic: Claude Opus 4.5anthropic/claude-opus-4.5 | 200K | $5 | $25 | — | |||
| Qwen: Qwen3 Coder Flashqwen/qwen3-coder-flash | 1M | $0.195 | $0.975 | — | |||
| OpenAI: GPT-4 Turbo Previewopenai/gpt-4-turbo-preview | 128K | $10 | $30 | — | |||
| NVIDIA: Nemotron Nano 12B 2 VL (free)nvidia/nemotron-nano-12b-v2-vl:free | 128K | Free | Free | — | |||
| IBM: Granite 4.0 Microibm-granite/granite-4.0-h-micro | 131K | $0.017 | $0.112 | — | |||
| Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5 | 200K | $1 | $5 | — | |||
| Qwen: Qwen3 VL 8B Thinkingqwen/qwen3-vl-8b-thinking | 131.072K | $0.18 | $2.1 | — | |||
| Qwen: Qwen3 Next 80B A3B Instructqwen/qwen3-next-80b-a3b-instruct | 262.144K | $0.09 | $1.1 | — | |||
| Qwen: Qwen3 Coder 30B A3B Instructqwen/qwen3-coder-30b-a3b-instruct | 262.144K | $0.07 | $0.28 | — | |||
| Qwen: Qwen3 235B A22B Instruct 2507qwen/qwen3-235b-a22b-2507 | 262.144K | $0.087 | $0.35 | — | |||
| DeepSeek: R1 0528deepseek/deepseek-r1-0528 | 163.84K | $0.5 | $2.15 | — |