No provider description is available for this model yet.
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
No provider description is available for this model yet.
No provider description is available for this model yet.
GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...
Transform your natural language requests into structured OpenRouter API request objects. Describe what you want to accomplish with AI models, and Body Builder will construct the appropriate API calls. Example:...
The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.
Olmo 3 32B Think is a large-scale, 32-billion-parameter model purpose-built for deep reasoning, complex logic chains and advanced instruction-following scenarios. Its capacity enables strong performance on demanding evaluation tasks and...
Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling custom models.
Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed for deeper reasoning and analysis. Pricing is based...
Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...
Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction following, multilingual understanding, and...
[GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,...
Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Thinking variant enhances reasoning in STEM, math, and complex tasks. It excels...
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...
Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code> <update>{edit_snippet}</update>...
Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in collaboration with Venice.ai. This model is designed as an “uncensored” instruct-tuned LLM, preserving...
Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B and support for reasoning via Chain-of-Thought. It offers competitive benchmark...
Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code>...
Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets. It supports multimodal inputs—including text, visual data, and audio—enabling diverse tasks...
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
Skyfall 36B v2 is an enhanced iteration of Mistral Small 2501, specifically fine-tuned for improved creativity, nuanced writing, role-playing, and coherent storytelling.
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
Mistral Saba is a 24B-parameter language model specifically designed for the Middle East and South Asia, delivering accurate and contextually relevant responses while maintaining efficient performance. Trained on curated regional...
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...
MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, with 45.9 billion parameters activated per inference, and can handle a context...
[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion...
Euryale L3.3 70B is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.2](/models/sao10k/l3-euryale-70b).
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output. Amazon Nova Lite...
Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost. With a context length...
Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This...
Euryale L3.1 70B v2.2 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.1](/models/sao10k/l3-euryale-70b).
Hermes 3 is a generalist language model with many improvements over [Hermes 2](/models/nousresearch/nous-hermes-2-mistral-7b-dpo), including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...
Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...
Lunaris 8B is a versatile generalist and roleplaying model based on Llama 3. It's a strategic merge of multiple models, designed to balance creativity with improved logic and general knowledge....
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases. It has demonstrated strong...
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese,...
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
Gemma 2 27B by Google is an open model built from the same research and technology used to create the [Gemini models](/models?q=gemini). Gemma models are well-suited for a variety of...
Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See the launch announcement and benchmark results [here](https://www.anthropic.com/news/claude-3-haiku) #multimodal
This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement [here](https://mistral.ai/news/mistral-large-2407/)....
Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it...
UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile systems, and games. Built by ByteDance, it builds upon the UI-TARS framework with reinforcement...
Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...
This model always redirects to the latest model in the Claude Haiku family.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| anthropic.claude-v1bedrock/anthropic.claude-v1 | 100K | $8 | $24 | — | |||
| anthropic.claude-v2:1bedrock/anthropic.claude-v2:1 | 100K | $8 | $24 | — | |||
| HuggingFaceH4/zephyr-7b-betaanyscale/huggingfaceh4/zephyr-7b-beta | 16.384K | $0.15 | $0.15 | — | |||
| Z.ai: GLM 4.6Vz-ai/glm-4.6v | 131.072K | $0.3 | $0.9 | — | |||
| Body Builder (beta)openrouter/bodybuilder | 128K | — | — | — | |||
| Mistral: Ministral 3 3B 2512mistralai/ministral-3b-2512 | 131.072K | $0.1 | $0.1 | — | |||
| AllenAI: Olmo 3 32B Thinkallenai/olmo-3-32b-think | 65.536K | $0.15 | $0.5 | — | |||
| Amazon: Nova Premier 1.0amazon/nova-premier-v1 | 1M | $2.5 | $12.5 | — | |||
| Perplexity: Sonar Pro Searchperplexity/sonar-pro-search | 200K | $3 | $15 | — | |||
| Qwen: Qwen3 VL 30B A3B Instructqwen/qwen3-vl-30b-a3b-instruct | 262.144K | $0.15 | $0.6 | — | |||
| Qwen: Qwen3 30B A3B Instruct 2507qwen/qwen3-30b-a3b-instruct-2507 | 128K | $0.048 | $0.193 | — | |||
| OpenAI: GPT-5 Imageopenai/gpt-5-image | 400K | $10 | $10 | — | |||
| Qwen: Qwen3 VL 30B A3B Thinkingqwen/qwen3-vl-30b-a3b-thinking | 131.072K | $0.2 | $2.4 | — | |||
| DeepSeek: DeepSeek V3.1deepseek/deepseek-chat-v3.1 | 163.84K | $0.25 | $0.95 | — | |||
| OpenAI: gpt-oss-20b (free)openai/gpt-oss-20b:free | 131.072K | Free | Free | — | |||
| Z.ai: GLM 4.5z-ai/glm-4.5 | 131.072K | $0.6 | $2.2 | — | |||
| Morph: Morph V3 Fastmorph/morph-v3-fast | 81.92K | $0.8 | $1.2 | — | |||
| Venice: Uncensoredcognitivecomputations/dolphin-mistral-24b-venice-edition | 128K | $0.2 | $0.9 | — | |||
| Tencent: Hunyuan A13B Instructtencent/hunyuan-a13b-instruct | 131.072K | $0.14 | $0.57 | — | |||
| Morph: Morph V3 Largemorph/morph-v3-large | 262.144K | $0.9 | $1.9 | — | |||
| Google: Gemma 3n 4Bgoogle/gemma-3n-e4b-it | 32.768K | $0.06 | $0.12 | — | |||
| Google: Gemini 2.5 Pro Preview 06-05google/gemini-2.5-pro-preview | 1.04858M | $1.25 | $10 | — | |||
| TheDrummer: Skyfall 36B V2thedrummer/skyfall-36b-v2 | 32.768K | $0.55 | $0.8 | — | |||
| Meta: Llama 4 Maverickmeta-llama/llama-4-maverick | 128K | $0.188 | $0.652 | — | |||
| Mistral: Sabamistralai/mistral-saba | 32.768K | $0.2 | $0.6 | — | |||
| OpenAI: o3 Mini Highopenai/o3-mini-high | 200K | $1.1 | $4.4 | — | |||
| DeepSeek: R1 Distill Llama 70Bdeepseek/deepseek-r1-distill-llama-70b | 8.192K | $0.8 | $0.8 | — | |||
| MiniMax: MiniMax-01minimax/minimax-01 | 1.00019M | $0.2 | $1.1 | — | |||
| Microsoft: Phi 4microsoft/phi-4 | 16.384K | $0.07 | $0.14 | — | |||
| Sao10K: Llama 3.3 Euryale 70Bsao10k/l3.3-euryale-70b | 131.072K | $0.65 | $0.75 | — | |||
| Meta: Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | 131.072K | $0.1 | $0.32 | — | |||
| Amazon: Nova Lite 1.0amazon/nova-lite-v1 | 300K | $0.06 | $0.24 | — | |||
| Amazon: Nova Micro 1.0amazon/nova-micro-v1 | 128K | $0.035 | $0.14 | — | |||
| Qwen: Qwen3.5 Plus 2026-04-20qwen/qwen3.5-plus-20260420 | 1M | $0.3 | $1.8 | — | |||
| Sao10K: Llama 3.1 Euryale 70B v2.2sao10k/l3.1-euryale-70b | 131.072K | $0.85 | $0.85 | — | |||
| Nous: Hermes 3 70B Instructnousresearch/hermes-3-llama-3.1-70b | 131.072K | $0.7 | $0.7 | — | |||
| Nous: Hermes 3 405B Instructnousresearch/hermes-3-llama-3.1-405b | 131.072K | $1 | $1 | — | |||
| Sao10K: Llama 3 8B Lunarissao10k/l3-lunaris-8b | 8.192K | $0.04 | $0.05 | — | |||
| Meta: Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instruct | 131.072K | $0.4 | $0.4 | — | |||
| Meta: Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | 131.072K | $0.05 | $0.08 | — | |||
| Mistral: Mistral Nemomistralai/mistral-nemo | 131.072K | $0.019 | $0.03 | — | |||
| OpenAI: GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18 | 128K | $0.15 | $0.6 | — | |||
| Google: Gemma 2 27Bgoogle/gemma-2-27b-it | 8.192K | $0.65 | $0.65 | — | |||
| Anthropic: Claude 3 Haikuanthropic/claude-3-haiku | 200K | $0.25 | $1.25 | — | |||
| Mistral Largemistralai/mistral-large | 128K | $2 | $6 | — | |||
| Qwen: Qwen3 Max Thinkingqwen/qwen3-max-thinking | 262.144K | $0.78 | $3.9 | — | |||
| ByteDance: UI-TARS 7B bytedance/ui-tars-1.5-7b | 128K | $0.1 | $0.2 | — | |||
| Meta: Llama Guard 4 12Bmeta-llama/llama-guard-4-12b | 163.84K | $0.18 | $0.18 | — | |||
| NVIDIA: Nemotron 3 Nano Omni (free)nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | 256K | Free | Free | — | |||
| Anthropic: Claude Haiku Latest~anthropic/claude-haiku-latest | 200K | $1 | $5 | — |