No provider description is available for this model yet.
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Olmo 3 32B Think is a large-scale, 32-billion-parameter model purpose-built for deep reasoning, complex logic chains and advanced instruction-following scenarios. Its capacity enables strong performance on demanding evaluation tasks and...
Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling custom models.
Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed for deeper reasoning and analysis. Pricing is based...
No provider description is available for this model yet.
Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction following, multilingual understanding, and...
[GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,...
Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Thinking variant enhances reasoning in STEM, math, and complex tasks. It excels...
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...
Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code> <update>{edit_snippet}</update>...
No provider description is available for this model yet.
Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B and support for reasoning via Chain-of-Thought. It offers competitive benchmark...
Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code>...
Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets. It supports multimodal inputs—including text, visual data, and audio—enabling diverse tasks...
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
Skyfall 36B v2 is an enhanced iteration of Mistral Small 2501, specifically fine-tuned for improved creativity, nuanced writing, role-playing, and coherent storytelling.
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
Mistral Saba is a 24B-parameter language model specifically designed for the Middle East and South Asia, delivering accurate and contextually relevant responses while maintaining efficient performance. Trained on curated regional...
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, with 45.9 billion parameters activated per inference, and can handle a context...
Euryale L3.3 70B is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.2](/models/sao10k/l3-euryale-70b).
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output. Amazon Nova Lite...
Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost. With a context length...
Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This...
Euryale L3.1 70B v2.2 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.1](/models/sao10k/l3-euryale-70b).
Hermes 3 is a generalist language model with many improvements over [Hermes 2](/models/nousresearch/nous-hermes-2-mistral-7b-dpo), including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...
Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases. It has demonstrated strong...
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
No provider description is available for this model yet.
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See the launch announcement and benchmark results [here](https://www.anthropic.com/news/claude-3-haiku) #multimodal
This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement [here](https://mistral.ai/news/mistral-large-2407/)....
Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it...
UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile systems, and games. Built by ByteDance, it builds upon the UI-TARS framework with reinforcement...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...
No provider description is available for this model yet.
This model always redirects to the latest model in the Claude Haiku family.
No provider description is available for this model yet.
This model always redirects to the latest model in the GPT Mini family.
This model always redirects to the latest model in the Gemini Pro family.
No provider description is available for this model yet.
No provider description is available for this model yet.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| deepseek-ai/DeepSeek-V4-Pro-0813deepinfra/deepseek-ai/deepseek-v4-pro-0813 | 1.04858M | $1.3 | $2.6 | — | |||
| AllenAI: Olmo 3 32B Thinkallenai/olmo-3-32b-think | 65.536K | $0.15 | $0.5 | — | |||
| Amazon: Nova Premier 1.0amazon/nova-premier-v1 | 1M | $2.5 | $12.5 | — | |||
| Perplexity: Sonar Pro Searchperplexity/sonar-pro-search | 200K | $3 | $15 | — | |||
| minimax/minimax-m3novita/minimax/minimax-m3 | 1M | $0.3 | $1.2 | — | |||
| Qwen: Qwen3 30B A3B Instruct 2507qwen/qwen3-30b-a3b-instruct-2507 | 128K | $0.048 | $0.193 | — | |||
| OpenAI: GPT-5 Imageopenai/gpt-5-image | 400K | $10 | $10 | — | |||
| Qwen: Qwen3 VL 30B A3B Thinkingqwen/qwen3-vl-30b-a3b-thinking | 131.072K | $0.2 | $2.4 | — | |||
| DeepSeek: DeepSeek V3.1deepseek/deepseek-chat-v3.1 | 163.84K | $0.25 | $0.95 | — | |||
| OpenAI: gpt-oss-20b (free)openai/gpt-oss-20b:free | 131.072K | Free | Free | — | |||
| Z.ai: GLM 4.5z-ai/glm-4.5 | 131.072K | $0.6 | $2.2 | — | |||
| Morph: Morph V3 Fastmorph/morph-v3-fast | 81.92K | $0.8 | $1.2 | — | |||
| ps/glm-4.5-airpinstripes/ps/glm-4.5-air | 128K | $0.125 | $0.45 | — | |||
| Tencent: Hunyuan A13B Instructtencent/hunyuan-a13b-instruct | 131.072K | $0.14 | $0.57 | — | |||
| Morph: Morph V3 Largemorph/morph-v3-large | 262.144K | $0.9 | $1.9 | — | |||
| Google: Gemma 3n 4Bgoogle/gemma-3n-e4b-it | 32.768K | $0.06 | $0.12 | — | |||
| Google: Gemini 2.5 Pro Preview 06-05google/gemini-2.5-pro-preview | 1.04858M | $1.25 | $10 | — | |||
| TheDrummer: Skyfall 36B V2thedrummer/skyfall-36b-v2 | 32.768K | $0.55 | $0.8 | — | |||
| Meta: Llama 4 Maverickmeta-llama/llama-4-maverick | 128K | $0.2 | $0.696 | — | |||
| Mistral: Sabamistralai/mistral-saba | 32.768K | $0.2 | $0.6 | — | |||
| OpenAI: o3 Mini Highopenai/o3-mini-high | 200K | $1.1 | $4.4 | — | |||
| MiniMax: MiniMax-01minimax/minimax-01 | 1.00019M | $0.2 | $1.1 | — | |||
| Sao10K: Llama 3.3 Euryale 70Bsao10k/l3.3-euryale-70b | 131.072K | $0.65 | $0.75 | — | |||
| Meta: Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | 131.072K | $0.1 | $0.32 | — | |||
| Amazon: Nova Lite 1.0amazon/nova-lite-v1 | 300K | $0.06 | $0.24 | — | |||
| Amazon: Nova Micro 1.0amazon/nova-micro-v1 | 128K | $0.035 | $0.14 | — | |||
| Qwen: Qwen3.5 Plus 2026-04-20qwen/qwen3.5-plus-20260420 | 1M | $0.3 | $1.8 | — | |||
| Sao10K: Llama 3.1 Euryale 70B v2.2sao10k/l3.1-euryale-70b | 131.072K | $0.85 | $0.85 | — | |||
| Nous: Hermes 3 70B Instructnousresearch/hermes-3-llama-3.1-70b | 131.072K | $0.7 | $0.7 | — | |||
| Nous: Hermes 3 405B Instructnousresearch/hermes-3-llama-3.1-405b | 131.072K | $1 | $1 | — | |||
| Meta: Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instruct | 131.072K | $0.4 | $0.4 | — | |||
| Meta: Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | 131.072K | $0.05 | $0.08 | — | |||
| google.gemma-4-31bbedrock_mantle/google.gemma-4-31b | 256K | $0.14 | $0.4 | — | |||
| OpenAI: GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18 | 128K | $0.15 | $0.6 | — | |||
| Anthropic: Claude 3 Haikuanthropic/claude-3-haiku | 200K | $0.25 | $1.25 | — | |||
| Mistral Largemistralai/mistral-large | 128K | $2 | $6 | — | |||
| Qwen: Qwen3 Max Thinkingqwen/qwen3-max-thinking | 262.144K | $0.78 | $3.9 | — | |||
| ByteDance: UI-TARS 7B bytedance/ui-tars-1.5-7b | 128K | $0.1 | $0.2 | — | |||
| databricks-gpt-5-6-terradatabricks/databricks-gpt-5-6-terra | 922K | $2.5 | $15 | — | |||
| databricks-gpt-5-6-soldatabricks/databricks-gpt-5-6-sol | 922K | $4 | $20 | — | |||
| moonshotai/Kimi-K2.6deepinfra/moonshotai/kimi-k2.6 | 262.144K | $0.75 | $3.5 | — | |||
| Meta: Llama Guard 4 12Bmeta-llama/llama-guard-4-12b | 163.84K | $0.18 | $0.18 | — | |||
| NVIDIA: Nemotron 3 Nano Omni (free)nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | 256K | Free | Free | — | |||
| databricks-glm-5-3databricks/databricks-glm-5-3 | 1.04858M | $1.4 | $4.4 | — | |||
| Anthropic: Claude Haiku Latest~anthropic/claude-haiku-latest | 200K | $1 | $5 | — | |||
| databricks-gemini-3-5-flash-litedatabricks/databricks-gemini-3-5-flash-lite | 1.04858M | $0.375 | $3.125 | — | |||
| OpenAI: GPT Mini Latest~openai/gpt-mini-latest | 400K | $0.75 | $4.5 | — | |||
| Google: Gemini Pro Latest~google/gemini-pro-latest | 1.04858M | $2 | $12 | — | |||
| thinkingmachines/Inklingdeepinfra/thinkingmachines/inkling | 524.288K | $0.95 | $4.05 | — | |||
| mindai/macaron-v1-ventinovita/mindai/macaron-v1-venti | 1.04858M | $1.5 | $4.5 | — |