Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction following, multilingual understanding, and...

qwen/qwen3-30b-a3b-instruct-2507 128K context $0.048/M input $0.193/M output

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...

deepseek/deepseek-chat-v3.1 163.84K context $0.25/M input $0.95/M output

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...

openai/gpt-oss-20b:free 131.072K context Free input Free output

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...

z-ai/glm-4.5 131.072K context $0.6/M input $2.2/M output

Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in collaboration with Venice.ai. This model is designed as an “uncensored” instruct-tuned LLM, preserving...

cognitivecomputations/dolphin-mistral-24b-venice-edition 128K context $0.2/M input $0.9/M output

Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B and support for reasoning via Chain-of-Thought. It offers competitive benchmark...

tencent/hunyuan-a13b-instruct 131.072K context $0.14/M input $0.57/M output

Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code>...

morph/morph-v3-large 262.144K context $0.9/M input $1.9/M output

Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code> <update>{edit_snippet}</update>...

morph/morph-v3-fast 81.92K context $0.8/M input $1.2/M output

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

google/gemini-2.5-pro-preview 1.04858M context $1.25/M input $10/M output

Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets. It supports multimodal inputs—including text, visual data, and audio—enabling diverse tasks...

google/gemma-3n-e4b-it 32.768K context $0.06/M input $0.12/M output

Skyfall 36B v2 is an enhanced iteration of Mistral Small 2501, specifically fine-tuned for improved creativity, nuanced writing, role-playing, and coherent storytelling.

thedrummer/skyfall-36b-v2 32.768K context $0.55/M input $0.8/M output

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...

meta-llama/llama-4-maverick 128K context $0.188/M input $0.652/M output

Mistral Saba is a 24B-parameter language model specifically designed for the Middle East and South Asia, delivering accurate and contextually relevant responses while maintaining efficient performance. Trained on curated regional...

mistralai/mistral-saba 32.768K context $0.2/M input $0.6/M output

OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...

openai/o3-mini-high 200K context $1.1/M input $4.4/M output

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...

deepseek/deepseek-r1-distill-llama-70b 8.192K context $0.8/M input $0.8/M output

MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, with 45.9 billion parameters activated per inference, and can handle a context...

minimax/minimax-01 1.00019M context $0.2/M input $1.1/M output

Euryale L3.3 70B is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.2](/models/sao10k/l3-euryale-70b).

sao10k/l3.3-euryale-70b 131.072K context $0.65/M input $0.75/M output

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

meta-llama/llama-3.3-70b-instruct 131.072K context $0.1/M input $0.32/M output

Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output. Amazon Nova Lite...

amazon/nova-lite-v1 300K context $0.06/M input $0.24/M output

Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost. With a context length...

amazon/nova-micro-v1 128K context $0.035/M input $0.14/M output

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This...

qwen/qwen3.5-plus-20260420 1M context $0.3/M input $1.8/M output

Euryale L3.1 70B v2.2 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.1](/models/sao10k/l3-euryale-70b).

sao10k/l3.1-euryale-70b 131.072K context $0.85/M input $0.85/M output

Hermes 3 is a generalist language model with many improvements over [Hermes 2](/models/nousresearch/nous-hermes-2-mistral-7b-dpo), including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...

nousresearch/hermes-3-llama-3.1-70b 131.072K context $0.7/M input $0.7/M output

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...

nousresearch/hermes-3-llama-3.1-405b 131.072K context $1/M input $1/M output

Lunaris 8B is a versatile generalist and roleplaying model based on Llama 3. It's a strategic merge of multiple models, designed to balance creativity with improved logic and general knowledge....

sao10k/l3-lunaris-8b 8.192K context $0.04/M input $0.05/M output

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases. It has demonstrated strong...

meta-llama/llama-3.1-70b-instruct 131.072K context $0.4/M input $0.4/M output

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...

meta-llama/llama-3.1-8b-instruct 131.072K context $0.05/M input $0.08/M output

A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese,...

mistralai/mistral-nemo 131.072K context $0.019/M input $0.03/M output

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...

openai/gpt-4o-mini-2024-07-18 128K context $0.15/M input $0.6/M output

Gemma 2 27B by Google is an open model built from the same research and technology used to create the [Gemini models](/models?q=gemini). Gemma models are well-suited for a variety of...

google/gemma-2-27b-it 8.192K context $0.65/M input $0.65/M output

Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See the launch announcement and benchmark results [here](https://www.anthropic.com/news/claude-3-haiku) #multimodal

anthropic/claude-3-haiku 200K context $0.25/M input $1.25/M output

This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement [here](https://mistral.ai/news/mistral-large-2407/)....

mistralai/mistral-large 128K context $2/M input $6/M output

Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it...

qwen/qwen3-max-thinking 262.144K context $0.78/M input $3.9/M output

UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile systems, and games. Built by ByteDance, it builds upon the UI-TARS framework with reinforcement...

bytedance/ui-tars-1.5-7b 128K context $0.1/M input $0.2/M output

Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...

meta-llama/llama-guard-4-12b 163.84K context $0.18/M input $0.18/M output

NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free 256K context Free input Free output

This model always redirects to the latest model in the GPT Mini family.

~openai/gpt-mini-latest 400K context $0.75/M input $4.5/M output

This model always redirects to the latest model in the Gemini Pro family.

~google/gemini-pro-latest 1.04858M context $2/M input $12/M output

This model always redirects to the latest model in the Kimi family.

~moonshotai/kimi-latest 1.04858M context $1.975/M input $11.06/M output

This model always redirects to the latest model in the Gemini Flash family.

~google/gemini-flash-latest 1.04858M context $0.75/M input $3.75/M output

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...

poolside/laguna-xs-2.1:free 262.144K context Free input Free output

This model always redirects to the latest model in the OpenAI GPT family.

~openai/gpt-latest 1.05M context $2/M input $10/M output

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

nvidia/nemotron-3.5-content-safety:free 128K context Free input Free output

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

mistralai/mistral-medium-3-5 262.144K context $1.5/M input $7.5/M output

Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts text and image input and produces...

nex-agi/nex-n2-pro 262.144K context $0.25/M input $1/M output

GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rolls out new Instant model updates...

openai/gpt-chat-latest 400K context $5/M input $30/M output

Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and operational efficiency. It is optimized for coding agents, tool...

inclusionai/ring-2.6-1t 262.144K context $0.075/M input $0.625/M output

Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal capabilities. It provides state-of-the-art performance in text-based reasoning and...

mistralai/mistral-small-3.1-24b-instruct 128K context $0.351/M input $0.555/M output