GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See the launch announcement and benchmark results [here](https://www.anthropic.com/news/claude-3-haiku) #multimodal
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
No provider description is available for this model yet.
A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese,...
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
No provider description is available for this model yet.
KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...
No provider description is available for this model yet.
No provider description is available for this model yet.
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...
No provider description is available for this model yet.
No provider description is available for this model yet.
Laguna M.1 is the flagship coding agent model from [Poolside](https://poolside.ai/), optimized for complex software engineering tasks. Designed for agentic coding workflows, it supports tool calling and reasoning, with a 256K...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases. It has demonstrated strong...
No provider description is available for this model yet.
Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
No provider description is available for this model yet.
Hermes 3 is a generalist language model with many improvements over [Hermes 2](/models/nousresearch/nous-hermes-2-mistral-7b-dpo), including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...
Euryale L3.1 70B v2.2 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.1](/models/sao10k/l3-euryale-70b).
No provider description is available for this model yet.
No provider description is available for this model yet.
Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...
Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost. With a context length...
No provider description is available for this model yet.
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...
Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output. Amazon Nova Lite...
No provider description is available for this model yet.
No provider description is available for this model yet.
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| OpenAI: GPT-5 Nano (batch)openai/gpt-5-nano:batch | 400K | $0.025 | $0.2 | — | |||
| Anthropic: Claude 3 Haikuanthropic/claude-3-haiku | 200K | $0.25 | $1.25 | — | |||
| OpenAI: GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18 | 128K | $0.15 | $0.6 | — | |||
| NousResearch/Hermes-4-70Bnebius/nousresearch/hermes-4-70b | 131.072K | $0.13 | $0.4 | — | |||
| Mistral: Mistral Nemomistralai/mistral-nemo | 131.072K | $0.019 | $0.03 | — | |||
| Meta: Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | 131.072K | $0.05 | $0.08 | — | |||
| NousResearch/Hermes-4-405Bnebius/nousresearch/hermes-4-405b | 131.072K | $1 | $3 | — | |||
| Kwaipilot: KAT-Coder-Air V2.5 (free)kwaipilot/kat-coder-air-v2.5:free | 256K | Free | Free | — | |||
| ap-south-1/minimax.minimax-m2.5bedrock/ap-south-1/minimax.minimax-m2.5 | 1M | $0.36 | $1.44 | — | |||
| ap-south-1/moonshotai.kimi-k2-thinkingbedrock/ap-south-1/moonshotai.kimi-k2-thinking | 262.144K | $0.71 | $2.94 | — | |||
| Tencent: Hy3 (free)tencent/hy3:free | 262.144K | Free | Free | — | |||
| qwen3-max-previewqwen_ai_platform/qwen3-max-preview | 258.048K | — | — | — | |||
| ap-south-1/moonshotai.kimi-k2.5bedrock/ap-south-1/moonshotai.kimi-k2.5 | 262.144K | $0.72 | $3.6 | — | |||
| Poolside: Laguna M.1 (free)poolside/laguna-m.1:free | 262.144K | Free | Free | — | |||
| ap-south-1/qwen.qwen3-coder-nextbedrock/ap-south-1/qwen.qwen3-coder-next | 262.144K | $0.6 | $1.44 | — | |||
| ap-southeast-2/minimax.minimax-m2.5bedrock/ap-southeast-2/minimax.minimax-m2.5 | 1M | $0.309 | $1.236 | — | |||
| qwen-flashqwencloud/qwen-flash | 997.952K | — | — | — | |||
| deepseek-ai/DeepSeek-V4-Flashdeepinfra/deepseek-ai/deepseek-v4-flash | 1.04858M | $0.09 | $0.18 | — | |||
| Meta: Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instruct | 131.072K | $0.4 | $0.4 | — | |||
| ap-southeast-3/minimax.minimax-m2.1bedrock/ap-southeast-3/minimax.minimax-m2.1 | 196K | $0.36 | $1.44 | — | |||
| Nous: Hermes 3 405B Instructnousresearch/hermes-3-llama-3.1-405b | 131.072K | $1 | $1 | — | |||
| ap-southeast-3/minimax.minimax-m2.5bedrock/ap-southeast-3/minimax.minimax-m2.5 | 1M | $0.36 | $1.44 | — | |||
| ap-southeast-3/moonshotai.kimi-k2.5bedrock/ap-southeast-3/moonshotai.kimi-k2.5 | 262.144K | $0.72 | $3.6 | — | |||
| ap-southeast-3/qwen.qwen3-coder-nextbedrock/ap-southeast-3/qwen.qwen3-coder-next | 262.144K | $0.6 | $1.44 | — | |||
| eu-north-1/deepseek.v3.2bedrock/eu-north-1/deepseek.v3.2 | 163.84K | $0.74 | $2.22 | — | |||
| eu-north-1/minimax.minimax-m2.1bedrock/eu-north-1/minimax.minimax-m2.1 | 196K | $0.36 | $1.44 | — | |||
| eu-north-1/minimax.minimax-m2.5bedrock/eu-north-1/minimax.minimax-m2.5 | 1M | $0.36 | $1.44 | — | |||
| eu-north-1/moonshotai.kimi-k2.5bedrock/eu-north-1/moonshotai.kimi-k2.5 | 262.144K | $0.72 | $3.6 | — | |||
| OpenAI: gpt-oss-20b (batch)openai/gpt-oss-20b:batch | 131.072K | $0.05 | $0.2 | — | |||
| moonshotai/Kimi-K3nebius/moonshotai/kimi-k3 | 1.024M | $3 | $15 | — | |||
| Nous: Hermes 3 70B Instructnousresearch/hermes-3-llama-3.1-70b | 131.072K | $0.7 | $0.7 | — | |||
| Sao10K: Llama 3.1 Euryale 70B v2.2sao10k/l3.1-euryale-70b | 131.072K | $0.85 | $0.85 | — | |||
| moonshotai/Kimi-K2.7-Codenebius/moonshotai/kimi-k2.7-code | 262.144K | $0.95 | $4 | — | |||
| qwen3-coder-plus-2025-07-22qwen_ai_platform/qwen3-coder-plus-2025-07-22 | 997.952K | — | — | — | |||
| Qwen: Qwen3.5 Plus 2026-04-20qwen/qwen3.5-plus-20260420 | 1M | $0.3 | $1.8 | — | |||
| eu-central-1/minimax.minimax-m2.1bedrock/eu-central-1/minimax.minimax-m2.1 | 196K | $0.36 | $1.44 | — | |||
| eu-central-1/minimax.minimax-m2.5bedrock/eu-central-1/minimax.minimax-m2.5 | 1M | $0.36 | $1.44 | — | |||
| eu-central-1/qwen.qwen3-coder-nextbedrock/eu-central-1/qwen.qwen3-coder-next | 262.144K | $0.6 | $1.44 | — | |||
| eu-west-1/minimax.minimax-m2.1bedrock/eu-west-1/minimax.minimax-m2.1 | 196K | $0.36 | $1.44 | — | |||
| eu-west-1/minimax.minimax-m2.5bedrock/eu-west-1/minimax.minimax-m2.5 | 1M | $0.36 | $1.44 | — | |||
| Claude Opus 5 (batch)anthropic/claude-opus-5:batch | 1M | $2.5 | $12.5 | — | |||
| SpaceXAI: Grok 4.5x-ai/grok-4.5 | 500K | $2 | $6 | — | |||
| SpaceXAI: Grok 4.3 (batch)x-ai/grok-4.3:batch | 1M | $1 | $2 | — | |||
| Amazon: Nova Micro 1.0amazon/nova-micro-v1 | 128K | $0.035 | $0.14 | — | |||
| moonshotai/Kimi-K2.6nebius/moonshotai/kimi-k2.6 | 262.144K | $0.95 | $4 | — | |||
| Inception: Mercury 2.5inception/mercury-2.5 | 260K | $0.04 | $0.15 | — | |||
| Amazon: Nova Lite 1.0amazon/nova-lite-v1 | 300K | $0.06 | $0.24 | — | |||
| eu-west-2/minimax.minimax-m2.1bedrock/eu-west-2/minimax.minimax-m2.1 | 196K | $0.47 | $1.86 | — | |||
| eu-west-2/minimax.minimax-m2.5bedrock/eu-west-2/minimax.minimax-m2.5 | 1M | $0.47 | $1.86 | — | |||
| Meta: Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | 131.072K | $0.1 | $0.32 | — |