Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B and support for reasoning via Chain-of-Thought. It offers competitive benchmark...
GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K...
No provider description is available for this model yet.
No provider description is available for this model yet.
Fast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode
No provider description is available for this model yet.
Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and multilingual text analysis. Its smaller size allows it to operate...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Inflection 3 Productivity is optimized for following instructions. It is better for tasks requiring JSON output or precise adherence to provided guidelines. It has access to recent news. For emotional...
Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It excels in tasks such as image captioning and...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather...
Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts. It is suited for Japanese instruction following,...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5 | 200K | $1 | $5 | — | |||
| claude-fable-5azure_ai/claude-fable-5 | 1M | $10 | $50 | — | |||
| Qwen/Qwen-AgentWorld-35B-A3BQwen/Qwen-AgentWorld-35B-A3B | Not documented | — | — | — | |||
| claude-opus-5azure_ai/claude-opus-5 | 1M | $5 | $25 | — | |||
| Tencent: Hunyuan A13B Instructtencent/hunyuan-a13b-instruct | 131.072K | $0.14 | $0.57 | — | |||
| OpenAI: GPT-5.4 Pro (batch)openai/gpt-5.4-pro:batch | 1.05M | $15 | $90 | — | |||
| mistralai/Mistral-7B-Instruct-v0.2mistralai/Mistral-7B-Instruct-v0.2 | Not documented | — | — | — | |||
| mistralai/Mistral-7B-Instruct-v0.1mistralai/Mistral-7B-Instruct-v0.1 | Not documented | — | — | — | |||
| Anthropic: Claude Opus 4.7 (Fast)anthropic/claude-opus-4.7-fast | 1M | $30 | $150 | — | |||
| Qwen/WebWorld-14BQwen/WebWorld-14B | Not documented | — | — | — | |||
| Meta: Llama 3.2 1B Instructmeta-llama/llama-3.2-1b-instruct | 60K | $0.027 | $0.201 | — | |||
| Qwen/WebWorld-8BQwen/WebWorld-8B | Not documented | — | — | — | |||
| Qwen/Qwen3.5-397B-A17BQwen/Qwen3.5-397B-A17B | Not documented | — | — | — | |||
| claude-opus-4-8azure_ai/claude-opus-4-8 | 1M | $5 | $25 | — | |||
| Inflection: Inflection 3 Productivityinflection/inflection-3-productivity | 8K | $2.5 | $10 | — | |||
| Meta: Llama 3.2 11B Vision Instructmeta-llama/llama-3.2-11b-vision-instruct | 131.072K | $0.345 | $0.345 | — | |||
| anthropic.claude-opus-4-5-20251101-v1:0bedrock_converse/anthropic.claude-opus-4-5-20251101-v1:0 | 200K | $5 | $25 | — | |||
| claude-opus-4-1azure_ai/claude-opus-4-1 | 200K | $15 | $75 | — | |||
| Qwen/Qwen3.5-27BQwen/Qwen3.5-27B | Not documented | — | — | — | |||
| Qwen/Qwen3.6-35B-A3BQwen/Qwen3.6-35B-A3B | Not documented | — | — | — | |||
| Inference.net: Schematron V2 Turboinference-net/schematron-v2-turbo | 128K | $0.03 | $0.15 | — | |||
| Sakana: Namazusakana/namazu | 262.144K | $0.95 | $4 | — | |||
| claude-sonnet-4-5azure_ai/claude-sonnet-4-5 | 200K | $3 | $15 | — | |||
| zai-glm-4.7cerebras/zai-glm-4.7 | 128K | $2.25 | $2.75 | — | |||
| llama3ollama/llama3 | 8.192K | — | — | — | |||
| Qwen/Qwen3.5-35B-A3BQwen/Qwen3.5-35B-A3B | Not documented | — | — | — | |||
| @cf/deepseek-ai/deepseek-r1-distill-qwen-32bcloudflare/@cf/deepseek-ai/deepseek-r1-distill-qwen-32b | 80K | $0.497 | $4.881 | — | |||
| Qwen/Qwen3.5-9B-BaseQwen/Qwen3.5-9B-Base | Not documented | — | — | — | |||
| Qwen/Qwen3.5-35B-A3B-BaseQwen/Qwen3.5-35B-A3B-Base | Not documented | — | — | — | |||
| Qwen/Qwen3.5-0.8B-BaseQwen/Qwen3.5-0.8B-Base | Not documented | — | — | — | |||
| Qwen/Qwen3.5-2B-BaseQwen/Qwen3.5-2B-Base | Not documented | — | — | — | |||
| Qwen/Qwen3.5-4B-BaseQwen/Qwen3.5-4B-Base | Not documented | — | — | — | |||
| anthropic.claude-opus-4-20250514-v1:0bedrock_converse/anthropic.claude-opus-4-20250514-v1:0 | 200K | $15 | $75 | — | |||
| claude-sonnet-5azure_ai/claude-sonnet-5 | 1M | $2 | $10 | — | |||
| deepseek-ai/ESFT-token-summary-litedeepseek-ai/ESFT-token-summary-lite | Not documented | — | — | — | |||
| computer-use-previewazure/computer-use-preview | 8.192K | $3 | $12 | — | |||
| Qwen/Qwen3.5-0.8BQwen/Qwen3.5-0.8B | Not documented | — | — | — | |||
| containerazure/container | Not documented | — | — | — | |||
| gemini-3.8-flashvertex_ai/gemini-3.8-flash | 1.04858M | $0.75 | $3.75 | — | |||
| gpt-oss-120bazure_ai/gpt-oss-120b | 131.072K | $0.15 | $0.6 | — | |||
| llama2:7bollama/llama2:7b | 4.096K | — | — | — | |||
| Qwen/Qwen3.5-2BQwen/Qwen3.5-2B | Not documented | — | — | — | |||
| Qwen/Qwen3.5-4BQwen/Qwen3.5-4B | Not documented | — | — | — | |||
| @cf/meta/llama-3.1-8b-instruct-fp8cloudflare/@cf/meta/llama-3.1-8b-instruct-fp8 | 32K | $0.152 | $0.287 | — | |||
| Qwen/Qwen3.5-9BQwen/Qwen3.5-9B | Not documented | — | — | — | |||
| Qwen/Qwen3-Coder-NextQwen/Qwen3-Coder-Next | Not documented | — | — | — | |||
| Qwen/Qwen3-Coder-Next-BaseQwen/Qwen3-Coder-Next-Base | Not documented | — | — | — | |||
| Qwen/Qwen-Image-2512Qwen/Qwen-Image-2512 | Not documented | — | — | — | |||
| anthropic.claude-opus-4-1-20250805-v1:0bedrock_converse/anthropic.claude-opus-4-1-20250805-v1:0 | 200K | $15 | $75 | — | |||
| meta-llama/CodeLlama-34b-hfmeta-llama/CodeLlama-34b-hf | Not documented | — | — | — |