No provider description is available for this model yet.
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...
No provider description is available for this model yet.
Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization. Designed with the latest transformer architecture, it...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....
No provider description is available for this model yet.
No provider description is available for this model yet.
Schematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex schemas and long pages. Extraction instructions must be supplied through a JSON schema...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous programming via tool calling and...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| us-west-2/anthropic.claude-v2:1bedrock/us-west-2/anthropic.claude-v2:1 | 100K | $8 | $24 | — | |||
| Google: Gemini 2.5 Flash Lite (batch)google/gemini-2.5-flash-lite:batch | 1.04858M | $0.05 | $0.2 | — | |||
| us-gov-east-1/nvidia.nemotron-nano-3-30bbedrock/us-gov-east-1/nvidia.nemotron-nano-3-30b | 262.144K | $0.072 | $0.288 | — | |||
| Meta: Llama 3.2 3B Instruct (free)meta-llama/llama-3.2-3b-instruct:free | 131.072K | Free | Free | — | |||
| us-west-2/mistral.mistral-7b-instruct-v0:2bedrock/us-west-2/mistral.mistral-7b-instruct-v0:2 | 32K | $0.15 | $0.2 | — | |||
| gpt-5.4-2026-03-05azure_ai/gpt-5.4-2026-03-05 | 1.05M | $2.5 | $15 | — | |||
| us-gov-east-1/nvidia.nemotron-nano-12b-v2bedrock/us-gov-east-1/nvidia.nemotron-nano-12b-v2 | 128K | $0.24 | $0.72 | — | |||
| us-west-2/mistral.mistral-large-2402-v1:0bedrock/us-west-2/mistral.mistral-large-2402-v1:0 | 32K | $8 | $24 | — | |||
| us-gov-east-1/nvidia.nemotron-super-3-120bbedrock/us-gov-east-1/nvidia.nemotron-super-3-120b | 256K | $0.18 | $0.78 | — | |||
| us-west-2/mistral.mixtral-8x7b-instruct-v0:1bedrock/us-west-2/mistral.mixtral-8x7b-instruct-v0:1 | 32K | $0.45 | $0.7 | — | |||
| ByteDance Seed: Seed 2.1 Turbobytedance-seed/seed-2-1-turbo | 262.144K | $0.5 | $2.5 | — | |||
| gpt-5.4-miniazure_ai/gpt-5.4-mini | 272K | $0.75 | $4.5 | — | |||
| us-gov-east-1/openai.gpt-oss-20b-1:0bedrock/us-gov-east-1/openai.gpt-oss-20b-1:0 | 128K | $0.084 | $0.36 | — | |||
| gpt-5.4-mini-2026-03-17azure_ai/gpt-5.4-mini-2026-03-17 | 272K | $0.75 | $4.5 | — | |||
| us-gov-east-1/openai.gpt-oss-120b-1:0bedrock/us-gov-east-1/openai.gpt-oss-120b-1:0 | 128K | $0.18 | $0.72 | — | |||
| gpt-5.4-nanoazure_ai/gpt-5.4-nano | 272K | $0.2 | $1.25 | — | |||
| us-gov-east-1/anthropic.claude-sonnet-5bedrock/us-gov-east-1/anthropic.claude-sonnet-5 | 1M | $2.4 | $12 | — | |||
| gpt-5.4-nano-2026-03-17azure_ai/gpt-5.4-nano-2026-03-17 | 272K | $0.2 | $1.25 | — | |||
| us-gov-east-1/anthropic.claude-opus-4-8bedrock/us-gov-east-1/anthropic.claude-opus-4-8 | 1M | $6 | $30 | — | |||
| eu/gpt-4o-2024-08-06azure/eu/gpt-4o-2024-08-06 | 128K | $2.75 | $11 | — | |||
| eu/gpt-4o-2024-11-20azure/eu/gpt-4o-2024-11-20 | 128K | $2.75 | $11 | — | |||
| @cf/meta/llama-3.2-1b-instructcloudflare/@cf/meta/llama-3.2-1b-instruct | 60K | $0.027 | $0.201 | — | |||
| eu/gpt-4o-mini-2024-07-18azure/eu/gpt-4o-mini-2024-07-18 | 128K | $0.165 | $0.66 | — | |||
| eu/gpt-5-2025-08-07azure/eu/gpt-5-2025-08-07 | 272K | $1.375 | $11 | — | |||
| eu/gpt-5-mini-2025-08-07azure/eu/gpt-5-mini-2025-08-07 | 272K | $0.275 | $2.2 | — | |||
| eu/gpt-5.1azure/eu/gpt-5.1 | 272K | $1.38 | $11 | — | |||
| eu/gpt-5.1-chatazure/eu/gpt-5.1-chat | 128K | $1.38 | $11 | — | |||
| eu/gpt-5-nano-2025-08-07azure/eu/gpt-5-nano-2025-08-07 | 272K | $0.055 | $0.44 | — | |||
| eu/o1-2024-12-17azure/eu/o1-2024-12-17 | 200K | $16.5 | $66 | — | |||
| eu/o1-mini-2024-09-12azure/eu/o1-mini-2024-09-12 | 128K | $1.21 | $4.84 | — | |||
| us-gov-west-1/xai.grok-4.3bedrock_mantle/us-gov-west-1/xai.grok-4.3 | 131.072K | $1.5 | $3 | — | |||
| eu/o1-preview-2024-09-12azure/eu/o1-preview-2024-09-12 | 128K | $16.5 | $66 | — | |||
| OpenAI: o3 (batch)openai/o3:batch | 200K | $1 | $4 | — | |||
| us-gov/gpt-5.1azure/us-gov/gpt-5.1 | 272K | $1.719 | $13.75 | — | |||
| us-gov/o3-miniazure/us-gov/o3-mini | 200K | $1.513 | $6.05 | — | |||
| Inference.net: Schematron V2 Smallinference-net/schematron-v2-small | 128K | $0.05 | $0.23 | — | |||
| eu/o3-mini-2025-01-31azure/eu/o3-mini-2025-01-31 | 200K | $1.21 | $4.84 | — | |||
| global-standard/gpt-4o-2024-08-06azure/global-standard/gpt-4o-2024-08-06 | 128K | $2.5 | $10 | — | |||
| global-standard/gpt-4o-2024-11-20azure/global-standard/gpt-4o-2024-11-20 | 128K | $2.5 | $10 | — | |||
| inclusionAI: Ling 3.0 Flash Fin (free)inclusionai/ling-3.0-flash-fin:free | 262.144K | Free | Free | — | |||
| global-standard/gpt-4o-miniazure/global-standard/gpt-4o-mini | 128K | $0.15 | $0.6 | — | |||
| global/gpt-4o-2024-08-06azure/global/gpt-4o-2024-08-06 | 128K | $2.5 | $10 | — | |||
| global/gpt-4o-2024-11-20azure/global/gpt-4o-2024-11-20 | 128K | $2.5 | $10 | — | |||
| global/gpt-5.1azure/global/gpt-5.1 | 272K | $1.25 | $10 | — | |||
| global/gpt-5.1-chatazure/global/gpt-5.1-chat | 128K | $1.25 | $10 | — | |||
| databricks-claude-fable-5databricks/databricks-claude-fable-5 | 1M | $10 | $50 | — | |||
| Qwen: Qwen3 Coder Plusqwen/qwen3-coder-plus | 1M | $0.65 | $3.25 | — | |||
| databricks-claude-opus-4-7databricks/databricks-claude-opus-4-7 | 1M | $5 | $25 | — | |||
| gpt-4-0125-previewazure/gpt-4-0125-preview | 128K | $10 | $30 | — | |||
| databricks-claude-opus-4-8databricks/databricks-claude-opus-4-8 | 1M | $5 | $25 | — |