No provider description is available for this model yet.
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user request. In contrast to RAG, relace-search performs agentic...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on...
No provider description is available for this model yet.
Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in collaboration with Venice.ai. This model is designed as an “uncensored” instruct-tuned LLM, preserving...
No provider description is available for this model yet.
GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
MiniMax M2-her is a dialogue-first large language model built for immersive roleplay, character-driven chat, and expressive multi-turn conversations. Designed to stay consistent in tone and personality, it supports rich message...
Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| us-gov.nvidia.nemotron-nano-12b-v2bedrock_converse/us-gov.nvidia.nemotron-nano-12b-v2 | 128K | $0.24 | $0.72 | — | |||
| FW-GLM-5.2-Fastazure_ai/fw-glm-5.2-fast | 1.04858M | $2.1 | $6.6 | — | |||
| FW-Inklingazure_ai/fw-inkling | 1.04858M | $1 | $4.05 | — | |||
| deepseek-v4-flash-0731fireworks_ai/deepseek-v4-flash-0731 | 1.04858M | $0.14 | $0.28 | — | |||
| Relace: Relace Searchrelace/relace-search | 256K | $1 | $3 | — | |||
| zai-org/GLM-5.1friendliai/zai-org/glm-5.1 | 202.752K | $1.4 | $4.4 | — | |||
| us-gov.nvidia.nemotron-nano-9b-v2bedrock_converse/us-gov.nvidia.nemotron-nano-9b-v2 | 128K | $0.072 | $0.276 | — | |||
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3Bnebius/nvidia/nvidia-nemotron-3-nano-30b-a3b | 262.144K | $0.06 | $0.24 | — | |||
| qwen3-30b-a3bqwen_ai_platform/qwen3-30b-a3b | 129.024K | — | — | — | |||
| glm-5p2-fast-usfireworks_ai/glm-5p2-fast-us | 1.04858M | $2.1 | $6.6 | — | |||
| kimi-k3fireworks_ai/kimi-k3 | 1.04858M | $3 | $15 | — | |||
| us-gov.nvidia.nemotron-super-3-120bbedrock_converse/us-gov.nvidia.nemotron-super-3-120b | 256K | $0.18 | $0.78 | — | |||
| OpenAI: GPT-5.2 Chatopenai/gpt-5.2-chat | 128K | $1.75 | $14 | — | |||
| nvidia/Riva-Translate-4B-Instruct-v2nvidia/Riva-Translate-4B-Instruct-v2 | Not documented | — | — | — | |||
| Venice: Uncensored (free)cognitivecomputations/dolphin-mistral-24b-venice-edition:free | 32.768K | Free | Free | — | |||
| us-gov.openai.gpt-oss-120b-1:0bedrock_converse/us-gov.openai.gpt-oss-120b-1:0 | 128K | $0.18 | $0.72 | — | |||
| Z.ai: GLM 4.7z-ai/glm-4.7 | 202.752K | $0.4 | $1.75 | — | |||
| kimi-k3-fastfireworks_ai/kimi-k3-fast | 1.04858M | $4.5 | $22.5 | — | |||
| FW-Kimi-K2.6azure_ai/fw-kimi-k2.6 | 262.144K | $1.045 | $4.4 | — | |||
| FW-Kimi-K2.7-Codeazure_ai/fw-kimi-k2.7-code | 262.144K | $1.05 | $4.4 | — | |||
| nvidia/Llama-3_1-Nemotron-Ultra-253B-v1nebius/nvidia/llama-3_1-nemotron-ultra-253b-v1 | 131.072K | $0.6 | $1.8 | — | |||
| qwen3p8-maxfireworks_ai/qwen3p8-max | 262.144K | $2 | $6 | — | |||
| muse-glimmer-30bfireworks_ai/muse-glimmer-30b | 131.072K | $0.35 | $1.5 | — | |||
| us-gov-west-1/nvidia.nemotron-nano-9b-v2bedrock/us-gov-west-1/nvidia.nemotron-nano-9b-v2 | 128K | $0.072 | $0.276 | — | |||
| us-gov-west-1/anthropic.claude-opus-5bedrock/us-gov-west-1/anthropic.claude-opus-5 | 1M | $6 | $30 | — | |||
| nemotron-lightning-3p5-30b-a3bfireworks_ai/nemotron-lightning-3p5-30b-a3b | 262.144K | $0.05 | $0.2 | — | |||
| us-gov-west-1/anthropic.claude-fable-5-1bedrock/us-gov-west-1/anthropic.claude-fable-5-1 | 1M | $12 | $60 | — | |||
| FW-Kimi-K3azure_ai/fw-kimi-k3 | 1.04858M | $3.3 | $16.5 | — | |||
| MiniMax: MiniMax M2.1minimax/minimax-m2.1 | 204.8K | $0.3 | $1.2 | — | |||
| FW-MiniMax-M3azure_ai/fw-minimax-m3 | 512K | $0.33 | $1.32 | — | |||
| FW-Nemotron-3-Ultra-NVFP4azure_ai/fw-nemotron-3-ultra-nvfp4 | 262.144K | $0.6 | $2.4 | — | |||
| us-gov-east-1/nvidia.nemotron-nano-9b-v2bedrock/us-gov-east-1/nvidia.nemotron-nano-9b-v2 | 128K | $0.072 | $0.276 | — | |||
| OpenAI: GPT Audioopenai/gpt-audio | 128K | $2.5 | $10 | — | |||
| accounts/fireworks/models/muse-glimmer-30bfireworks_ai/accounts/fireworks/models/muse-glimmer-30b | 131.072K | $0.35 | $1.5 | — | |||
| accounts/fireworks/models/nemotron-lightning-3p5-30b-a3bfireworks_ai/accounts/fireworks/models/nemotron-lightning-3p5-30b-a3b | 262.144K | $0.05 | $0.2 | — | |||
| grok-4.3azure_ai/grok-4.3 | 200K | $1.25 | $2.5 | — | |||
| nvidia/Cosmos3-Super-Reasonernebius/nvidia/cosmos3-super-reasoner | 262.144K | $0.1 | $0.3 | — | |||
| Anthropic: Claude Sonnet 4anthropic/claude-sonnet-4 | 200K | $3 | $15 | — | |||
| Google: Gemini 2.5 Pro Preview 05-06google/gemini-2.5-pro-preview-05-06 | 1.04858M | $1.25 | $10 | — | |||
| gpt-4-o-previewgithub_copilot/gpt-4-o-preview | 64K | — | — | — | |||
| qwen-turbo-latestqwen_ai_platform/qwen-turbo-latest | 1M | $0.05 | $0.2 | — | |||
| us-gov-east-1/anthropic.claude-opus-5bedrock/us-gov-east-1/anthropic.claude-opus-5 | 1M | $6 | $30 | — | |||
| us-east-1/minimax.minimax-m2.1bedrock/us-east-1/minimax.minimax-m2.1 | 196K | $0.3 | $1.2 | — | |||
| gpt-4.1github_copilot/gpt-4.1 | 128K | — | — | — | |||
| kimi-k2.7-codeqwencloud/kimi-k2.7-code | 229.376K | $0.95 | $4 | — | |||
| accounts/fireworks/models/qwen3p8-maxfireworks_ai/accounts/fireworks/models/qwen3p8-max | 262.144K | $2 | $6 | — | |||
| google/gemma-4-26B-A4B-itdeepinfra/google/gemma-4-26b-a4b-it | 262.144K | $0.07 | $0.34 | — | |||
| accounts/fireworks/routers/glm-5p2-fastfireworks_ai/accounts/fireworks/routers/glm-5p2-fast | 1.04858M | $2.1 | $6.6 | — | |||
| MiniMax: MiniMax M2-herminimax/minimax-m2-her | 65.536K | $0.3 | $1.2 | — | |||
| Upstage: Solar Pro 3upstage/solar-pro-3 | 131.072K | $0.15 | $0.6 | — |