Fast web-grounded Sonar for current answers, citations, and lightweight retrieval
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Compact GPT model for low-latency assistance and high-volume workloads
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
No provider description is available for this model yet.
No provider description is available for this model yet.
Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)
No provider description is available for this model yet.
Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...
No provider description is available for this model yet.
No provider description is available for this model yet.
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...
OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model supports the `reasoning_effort` parameter, which can be set to...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...
No provider description is available for this model yet.
Virtuoso‑Large is Arcee's top‑tier general‑purpose LLM at 72 B parameters, tuned to tackle cross‑domain reasoning, creative writing and enterprise QA. Unlike many 70 B peers, it retains the 128 k...
No provider description is available for this model yet.
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...
Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...
The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to think harder and provide consistently...
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
No provider description is available for this model yet.
Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information...
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
No provider description is available for this model yet.
No provider description is available for this model yet.
Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
No provider description is available for this model yet.
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...
No provider description is available for this model yet.
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Sonarperplexity/sonar | 128K | $1 | $1 | 2024-01-01 | |||
| GPT-4 Turboopenai/gpt-4-turbo | 128K | $10 | $30 | 2023-11-06 | |||
| qwen3-vl-plusqwencloud/qwen3-vl-plus | 260.096K | — | — | — | |||
| qwen3-vl-32b-thinkingqwencloud/qwen3-vl-32b-thinking | 131.072K | $0.16 | $2.87 | — | |||
| qwen3-vl-32b-instructqwencloud/qwen3-vl-32b-instruct | 131.072K | $0.16 | $0.64 | — | |||
| Mistral: Ministral 3 8B 2512 (batch)mistralai/ministral-8b-2512:batch | 262.144K | $0.075 | $0.075 | — | |||
| qwen3-vl-235b-a22b-thinkingqwencloud/qwen3-vl-235b-a22b-thinking | 131.072K | $0.4 | $4 | — | |||
| deepseek-ai/DeepSeek-V4-Prodeepinfra/deepseek-ai/deepseek-v4-pro | 1.04858M | $1.3 | $2.6 | — | |||
| Mistral: Codestral 2508mistralai/codestral-2508 | 256K | $0.3 | $0.9 | — | |||
| qwen3-vl-235b-a22b-instructqwencloud/qwen3-vl-235b-a22b-instruct | 131.072K | $0.4 | $1.6 | — | |||
| Anthropic: Claude Opus 4.1anthropic/claude-opus-4.1 | 200K | $15 | $75 | — | |||
| Qwen: Qwen3.8 2.4T A95Bqwen/qwen3.8-2.4t-a95b | 1M | $2 | $6 | — | |||
| qwen3-next-80b-a3b-thinkingqwencloud/qwen3-next-80b-a3b-thinking | 262.144K | $0.15 | $1.2 | — | |||
| qwen3-next-80b-a3b-instructqwencloud/qwen3-next-80b-a3b-instruct | 262.144K | $0.15 | $1.2 | — | |||
| Anthropic: Claude Opus 4.7anthropic/claude-opus-4.7 | 1M | $5 | $25 | — | |||
| OpenAI: o3 Mini (batch)openai/o3-mini:batch | 200K | $0.55 | $2.2 | — | |||
| qwen3-max-2026-01-23qwencloud/qwen3-max-2026-01-23 | 258.048K | — | — | — | |||
| qwen3-maxqwencloud/qwen3-max | 258.048K | — | — | — | |||
| qwen3-max-previewqwencloud/qwen3-max-preview | 258.048K | — | — | — | |||
| Google: Gemini 2.5 Flash Lite (batch)google/gemini-2.5-flash-lite:batch | 1.04858M | $0.05 | $0.2 | — | |||
| qwen3-coder-plus-2025-07-22qwencloud/qwen3-coder-plus-2025-07-22 | 997.952K | — | — | — | |||
| Arcee AI: Virtuoso Largearcee-ai/virtuoso-large | 131.072K | $0.75 | $1.2 | — | |||
| accounts/fireworks/models/glm-5p3fireworks_ai/accounts/fireworks/models/glm-5p3 | 1.04858M | $1.4 | $4.4 | — | |||
| OpenAI: gpt-oss-120b (batch)openai/gpt-oss-120b:batch | 131.072K | $0.15 | $0.6 | — | |||
| OpenAI: GPT-5.6 Luna (batch)openai/gpt-5.6-luna:batch | 1.05M | $0.1 | $0.6 | — | |||
| OpenAI: GPT-5.4 Mini (batch)openai/gpt-5.4-mini:batch | 400K | $0.375 | $2.25 | — | |||
| SpaceXAI: Grok 4.20x-ai/grok-4.20 | 2M | $1.25 | $2.5 | — | |||
| OpenAI: o3 Pro (batch)openai/o3-pro:batch | 200K | $10 | $40 | — | |||
| Mistral: Mistral Large 3 2512mistralai/mistral-large-2512 | 262.144K | $0.5 | $1.5 | — | |||
| qwen3-coder-plusqwencloud/qwen3-coder-plus | 997.952K | — | — | — | |||
| SpaceXAI: Grok 4.20 Multi-Agentx-ai/grok-4.20-multi-agent | 2M | $1.25 | $2.5 | — | |||
| Anthropic: Claude Sonnet 4.6 (batch)anthropic/claude-sonnet-4.6:batch | 1M | $1.5 | $7.5 | — | |||
| zai-org/GLM-5.1deepinfra/zai-org/glm-5.1 | 202.752K | $1.05 | $3.5 | — | |||
| mistral-code-latestmistral/mistral-code-latest | 128K | $0.3 | $0.9 | — | |||
| zai-org/GLM-5deepinfra/zai-org/glm-5 | 202.752K | $0.6 | $2.08 | — | |||
| inclusionAI: Ling 3.0 Flashinclusionai/ling-3.0-flash | 262.144K | $0.021 | $0.063 | — | |||
| qwen3-coder-flash-2025-07-28qwencloud/qwen3-coder-flash-2025-07-28 | 997.952K | — | — | — | |||
| qwen3-coder-flashqwencloud/qwen3-coder-flash | 997.952K | — | — | — | |||
| Anthropic: Claude Opus 4.1 (batch)anthropic/claude-opus-4.1:batch | 200K | $7.5 | $37.5 | — | |||
| qwen3-30b-a3bqwencloud/qwen3-30b-a3b | 129.024K | — | — | — | |||
| qwen-turbo-latestqwencloud/qwen-turbo-latest | 1M | $0.05 | $0.2 | — | |||
| Qwen/Qwen3.5-122B-A10Bdeepinfra/qwen/qwen3.5-122b-a10b | 262.144K | $0.29 | $2.4 | — | |||
| Z.ai: GLM 5.3 Flash (batch)z-ai/glm-5.3-flash:batch | 1.04858M | $0.075 | $0.25 | — | |||
| NVIDIA: Nemotron 3 Ultra (batch)nvidia/nemotron-3-ultra-550b-a55b:batch | 512.288K | $0.6 | $3.6 | — | |||
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | 1M | $3 | $15 | — | |||
| qwen-turbo-2025-04-28qwencloud/qwen-turbo-2025-04-28 | 1M | $0.05 | $0.2 | — | |||
| Mistral: Mistral Large 3 2512 (batch)mistralai/mistral-large-2512:batch | 262.144K | $0.25 | $0.75 | — | |||
| Google: Gemini 3 Flash Preview (batch)google/gemini-3-flash-preview:batch | 1.04858M | $0.25 | $1.5 | — | |||
| qwen-turbo-2024-11-01qwencloud/qwen-turbo-2024-11-01 | 1M | $0.05 | $0.2 | — | |||
| Google: Gemini 3.5 Flash (batch)google/gemini-3.5-flash:batch | 1.04858M | $0.75 | $4.5 | — |