2,824 models
JSON

Fast web-grounded Sonar for current answers, citations, and lightweight retrieval

perplexity/sonar 2024-01-01 128K context $1/M input $1/M output
9 providers

Compact GPT model for low-latency assistance and high-volume workloads

openai/gpt-4-turbo 2023-11-06 128K context $10/M input $30/M output
15 providers

No provider description is available for this model yet.

qwencloud/qwen3-vl-plus 260.096K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen3-vl-32b-thinking 131.072K context $0.16/M input $2.87/M output

No provider description is available for this model yet.

qwencloud/qwen3-vl-32b-instruct 131.072K context $0.16/M input $0.64/M output

A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

mistralai/ministral-8b-2512:batch 262.144K context $0.075/M input $0.075/M output

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)

mistralai/codestral-2508 256K context $0.3/M input $0.9/M output

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...

anthropic/claude-opus-4.1 200K context $15/M input $75/M output

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

qwen/qwen3.8-2.4t-a95b 1M context $2/M input $6/M output

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

anthropic/claude-opus-4.7 1M context $5/M input $25/M output

OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model supports the `reasoning_effort` parameter, which can be set to...

openai/o3-mini:batch 200K context $0.55/M input $2.2/M output

No provider description is available for this model yet.

qwencloud/qwen3-max-2026-01-23 258.048K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen3-max 258.048K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen3-max-preview 258.048K context Input not listed Output not listed

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...

google/gemini-2.5-flash-lite:batch 1.04858M context $0.05/M input $0.2/M output

No provider description is available for this model yet.

qwencloud/qwen3-coder-plus-2025-07-22 997.952K context Input not listed Output not listed

Virtuoso‑Large is Arcee's top‑tier general‑purpose LLM at 72 B parameters, tuned to tackle cross‑domain reasoning, creative writing and enterprise QA. Unlike many 70 B peers, it retains the 128 k...

arcee-ai/virtuoso-large 131.072K context $0.75/M input $1.2/M output

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...

openai/gpt-oss-120b:batch 131.072K context $0.15/M input $0.6/M output

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

openai/gpt-5.6-luna:batch 1.05M context $0.1/M input $0.6/M output

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...

openai/gpt-5.4-mini:batch 400K context $0.375/M input $2.25/M output

Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...

x-ai/grok-4.20 2M context $1.25/M input $2.5/M output

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to think harder and provide consistently...

openai/o3-pro:batch 200K context $10/M input $40/M output

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

mistralai/mistral-large-2512 262.144K context $0.5/M input $1.5/M output

No provider description is available for this model yet.

qwencloud/qwen3-coder-plus 997.952K context Input not listed Output not listed

Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information...

x-ai/grok-4.20-multi-agent 2M context $1.25/M input $2.5/M output

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...

anthropic/claude-sonnet-4.6:batch 1M context $1.5/M input $7.5/M output

No provider description is available for this model yet.

deepinfra/zai-org/glm-5.1 202.752K context $1.05/M input $3.5/M output

No provider description is available for this model yet.

mistral/mistral-code-latest 128K context $0.3/M input $0.9/M output

No provider description is available for this model yet.

deepinfra/zai-org/glm-5 202.752K context $0.6/M input $2.08/M output

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...

inclusionai/ling-3.0-flash 262.144K context $0.021/M input $0.063/M output

No provider description is available for this model yet.

qwencloud/qwen3-coder-flash 997.952K context Input not listed Output not listed

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...

anthropic/claude-opus-4.1:batch 200K context $7.5/M input $37.5/M output

No provider description is available for this model yet.

qwencloud/qwen3-30b-a3b 129.024K context Input not listed Output not listed

No provider description is available for this model yet.

qwencloud/qwen-turbo-latest 1M context $0.05/M input $0.2/M output

No provider description is available for this model yet.

deepinfra/qwen/qwen3.5-122b-a10b 262.144K context $0.29/M input $2.4/M output

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

z-ai/glm-5.3-flash:batch 1.04858M context $0.075/M input $0.25/M output

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

nvidia/nemotron-3-ultra-550b-a55b:batch 512.288K context $0.6/M input $3.6/M output

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

anthropic/claude-sonnet-4.5 1M context $3/M input $15/M output

No provider description is available for this model yet.

qwencloud/qwen-turbo-2025-04-28 1M context $0.05/M input $0.2/M output

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

mistralai/mistral-large-2512:batch 262.144K context $0.25/M input $0.75/M output

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...

google/gemini-3-flash-preview:batch 1.04858M context $0.25/M input $1.5/M output

No provider description is available for this model yet.

qwencloud/qwen-turbo-2024-11-01 1M context $0.05/M input $0.2/M output

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

google/gemini-3.5-flash:batch 1.04858M context $0.75/M input $4.5/M output