Multi-agent model for routing expert agents across complex analytical tasks
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...
Qwen frontier model tuned for agent frameworks, coding assistants, and long tasks
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
Coding-optimized GPT model for repository edits, reviews, and agentic software work
Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token....
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...
StepFun flash model for efficient multimodal reasoning, coding, and tool use
Late GLM-4 workhorse for coding agents, reasoning, and structured tasks
Flagship Qwen3 model for coding agents, complex reasoning, and tool use
Fast Mistral production model for chat, extraction, and cost-sensitive agents
Mistral's largest general model for enterprise agents, coding, and multilingual reasoning
DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass....
Hybrid-reasoning GLM release that made the 4.5 line broadly useful
The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improve relevance & readability. It’s also better at working with uploaded...
The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respone_format. Read more [here](https://openai.com/index/introducing-structured-outputs-in-the-api/). GPT-4o ("o" for "omni") is...
Mistral's coding-agent model for repository work, terminal tasks, and software fixes
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.
GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as...
Lighter GLM-4.5 variant for fast coding assistance and cheaper agents
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Smaller Qwen coder for efficient local agents and repo-level fixes
Popular open Llama workhorse for multilingual chat, coding, and self-hosting
Sonar is lightweight, affordable, fast, and simple to use — now featuring citations and the ability to customize sources. It is designed for companies seeking to integrate lightweight question-and-answer features...
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
Note: Sonar Pro pricing includes Perplexity search pricing. See [details here](https://docs.perplexity.ai/guides/pricing#detailed-pricing-breakdown-for-sonar-reasoning-pro-and-sonar-pro) For enterprises seeking more advanced capabilities, the Sonar Pro API can handle in-depth, multi-step queries with added extensibility, like...
GLM vision model for visual reasoning, documents, and multimodal agents
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| Fugusakana/fugu | 60.1 | 1M | — | — | 2026-06-15 | |||
| Sakana: Fugu Ultrasakana/fugu-ultra | 58.7 | 1M | $5 | $30 | 2026-06-15 | |||
| Qwen3.7 Maxalibaba/qwen3.7-max | 53.5 | 1M | $2.5 | $7.5 | 2026-05-21 | |||
| Google: Gemini 2.5 Progoogle/gemini-2.5-pro | 42.8 | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| GPT-5-Codexopenai/gpt-5-codex | 40.9 | 400K | $1.1 | $9 | 2025-09-15 | |||
| StepFun: Step 3.5 Flashstepfun/step-3.5-flash | 40.4 | 262.144K | $0.1 | $0.3 | 2026-01-29 | |||
| StepFun: Step 3.7 Flashstepfun/step-3.7-flash | 40.0 | 256K | $0.2 | $1.15 | 2026-05-29 | |||
| Google: Gemini 2.5 Flashgoogle/gemini-2.5-flash | 39.4 | 1.04858M | $0.3 | $2.5 | 2025-06-17 | |||
| Step 3.5 Flash 2603stepfun/step-3.5-flash-2603 | 38.5 | 256K | $0.1 | $0.3 | 2026-04-02 | |||
| GLM-4.6zhipuai/glm-4.6 | 38.4 | 204.8K | $0.6 | $2.2 | 2025-09-30 | |||
| Qwen3 Maxalibaba/qwen3-max | 38.3 | 262.144K | $1.2 | $6 | 2025-09-23 | |||
| Mistral Small 4mistral/mistral-small-2603 | 38.0 | 256K | $0.15 | $0.6 | 2026-03-16 | |||
| Mistral Large 3mistral/mistral-large-2512 | 36.2 | 262.144K | $0.5 | $1.5 | 2024-11-01 | |||
| DeepSeek: R1deepseek/deepseek-r1 | 35.7 | 64K | $0.7 | $2.5 | 2025-01-20 | |||
| GLM-4.5zhipuai/glm-4.5 | 34.8 | 131.072K | $0.6 | $2.2 | 2025-07-28 | |||
| OpenAI: GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | 33.3 | 128K | $2.5 | $10 | 2024-11-20 | |||
| OpenAI: GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | 33.1 | 128K | $2.5 | $10 | 2024-08-06 | |||
| Devstral 2mistral/devstral-2512 | 33.1 | 262.144K | $0.4 | $2 | 2025-12-09 | |||
| Mistral Medium 3mistral/mistral-medium-2505 | 33.1 | 131.072K | $0.4 | $2 | 2025-05-07 | |||
| OpenAI: GPT-4 Turboopenai/gpt-4-turbo | 31.9 | 128K | $10 | $30 | 2023-11-06 | |||
| OpenAI: GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | 30.9 | 128K | $5 | $15 | 2024-05-13 | |||
| GLM-4.5-Airzhipuai/glm-4.5-air | 30.6 | 131.072K | $0.2 | $1.1 | 2025-07-28 | |||
| Mistral Large 2.1mistral/mistral-large-2411 | 29.2 | 131.072K | $2 | $6 | 2024-11-18 | |||
| Qwen3-Coder 30B-A3B Instructalibaba/qwen3-coder-30b-a3b-instruct | 27.8 | 262.144K | $0.45 | $2.25 | 2025-04 | |||
| Llama-3.3-70B-Instructmeta/llama-3.3-70b-instruct | 26.0 | 128K | $0.1 | $0.32 | 2024-12-06 | |||
| Perplexity: Sonarperplexity/sonar | 22.9 | 127.072K | $1 | $1 | 2024-01-01 | |||
| OpenAI: GPT-4o-miniopenai/gpt-4o-mini | 22.9 | 128K | $0.15 | $0.6 | 2024-07-18 | |||
| Perplexity: Sonar Properplexity/sonar-pro | 22.6 | 200K | $3 | $15 | 2024-01-01 | |||
| GLM-4.5Vzhipuai/glm-4.5v | 22.1 | 64K | $0.6 | $1.8 | 2025-08-11 | |||
| Google: Gemini 2.5 Flash Litegoogle/gemini-2.5-flash-lite | 19.3 | 1.04858M | $0.1 | $0.4 | 2025-06-17 |