Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...
MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...
MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
| Model | Creator | Score | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|---|
| MoonshotAI: Kimi K3moonshotai/kimi-k3 | 76.2 | 1.04858M | $2.303 | $11.55 | 2026-07-16 | |||
| DeepSeek: DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | 69.1 | 1.04858M | $0.04 | $0.08 | 2026-07-31 | |||
| DeepSeek: DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 | 68.8 | 1.04858M | $0.578 | $1.734 | 2026-08-12 | |||
| MoonshotAI: Kimi K2.7 Codemoonshotai/kimi-k2.7-code | 60.8 | 262.144K | $0.71 | $3.5 | 2026-06-12 | |||
| Xiaomi: MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 60.2 | 1.04858M | $0.435 | $0.87 | 2026-04-22 | |||
| DeepSeek: DeepSeek V4 Pro 0423deepseek/deepseek-v4-pro | 59.4 | 1.024M | $0.799 | $1.598 | 2026-04-24 | |||
| Xiaomi: MiMo-V2.5xiaomi/mimo-v2.5 | 56.8 | 1.04858M | $0.14 | $0.28 | 2026-04-22 | |||
| Thinking Machines: Inkling Smallthinkingmachines/inkling-small | 52.9 | 524.288K | $0.45 | $1.2 | 2026-07-30 | |||
| Thinking Machines: Inklingthinkingmachines/inkling | 52.1 | 524.288K | $1 | $4.05 | 2026-07-15 | |||
| DeepSeek: DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | 52.0 | 1.024M | $0.067 | $0.134 | 2026-04-24 | |||
| Google: Gemma 4 31Bgoogle/gemma-4-31b-it | 43.4 | 262.144K | $0.09 | $0.34 | 2026-04-02 | |||
| StepFun: Step 3.7 Flashstepfun/step-3.7-flash | 39.6 | 256K | $0.2 | $1.15 | 2026-05-29 | |||
| Google: Gemma 4 26B A4B google/gemma-4-26b-a4b-it | 39.3 | 131.072K | $0.042 | $0.22 | 2026-04-02 | |||
| OpenAI: gpt-oss-120bopenai/gpt-oss-120b | 30.4 | 131.072K | $0.037 | $0.17 | 2025-08-05 | |||
| OpenAI: gpt-oss-20bopenai/gpt-oss-20b | 20.7 | 131.072K | $0.03 | $0.13 | 2025-08-05 | |||
| Google: Gemma 3 27Bgoogle/gemma-3-27b-it | 10.1 | 131.072K | $0.08 | $0.45 | 2025-03-12 | |||
| Google: Gemma 3 12Bgoogle/gemma-3-12b-it | 5.8 | 131.072K | $0.05 | $0.15 | 2025-03-12 | |||
| Google: Gemma 3 4Bgoogle/gemma-3-4b-it | 2.7 | 131.072K | $0.05 | $0.1 | 2025-03-12 |