Reasoning Tools JSON Open weights

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

deepseek/deepseek-v4.1-flash 2026-09-10 1.04858M context $0.15/M input $0.6/M output
17 providers
Reasoning Tools JSON Open weights

Flagship GLM model for long-horizon coding, agents, and complex project delivery

zhipuai/glm-5.3 2026-08-14 1M context $1.4/M input $4.4/M output
37 providers
Reasoning Tools JSON Open weights

Open-weight sparse MoE (2.4T total, 95B active), the open-weight twin of Qwen3.8 Max for coding, research, complex reasoning, and agentic workflows

alibaba/qwen3.8-2.4t-a95b 2026-08-12 262.144K context $2/M input $6/M output
12 providers
Reasoning Tools JSON Open weights

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

deepseek/deepseek-v4-pro-0813 2026-08-12 1.024M context $0.579/M input $1.738/M output
30 providers
Reasoning Tools JSON Open weights

Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads

nvidia/nemotron-3.5-lightning-30b-a3b 2026-08-11 262.144K context Input not listed Output not listed
2 providers
Reasoning Tools Open weights

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

nvidia/nemotron-3.5-lightning 2026-08-11 262.144K context $0.08/M input $0.2/M output
10 providers
Reasoning Tools JSON Open weights

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

deepseek/deepseek-v4-flash-0731 2026-07-31 1.04858M context $0.065/M input $0.18/M output
43 providers
Reasoning Tools JSON Open weights

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

moonshotai/kimi-k3 2026-07-16 1.04858M context $2.34/M input $11.7/M output
62 providers
Tools JSON Open weights

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

thinkingmachines/inkling 2026-07-15 1.04858M context $1/M input $4.05/M output
21 providers
Reasoning Tools JSON Open weights

Open flagship GLM for long-horizon coding agents and million-token context work

zhipuai/glm-5.2 2026-06-13 1M context $1.4/M input $4.4/M output
72 providers
Reasoning Tools JSON Open weights

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...

moonshotai/kimi-k2.7-code 2026-06-12 262.144K context $0.71/M input $3.5/M output
61 providers
Reasoning Tools Open weights

Lower-latency Kimi Code variant for interactive edits and coding-agent loops

moonshotai/kimi-k2.7-code-highspeed 2026-06-12 262.144K context $1.9/M input $8/M output
9 providers
Reasoning Tools Open weights

MiniMax multimodal model for long-context coding, perception, and agent planning

minimax/MiniMax-M3 2026-06-01 1.04858M context $0.3/M input $1.2/M output
37 providers
Tools JSON Open weights

Balanced Mistral model for enterprise assistants, multilingual work, and tools

mistral/mistral-medium-latest 2026-04-29 262.144K context $1.5/M input $7.5/M output
5 providers
Reasoning Tools JSON Open weights

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

deepseek/deepseek-v4-flash 2026-04-24 1.024M context $0.087/M input $0.174/M output
62 providers
Reasoning Tools JSON Open weights

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

deepseek/deepseek-v4-pro 2026-04-24 1.024M context $0.948/M input $1.896/M output
62 providers
Reasoning Tools JSON Open weights

DeepSeek V4 Pro initial snapshot with million-token context and support for thinking and non-thinking modes

deepseek/deepseek-v4-pro-0423 2026-04-23 1M context $1.32/M input $3.96/M output
2 providers
Reasoning Open weights

Initial DeepSeek V4 Flash snapshot for economical reasoning, coding, and million-token agent workloads

deepseek/deepseek-v4-flash-0423 2026-04-23 1M context $0.139/M input $0.278/M output
2 providers
Tools Open weights

Qwen vision-language model for visual reasoning, documents, and agent tasks

alibaba/qwen3.6-27b 2026-04-22 262.144K context $0.6/M input $3.6/M output
24 providers
Reasoning Tools JSON Open weights

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...

moonshotai/kimi-k2.6 2026-04-21 262.144K context $0.95/M input $4/M output
64 providers
Reasoning Tools JSON Open weights

Strong GLM coding model for agentic engineering, terminals, and repository generation

zhipuai/glm-5.1 2026-04-07 200K context $1.4/M input $4.4/M output
45 providers
Reasoning Tools JSON Open weights

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

google/gemma-4-26b-a4b-it 2026-04-02 131.072K context $0.042/M input $0.22/M output
22 providers
Reasoning Tools JSON Open weights

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

google/gemma-4-31b-it 2026-04-02 262.144K context $0.09/M input $0.34/M output
33 providers
Reasoning Tools Open weights

Low-latency M2.7 variant for interactive coding plans and agent loops

minimax/MiniMax-M2.7-highspeed 2026-03-18 204.8K context $0.6/M input $2.4/M output
13 providers
Reasoning Tools Open weights

Open MiniMax flagship for coding agents, office automation, and complex environments

minimax/MiniMax-M2.7 2026-03-18 204.8K context $0.3/M input $1.2/M output
35 providers
Tools Open weights

Efficient Mistral model for fast chat, extraction, and production assistants

mistral/mistral-small-latest 2026-03-16 256K context $0.15/M input $0.6/M output
5 providers
Reasoning Tools JSON Open weights

Qwen vision-language model for visual reasoning, documents, and agent tasks

alibaba/qwen3.5-122b-a10b 2026-02-23 262.144K context $0.4/M input $3.2/M output
16 providers
Reasoning Tools Open weights

Qwen instruction model for multilingual chat, reasoning, and tool use

alibaba/qwen3.5-9b 2026-02-23 262.144K context $0.04/M input $0.15/M output
16 providers
Tools Open weights

Qwen vision-language model for visual reasoning, documents, and agent tasks

alibaba/qwen3.5-27b 2026-02-23 262.144K context $0.3/M input $2.4/M output
14 providers
Reasoning Tools Open weights

High-speed MiniMax model for low-latency coding and agent workflows

minimax/MiniMax-M2.5-highspeed 2026-02-13 204.8K context $0.6/M input $2.4/M output
7 providers
Reasoning Tools Open weights

General GLM flagship for coding, analysis, and tool-heavy engineering workflows

zhipuai/glm-5 2026-02-12 204.8K context $1/M input $3.2/M output
37 providers
Reasoning Tools Open weights

Prior MiniMax coding model for agent workflows, office edits, and automation

minimax/MiniMax-M2.5 2026-02-12 204.8K context $0.3/M input $1.2/M output
34 providers
Tools JSON Open weights

Open-weight Qwen coding model for agents, repository edits, and multi-turn tool use

alibaba/qwen3-coder-next 2026-02-03 262.144K context $0.108/M input $0.675/M output
13 providers
Reasoning Tools Open weights

Efficient GLM model for fast reasoning, coding, and agent workflows

zhipuai/glm-4.7-flashx 2026-01-19 200K context $0.07/M input $0.4/M output
5 providers
Tools Open weights

Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed...

moonshotai/kimi-k2.5 2026-01 262.144K context $0.45/M input $2.25/M output
45 providers
Reasoning Tools Open weights

Earlier MiniMax agent model for practical coding and productivity tasks

minimax/MiniMax-M2.1 2025-12-23 204.8K context $0.3/M input $1.2/M output
15 providers
Reasoning Tools Open weights

Mature GLM model for dependable coding, reasoning, and structured agent tasks

zhipuai/glm-4.7 2025-12-22 204.8K context $0.6/M input $2.2/M output
30 providers
Tools Open weights

Mistral's coding-agent model for repository work, terminal tasks, and software fixes

mistral/devstral-2512 2025-12-09 262.144K context $0.4/M input $2/M output
13 providers
Tools Open weights

Mistral coding agent model for repository tasks and software engineering workflows

mistral/devstral-medium-latest 2025-12-02 262.144K context $0.4/M input $2/M output
3 providers
Reasoning Tools JSON Open weights

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

deepseek/deepseek-v3.2 2025-12-01 163.84K context $0.269/M input $0.4/M output
23 providers
Tools Open weights

Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning. Built on the trillion-parameter Mixture-of-Experts (MoE) architecture introduced in...

moonshotai/kimi-k2-thinking 2025-11-06 262.144K context $0.6/M input $2.5/M output
22 providers
Reasoning Tools JSON Open weights

Safety model for policy screening, moderation, and risk-aware routing workflows

openai/gpt-oss-safeguard-120b 2025-10-29 131.072K context $0.15/M input $0.6/M output
4 providers
Reasoning Open weights

gpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b. This open-weight, 21B-parameter Mixture-of-Experts (MoE) model offers lower latency for safety tasks like content classification, LLM filtering, and trust...

openai/gpt-oss-safeguard-20b 2025-10-29 131.072K context $0.075/M input $0.3/M output
7 providers
Reasoning Tools Open weights

Efficient open MiniMax model built for coding agents and tool-heavy workflows

minimax/MiniMax-M2 2025-10-27 204.8K context $0.3/M input $1.2/M output
13 providers
Reasoning Tools Open weights

Late GLM-4 workhorse for coding agents, reasoning, and structured tasks

zhipuai/glm-4.6 2025-09-30 204.8K context $0.6/M input $2.2/M output
18 providers
Reasoning Tools Open weights

Efficient Qwen thinking model for local reasoning, math, and coding agents

alibaba/qwen3-next-80b-a3b-thinking 2025-09 131.072K context $0.5/M input $6/M output
10 providers
Tools Open weights

Qwen instruction model for multilingual chat, reasoning, and tool use

alibaba/qwen3-next-80b-a3b-instruct 2025-09 131.072K context $0.5/M input $2/M output
12 providers
Reasoning Tools Open weights

Hybrid-reasoning DeepSeek model with thinking and non-thinking modes

deepseek/deepseek-v3.1 2025-08-21 131.072K context $0.19/M input $0.71/M output
10 providers
Reasoning Tools Open weights

Compact Nemotron model for efficient reasoning and deployable AI agents

nvidia/nemotron-nano-9b-v2 2025-08-18 131.072K context $0.06/M input $0.23/M output
4 providers
Reasoning Tools Open weights

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...

openai/gpt-oss-20b 2025-08-05 131.072K context $0.03/M input $0.13/M output
32 providers