Reasoning Tools Open weights

Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads

nvidia/nemotron-3.5-lightning 2026-08-11 262.144K context $0.05/M input $0.2/M output
10 providers
Reasoning Tools Open weights

Agentic coding model from Poolside in the XS size class for local deployment

poolside/laguna-s-2.1 2026-07-21 1.04858M context $0.09/M input $0.18/M output
7 providers
Reasoning Tools JSON Open weights

Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work

moonshotai/kimi-k3 2026-07-16 1.04858M context $3/M input $15/M output
62 providers
Reasoning Tools JSON Open weights

Open flagship GLM for long-horizon coding agents and million-token context work

zhipuai/glm-5.2 2026-06-13 1M context $1.4/M input $4.4/M output
72 providers
Reasoning Tools JSON Open weights

Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking

moonshotai/kimi-k2.7-code 2026-06-12 262.144K context $0.95/M input $4/M output
61 providers
Reasoning Tools Open weights

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

nvidia/nemotron-3-ultra-550b-a55b 2026-06-04 1M context $0.5/M input $2.5/M output
15 providers
Reasoning Tools Open weights

MiniMax multimodal model for long-context coding, perception, and agent planning

minimax/MiniMax-M3 2026-06-01 1.04858M context $0.3/M input $1.2/M output
37 providers
Reasoning Tools JSON Open weights

Balanced Mistral model for enterprise assistants, multilingual work, and tools

mistral/mistral-medium-2604 2026-04-29 262.144K context $1.5/M input $7.5/M output
9 providers
Reasoning Tools JSON Open weights

Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work

deepseek/deepseek-v4-flash 2026-04-24 1M context $0.15/M input $0.6/M output
62 providers
Reasoning Tools JSON Open weights

Open MoE flagship with million-token context for coding and long agent runs

deepseek/deepseek-v4-pro 2026-04-24 1M context $0.435/M input $0.87/M output
62 providers
Reasoning Tools Open weights

Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution

xiaomi/mimo-v2.5-pro 2026-04-22 1.04858M context $0.435/M input $0.87/M output
24 providers
Reasoning Tools JSON Open weights

Open MiMo model for multimodal coding agents and long-context automation

xiaomi/mimo-v2.5 2026-04-22 1.04858M context $0.14/M input $0.28/M output
23 providers
Tools Open weights

Qwen vision-language model for visual reasoning, documents, and agent tasks

alibaba/qwen3.6-27b 2026-04-22 262.144K context $0.6/M input $3.6/M output
24 providers
Reasoning Tools JSON Open weights

Multimodal Kimi workhorse for agent loops, coding tasks, and visual context

moonshotai/kimi-k2.6 2026-04-21 262.144K context $0.95/M input $4/M output
64 providers
Reasoning Tools JSON Open weights

Open multimodal Qwen MoE for local agents that need vision, audio, and code

alibaba/qwen3.6-35b-a3b 2026-04-17 262.144K context $0.248/M input $1.485/M output
19 providers
Reasoning Tools JSON Open weights

Strong GLM coding model for agentic engineering, terminals, and repository generation

zhipuai/glm-5.1 2026-04-07 200K context $1.4/M input $4.4/M output
45 providers
Reasoning Tools JSON Open weights

Largest Gemma 4 instruction model for open, self-hosted chat and reasoning

google/gemma-4-31b-it 2026-04-02 262.144K context $0.09/M input $0.34/M output
33 providers
Reasoning Tools JSON Open weights

Open Gemma instruction model for efficient chat and self-hosted deployments

google/gemma-4-E2B-it 2026-04-02 131.072K context $0.04/M input $0.08/M output
2 providers
Reasoning Tools JSON Open weights

Open Gemma instruction model for efficient chat and self-hosted deployments

google/gemma-4-E4B-it 2026-04-02 131.072K context $0.02/M input $0.1/M output
2 providers
Reasoning Tools Open weights

Open MiniMax flagship for coding agents, office automation, and complex environments

minimax/MiniMax-M2.7 2026-03-18 204.8K context $0.3/M input $1.2/M output
35 providers
Reasoning Tools Open weights

Fast Mistral production model for chat, extraction, and cost-sensitive agents

mistral/mistral-small-2603 2026-03-16 256K context $0.15/M input $0.6/M output
12 providers
Reasoning Open weights

Nemotron middle tier for collaborative agents and high-volume reasoning workloads

nvidia/nemotron-3-super-120b-a12b 2026-03-11 262.144K context $0.2/M input $0.8/M output
15 providers
Reasoning Tools Open weights

Qwen instruction model for multilingual chat, reasoning, and tool use

alibaba/qwen3.5-9b 2026-02-23 262.144K context $0.04/M input $0.15/M output
16 providers
Open weights

Small Nemotron 3 MoE for efficient coding, math, and long-context agents

nvidia/nemotron-3-nano-30b-a3b 2025-12-15 262.144K context $0.05/M input $0.2/M output
11 providers
Tools Open weights

Mistral's coding-agent model for repository work, terminal tasks, and software fixes

mistral/devstral-2512 2025-12-09 262.144K context $0.4/M input $2/M output
13 providers
Reasoning Tools Open weights

Hybrid-reasoning DeepSeek model with thinking and non-thinking modes

deepseek/deepseek-v3.1 2025-08-21 131.072K context $0.19/M input $0.71/M output
10 providers
Reasoning Tools JSON Open weights

Open GPT reasoning model for self-hosted agents and controllable deployments

openai/gpt-oss-120b 2025-08-05 131.072K context $0.03/M input $0.17/M output
53 providers
Reasoning Tools Open weights

Open GPT reasoning model for self-hosted agents and controllable deployments

openai/gpt-oss-20b 2025-08-05 131.072K context $0.02/M input $0.1/M output
32 providers
Tools Open weights

Updated large open Qwen3 MoE instruct model for multilingual chat, coding, and tool use

alibaba/qwen3-235b-a22b-instruct-2507 2025-07-21 262.144K context $0.069/M input $0.455/M output
6 providers
Reasoning Tools Open weights

Open Mistral reasoning model for transparent step-by-step problem solving

mistral/magistral-small-2506 2025-06-10 131.072K context $0.5/M input $1.5/M output
2 providers
Reasoning Tools JSON Open weights

Dense open Qwen model for self-hosted chat, reasoning, and coding

alibaba/qwen3-32b 2025-04 131.072K context $0.7/M input $2.8/M output
14 providers
Tools Open weights

Open DeepSeek MoE chat model for coding, math, and general reasoning

deepseek/deepseek-v3 2024-12-26 131.072K context $0.27/M input $1.12/M output
8 providers
Tools Open weights

Popular open Llama workhorse for multilingual chat, coding, and self-hosting

meta/llama-3.3-70b-instruct 2024-12-06 128K context $0.1/M input $0.32/M output
25 providers
Open weights

Tiny open Qwen code model for lightweight completion and on-device coding

alibaba/qwen2.5-coder-0.5b 2024-11-12 32.768K context $0.1/M input $0.1/M output
1 provider
Open weights

Mistral's largest general model for enterprise agents, coding, and multilingual reasoning

mistral/mistral-large-2512 2024-11-01 262.144K context $0.5/M input $1.5/M output
12 providers
Tools Open weights

Efficient open Mistral edge model for on-device chat and function calling

mistral/ministral-8b-instruct-2410 2024-10-16 131.072K context $0.15/M input $0.15/M output
1 provider
Open weights

Small open Llama base model for lightweight text generation and self-hosting

meta/llama-3.2-3b 2024-09-25 131.072K context $0.1/M input $0.1/M output
1 provider
Open weights

Compact open Llama base model for lightweight and on-device use

meta/llama-3.2-1b 2024-09-25 131.072K context $0.1/M input $0.1/M output
1 provider
Tools Open weights

Mistral vision-language model for image understanding and multimodal chat

mistral/pixtral-12b 2024-09-01 128K context $0.15/M input $0.15/M output
4 providers
Open weights

Efficient Mistral-NVIDIA open model for multilingual chat and local deployment

mistral/mistral-nemo 2024-07-01 128K context $0.15/M input $0.15/M output
11 providers
Open weights

Open Mistral code model for fill-in-the-middle and 80+ programming languages

mistral/codestral-22b-v0.1 2024-05-29 32.768K context $0.3/M input $0.9/M output
1 provider