Reasoning Tools JSON Open weights

DeepSeek V4 Pro snapshot with million-token context and support for thinking and non-thinking modes

deepseek/deepseek-v4-pro-0813 2026-08-12 1M context $0.442/M input $0.884/M output
30 providers
Reasoning Tools Open weights

Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads

nvidia/nemotron-3.5-lightning 2026-08-11 262.144K context $0.05/M input $0.2/M output
10 providers
Reasoning Tools JSON Open weights

Fast NVIDIA Nemotron MoE for reliable agentic tasks across enterprise workloads

nvidia/nemotron-3.5-lightning-30b-a3b 2026-08-11 262.144K context Input not listed Output not listed
2 providers
Reasoning Tools JSON Open weights

Muse Glimmer is a 30-billion-parameter open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark for always-on local agents, tool use, coding, and image understanding.

meta/muse-glimmer-30b 2026-08-10 131.072K context $0.2/M input $0.8/M output
9 providers
Reasoning Tools JSON Open weights

Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding

deepseek/deepseek-v4-flash-0731 2026-07-31 1M context $0.035/M input $0.07/M output
43 providers
Reasoning Tools JSON Open weights

Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work

moonshotai/kimi-k3 2026-07-16 1.04858M context $3/M input $15/M output
63 providers
Tools JSON Open weights

Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio

thinkingmachines/inkling 2026-07-15 1.04858M context $1.87/M input $4.68/M output
21 providers
Reasoning Tools Open weights

Agentic coding model from Poolside in the XS size class for local deployment

poolside/laguna-xs-2.1 2026-07-02 262.144K context $0.06/M input $0.12/M output
6 providers
Reasoning Tools JSON Open weights

Open flagship GLM for long-horizon coding agents and million-token context work

zhipuai/glm-5.2 2026-06-13 1M context $1.4/M input $4.4/M output
72 providers
Reasoning Tools Open weights

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

nvidia/nemotron-3-ultra-550b-a55b 2026-06-04 1M context $0.5/M input $2.5/M output
15 providers
Reasoning Tools Open weights

MiniMax multimodal model for long-context coding, perception, and agent planning

minimax/MiniMax-M3 2026-06-01 1.04858M context $0.3/M input $1.2/M output
37 providers
Reasoning Tools JSON Open weights

Newer StepFun flash model for faster agents, coding, and multimodal prompts

stepfun/step-3.7-flash 2026-05-29 256K context $0.185/M input $1.11/M output
17 providers
Reasoning Tools JSON Open weights

Balanced Mistral model for enterprise assistants, multilingual work, and tools

mistral/mistral-medium-2604 2026-04-29 262.144K context $1.5/M input $7.5/M output
9 providers
Reasoning Tools Open weights

Open Nemotron omni model combining reasoning with text, vision, and audio

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning 2026-04-28 256K context $0.2/M input $0.8/M output
5 providers
Reasoning Tools JSON Open weights

Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work

deepseek/deepseek-v4-flash 2026-04-24 1M context $0.15/M input $0.6/M output
62 providers
Reasoning Tools JSON Open weights

Open MoE flagship with million-token context for coding and long agent runs

deepseek/deepseek-v4-pro 2026-04-24 1M context $0.435/M input $0.87/M output
62 providers
Reasoning Tools JSON Open weights

Multimodal Kimi workhorse for agent loops, coding tasks, and visual context

moonshotai/kimi-k2.6 2026-04-21 262.144K context $0.95/M input $4/M output
64 providers
Open weights

Safety model for policy screening, moderation, and risk-aware routing workflows

nvidia/nemotron-3-content-safety 2026-04-16 128K context Input not listed Output not listed
1 provider
Reasoning Tools JSON Open weights

Largest Gemma 4 instruction model for open, self-hosted chat and reasoning

google/gemma-4-31b-it 2026-04-02 262.144K context $0.09/M input $0.34/M output
33 providers
Open weights

Reranking model for improving retrieval quality in search and recommendation systems

nvidia/llama-nemotron-rerank-vl-1b-v2 2026-03-31 128K context Input not listed Output not listed
1 provider
Reasoning Tools Open weights

Open MiniMax flagship for coding agents, office automation, and complex environments

minimax/MiniMax-M2.7 2026-03-18 204.8K context $0.3/M input $1.2/M output
35 providers
Tools Open weights

Nemotron multimodal model for visual reasoning and agentic AI workflows

nvidia/nemotron-voicechat 2026-03-16 128K context Input not listed Output not listed
1 provider
Reasoning Open weights

Nemotron middle tier for collaborative agents and high-volume reasoning workloads

nvidia/nemotron-3-super-120b-a12b 2026-03-11 262.144K context $0.2/M input $0.8/M output
15 providers
Reasoning Tools JSON Open weights

Qwen vision-language model for visual reasoning, documents, and agent tasks

alibaba/qwen3.5-122b-a10b 2026-02-23 262.144K context $0.4/M input $3.2/M output
16 providers
Open weights

Embedding model for semantic search, retrieval, clustering, and ranking pipelines

nvidia/llama-nemotron-embed-vl-1b-v2 2026-02-10 32.768K context Input not listed Output not listed
1 provider
Open weights

StepFun flash lane for quick multimodal reasoning and coding assistance

stepfun/step-3.5-flash 2026-01-29 256K context $0.1/M input $0.3/M output
14 providers
Reasoning Open weights

Safety model for policy screening, moderation, and risk-aware routing workflows

nvidia/nemotron-content-safety-reasoning-4b 2026-01-22 128K context Input not listed Output not listed
1 provider
Open weights

Small Nemotron 3 MoE for efficient coding, math, and long-context agents

nvidia/nemotron-3-nano-30b-a3b 2025-12-15 262.144K context $0.05/M input $0.2/M output
11 providers
Reasoning Tools Open weights

Nemotron multimodal model for visual reasoning and agentic AI workflows

nvidia/nemotron-nano-12b-v2-vl 2025-10-28 128K context $0.2/M input $0.6/M output
3 providers
Open weights

Safety model for policy screening, moderation, and risk-aware routing workflows

nvidia/llama-3.1-nemotron-safety-guard-8b-v3 2025-10-28 128K context Input not listed Output not listed
1 provider
Reasoning Tools Open weights

Open GPT reasoning model for self-hosted agents and controllable deployments

openai/gpt-oss-20b 2025-08-05 131.072K context $0.02/M input $0.1/M output
32 providers
Reasoning Tools JSON Open weights

Open GPT reasoning model for self-hosted agents and controllable deployments

openai/gpt-oss-120b 2025-08-05 131.072K context $0.03/M input $0.17/M output
53 providers
Reasoning Tools Open weights

Nemotron model for efficient reasoning, coding, and specialized AI agents

nvidia/llama-3.3-nemotron-super-49b-v1.5 2025-07-25 131.072K context $0.4/M input $0.4/M output
4 providers

Mistral model for multilingual chat, reasoning, and tool-assisted workflows

mistral/mistral-medium-2505 2025-05-07 131.072K context $0.4/M input $2/M output
10 providers
Tools Open weights

Nemotron model for efficient reasoning, coding, and specialized AI agents

nvidia/llama-3.1-nemotron-70b-instruct 2025-04-15 128K context Input not listed Output not listed
2 providers
Reasoning Tools Open weights

Nemotron model for efficient reasoning, coding, and specialized AI agents

nvidia/llama-3.3-nemotron-super-49b-v1 2025-04-07 131.072K context Input not listed Output not listed
1 provider
Reasoning Tools Open weights

Flagship Nemotron model for high-throughput reasoning and complex agents

nvidia/llama-3.1-nemotron-ultra-253b 2025-04-07 128K context Input not listed Output not listed
2 providers
Open weights

Open multimodal Gemma instruction model for multilingual text generation and image understanding

google/gemma-3-12b-it 2025-03-12 131.072K context $0.05/M input $0.1/M output
8 providers
Tools JSON Open weights

Open multimodal Gemma instruction model for efficient text generation and image understanding

google/gemma-3-4b-it 2025-03-12 131.072K context $0.04/M input $0.08/M output
7 providers
Tools Open weights

Popular open Llama workhorse for multilingual chat, coding, and self-hosting

meta/llama-3.3-70b-instruct 2024-12-06 128K context $0.1/M input $0.32/M output
25 providers
Tools JSON Open weights

Open multimodal Llama model for image understanding, captioning, and visual QA

meta/llama-3.2-11b-vision-instruct 2024-09-25 128K context $0.055/M input $0.055/M output
3 providers
Tools Open weights

Compact Nemotron model for efficient reasoning and deployable AI agents

nvidia/nemotron-mini-4b-instruct 2024-08-21 128K context Input not listed Output not listed
1 provider
Tools JSON Open weights

Open Llama instruction model for multilingual chat, reasoning, and coding

meta/llama-3.1-70b-instruct 2024-07-23 128K context $0.4/M input $0.4/M output
4 providers
Tools JSON Open weights

Compact open Llama model for lightweight chat, drafting, and self-hosting

meta/llama-3.1-8b-instruct 2024-07-23 128K context $0.02/M input $0.04/M output
7 providers