200 models
Reasoning Tools JSON Open weights

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

deepseek/deepseek-v4.1-flash 2026-09-10 1.04858M context $0.15/M input $0.6/M output
23 providers
Reasoning Tools JSON

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

openai/gpt-6-astra 2026-09-04 1.05M context $10/M input $50/M output
18 providers
Reasoning Tools JSON

2026-09-02 upgraded snapshot of Qwen3.8 Max with stronger coding, collaborative agents, and multimodal document understanding

alibaba/qwen3.8-max-0902 2026-09-02 1M context $1.71/M input $5.14/M output
7 providers
Reasoning Tools JSON

Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through...

meta/muse-spark-1.3 2026-09-02 1.04858M context $1.25/M input $4.25/M output
9 providers
Reasoning Tools JSON

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

google/gemini-3.8-flash 2026-09-02 1.04858M context $0.75/M input $3.75/M output
17 providers
Reasoning Tools JSON

Claude model for demanding reasoning and long-horizon agentic work

anthropic/claude-fable-5-1 2026-09-01 1M context $10/M input $50/M output
24 providers
Reasoning Tools JSON Open weights

Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that...

tencent/hy4-preview 2026-08-28 1.04858M context $0.834/M input $2.501/M output
9 providers
Reasoning Tools

Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...

inclusionai/ling-3.0-flash-fin 2026-08-27 262.144K context $0.06/M input $0.18/M output
3 providers
Reasoning Tools JSON

Qwen vision-language model for visual reasoning, documents, and agent tasks

alibaba/qwen3.8-flash 2026-08-26 1M context $0.15/M input $0.47/M output
21 providers
Reasoning Tools JSON Open weights

Native multimodal GLM model for efficient coding and long-horizon agent tasks

zhipuai/glm-5.3-flash 2026-08-26 1M context $0.075/M input $0.25/M output
36 providers
Reasoning Tools JSON

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...

deepseek/deepseek-v4-flash-vision-exp 2026-08-21 1.04858M context $0.22/M input $0.66/M output
19 providers
Reasoning Tools JSON Open weights

Flagship GLM model for long-horizon coding, agents, and complex project delivery

zhipuai/glm-5.3 2026-08-14 1M context $1.4/M input $4.4/M output
38 providers
Reasoning Tools JSON Open weights

Dense 27B vision-language model for coding, agent tasks, and image and video understanding

alibaba/qwen3.8-27b 2026-08-14 262.144K context $0.1/M input $0.4/M output
27 providers
Reasoning Tools JSON

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

google/gemini-3.7-flash 2026-08-13 1.04858M context $0.75/M input $3.75/M output
23 providers
Reasoning Tools JSON Open weights

Open-weight sparse MoE (2.4T total, 95B active), the open-weight twin of Qwen3.8 Max for coding, research, complex reasoning, and agentic workflows

alibaba/qwen3.8-2.4t-a95b 2026-08-12 262.144K context $2/M input $6/M output
12 providers
Reasoning Tools JSON

xAI's frontier model for long-running agents, coding, knowledge work, and visual projects

xai/grok-4.6 2026-08-12 500K context $2/M input $6/M output
27 providers
Reasoning Tools JSON Open weights

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

deepseek/deepseek-v4-pro-0813 2026-08-12 1.024M context $0.579/M input $1.738/M output
30 providers
Reasoning Tools Open weights

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

nvidia/nemotron-3.5-lightning 2026-08-11 262.144K context $0.08/M input $0.2/M output
10 providers
Reasoning Tools JSON Open weights

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon...

meta/muse-glimmer-30b 2026-08-10 131.072K context $0.3/M input $1.1/M output
9 providers
Tools JSON

Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive...

upstage/solar-pro4 2026-08-06 524.288K context $0.09/M input $0.36/M output
5 providers
Reasoning Tools JSON

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context...

meta/muse-spark-1.2 2026-08-05 1.04858M context $1.25/M input $4.25/M output
14 providers
Reasoning Tools JSON

Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts. It is suited for Japanese instruction following,...

sakana/sakana-namazu 2026-08-03 262.144K context $0.95/M input $4/M output
4 providers
Tools JSON

2.4-trillion-parameter MoE flagship for coding, professional work, multimodal understanding, and long-horizon agentic workflows

alibaba/qwen3.8-max 2026-08-03 1M context $2/M input $6/M output
28 providers
Reasoning Tools JSON Open weights

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

deepseek/deepseek-v4-flash-0731 2026-07-31 1.04858M context $0.065/M input $0.18/M output
43 providers
Tools Open weights

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

thinkingmachines/inkling-small 2026-07-30 524.288K context $0.45/M input $1.2/M output
11 providers
Reasoning Tools

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

anthropic/claude-opus-5 2026-07-24 1M context $5/M input $25/M output
32 providers
Reasoning Tools Open weights

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...

poolside/laguna-s-2.1 2026-07-21 1.04858M context $0.09/M input $0.18/M output
7 providers
Reasoning Tools JSON

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

google/gemini-3.6-flash 2026-07-21 1.04858M context $0.75/M input $3.75/M output
24 providers
Reasoning Tools JSON

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

google/gemini-3.5-flash-lite 2026-07-21 1.04858M context $0.3/M input $2.5/M output
23 providers
Reasoning Tools JSON Open weights

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

moonshotai/kimi-k3 2026-07-16 1.04858M context $1.796/M input $9.006/M output
63 providers
Tools JSON Open weights

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

thinkingmachines/inkling 2026-07-15 1.04858M context $1/M input $4.05/M output
21 providers
Reasoning Tools JSON

Lightweight multimodal Qwen model for high-throughput text, image, and video tasks

alibaba/qwen3.7-flash 2026-07-15 1M context $0.028/M input $0.113/M output
11 providers
Reasoning Tools JSON

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

openai/gpt-5.6-luna 2026-07-09 1.05M context $0.2/M input $1.2/M output
36 providers
Reasoning Tools JSON

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

openai/gpt-5.6-sol 2026-07-09 1.05M context $2/M input $10/M output
36 providers
Reasoning Tools JSON

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

openai/gpt-5.6-terra 2026-07-09 1.05M context $2/M input $12/M output
35 providers
Reasoning Tools JSON

xAI's Grok model for chat, coding, agentic tools, and lower hallucination risk

xai/grok-4.5 2026-07-08 500K context $2/M input $6/M output
26 providers
Reasoning Tools JSON Open weights

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...

tencent/hy3 2026-07-06 262.144K context $0.083/M input $0.33/M output
17 providers
Reasoning Tools Open weights

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...

poolside/laguna-xs-2.1 2026-07-02 262.144K context $0.06/M input $0.12/M output
6 providers
Reasoning Tools

LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic...

meituan/longcat-2.0 2026-06-30 1.04876M context $0.3/M input $1.2/M output
5 providers
Tools JSON

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

anthropic/claude-sonnet-5 2026-06-30 1M context $2/M input $10/M output
35 providers
Reasoning Tools

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation...

google/gemini-3.1-flash-lite-image 2026-06-30 65.536K context $0.25/M input $1.5/M output
5 providers
Reasoning Tools JSON

Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...

sakana/fugu-ultra 2026-06-15 1M context $5/M input $30/M output
11 providers
Reasoning Tools JSON Open weights

Open flagship GLM for long-horizon coding agents and million-token context work

zhipuai/glm-5.2 2026-06-13 1M context $1.4/M input $4.4/M output
72 providers
Reasoning Tools JSON Open weights

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...

moonshotai/kimi-k2.7-code 2026-06-12 262.144K context $0.71/M input $3.5/M output
61 providers
Reasoning Tools

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

anthropic/claude-fable-5 2026-06-09 1M context $10/M input $50/M output
37 providers
Reasoning Open weights

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

nvidia/nemotron-3.5-content-safety 2026-06-04 131.072K context $0.2/M input $0.2/M output
3 providers
Reasoning Tools Open weights

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

nvidia/nemotron-3-ultra-550b-a55b 2026-06-04 256K context $0.625/M input $3.125/M output
15 providers
Reasoning Tools

Multimodal Qwen workhorse for long-context agents, visual inputs, and coding

alibaba/qwen3.7-plus 2026-06-02 1M context $0.5/M input $3/M output
29 providers
Reasoning Tools Open weights

MiniMax multimodal model for long-context coding, perception, and agent planning

minimax/MiniMax-M3 2026-06-01 1.04858M context $0.3/M input $1.2/M output
37 providers
Reasoning Tools JSON Open weights

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...

stepfun/step-3.7-flash 2026-05-29 256K context $0.2/M input $1.15/M output
17 providers