109 models
Open weights

No provider description is available for this model yet.

deepseek-ai/Janus-Pro-1B Not documented context Input not listed Output not listed
Open weights

No provider description is available for this model yet.

deepseek-ai/Janus-Pro-7B Not documented context Input not listed Output not listed
Open weights

No provider description is available for this model yet.

deepseek-ai/JanusFlow-1.3B Not documented context Input not listed Output not listed
Open weights

No provider description is available for this model yet.

deepseek-ai/Janus-1.3B Not documented context Input not listed Output not listed

No provider description is available for this model yet.

deepseek-ai/deepseek-vl2 Not documented context Input not listed Output not listed

No provider description is available for this model yet.

deepseek-ai/DeepSeek-V2.5 Not documented context Input not listed Output not listed

No provider description is available for this model yet.

deepseek-ai/DeepSeek-V2 Not documented context Input not listed Output not listed

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

deepseek/deepseek-v4-flash-0731:batch 1.04858M context $0.11/M input $0.33/M output

No provider description is available for this model yet.

deepseek/deepseek-coder 128K context $0.14/M input $0.28/M output
Open weights

No provider description is available for this model yet.

deepseek-ai/DeepSeek-V4.1-Flash Not documented context Input not listed Output not listed

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...

deepseek/deepseek-chat-v3.1 163.84K context $0.25/M input $0.95/M output

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...

deepseek/deepseek-r1-distill-llama-70b 8.192K context $0.8/M input $0.8/M output

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...

deepseek/deepseek-r1-0528 163.84K context $0.5/M input $2.15/M output

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...

deepseek/deepseek-chat-v3-0324 163.84K context $0.25/M input $1/M output

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...

deepseek/deepseek-v4-flash-vision-exp:batch 1.04858M context $0.11/M input $0.33/M output