DeepSeek Model Family

Explore 7 curated models in the DeepSeek family and compare context, API pricing, modalities and capabilities.

7models tracked

Model directory

7 models in this view

Compare
Active

DeepSeek-V3.1 is a hybrid model that supports both thinking mode and non-thinking mode. Compared to the previous version, this upgrade brings improvements in multiple aspects:

Context
164K
Input
$0.25
Output
$0.95
Text
View model
Active

This update maintains the model's original capabilities while addressing issues reported by users, including:

Context
164K
Input
$0.27
Output
$1
Text
View model
Active

We introduce **DeepSeek-V3.2**, a model that harmonizes high computational efficiency with superior reasoning and agent performance. Our approach is built upon three key technical breakthroughs:

Context
164K
Input
$0.269
Output
$0.40
Text
View model
Active

We are excited to announce the official release of DeepSeek-V3.2-Exp, an experimental version of our model. As an intermediate step toward our next-generation architecture, V3.2-Exp builds upon V3.1-Terminus by introducing DeepSeek Sparse Attention—a sparse attention mechanism designed to explore and validate optimizations for training and inference efficiency in long-context scenarios.

Context
164K
Input
$0.27
Output
$0.41
Text
View model
Active

We present a preview version of **DeepSeek-V4** series, including two strong Mixture-of-Experts (MoE) language models — **DeepSeek-V4-Pro** with 1.6T parameters (49B activated) and **DeepSeek-V4-Flash** with 284B parameters (13B activated) — both supporting a context length of **one million tokens**.

Context
1.02M
Input
$0.0826
Output
$0.1652
Text
View model
Active

**DeepSeek-V4-Flash-0731** is the official release of **DeepSeek-V4-Flash**, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.

Context
1.05M
Input
$0.14
Output
$0.28
Text
View model
Active

We present a preview version of **DeepSeek-V4** series, including two strong Mixture-of-Experts (MoE) language models — **DeepSeek-V4-Pro** with 1.6T parameters (49B activated) and **DeepSeek-V4-Flash** with 284B parameters (13B activated) — both supporting a context length of **one million tokens**.

Context
1.05M
Input
$1.32
Output
$3.96
Text
View model
OUR METHOD

How AIToolly handles model data

Model-native facts and provider endpoint facts stay separate. Third-party catalogs support discovery and provider snapshots, while official documentation and verified model cards take priority for model facts.

01Sources and verification
02Verified
03Capability