
Note that this is a base model mostly meant for testing, you need to provide detailed prompts for the model to return useful responses.
DeepSeek-V3 Base is a 671B parameter open Mixture-of-Experts (MoE) language model with 37B active parameters per forward pass and a context length of 128K tokens. Trained on 14.8T tokens using FP8 mixed precision, it achieves high training efficiency and stability, with strong performance across language, reasoning, math, and coding tasks.
DeepSeek-V3 Base is the pre-trained model behind DeepSeek V3
Modalities
Context
131K
Released
Mar 29, 2025
Knowledge Cutoff
Jul 2024
Note that this is a base model mostly meant for testing, you need to provide detailed prompts for the model to return useful responses. DeepSeek-V3 Base is a 671B parameter open Mixture-of-Experts (MoE) language model with 37B active parameters per forward pass and a context length of 128K tokens.
DeepSeek V3 Base has a 131,072 token context window.
DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731, DeepSeek V4 Pro 0423 and 9 more are other text models from DeepSeek.
DeepSeek V3 Base was released on March 29, 2025. Its knowledge cutoff is July 31, 2024.
Token volume and request traffic to this model over time.