DeepSeek-V3
DeepSeek-V3 is a Mixture-of-Experts language model with 671B total parameters and 37B activated per token. It supports a 128K context and is available in Base and Chat variants with open weights for local deployment and API use.
Visit deepseek-ai/DeepSeek-V3Overview
DeepSeek-V3 is a Mixture-of-Experts (MoE) language model with 671B total parameters and 37B activated per token. It uses Multi-head Latent Attention and DeepSeekMoE, an auxiliary-loss-free load balancing strategy, and a Multi-Token Prediction objective. It was pre-trained on 14.8 trillion tokens and then fine-tuned and reinforced, with evaluations reporting strong performance on math and code tasks and a 128K context window.
Key Features
- Mixture-of-Experts architecture with 671B total parameters and 37B activated per token.
- Multi-head Latent Attention (MLA) and DeepSeekMoE architectures.
- Auxiliary-loss-free load balancing strategy for minimizing performance degradation.
- Multi-Token Prediction (MTP) training objective that can also support speculative decoding.
- FP8 mixed precision training framework validated on an extremely large-scale model.
- Pre-trained on 14.8 trillion tokens with 128K context length and Base and Chat variants.
Use Cases
- Chatting with DeepSeek-V3 through the official website or an OpenAI-compatible API.
- Running local inference and deployment with DeepSeek-Infer Demo, SGLang, LMDeploy, TensorRT-LLM, vLLM, or LightLLM.
- Handling code and math tasks where DeepSeek-V3 reports strong benchmark performance.
- Researching MoE architectures, FP8 training, Multi-Token Prediction, and reasoning distillation from DeepSeek-R1.
Getting Started
- Clone the DeepSeek-V3 GitHub repository and navigate to the inference folder.
- Install dependencies from requirements.txt, preferably in a fresh conda or uv virtual environment.
- Download model weights from Hugging Face and place them in a local folder.
- Convert Hugging Face model weights using convert.py with the desired expert and model-parallel settings.
- Run interactive or batch inference with torchrun and generate.py using the provided config.
Deployment & Requirements
- DeepSeek-Infer Demo requires Linux with Python 3.10 only; Mac and Windows are not supported.
- DeepSeek-Infer Demo dependencies include torch 2.4.1, triton 3.0.0, transformers 4.46.3, and safetensors 0.4.5.
- FP8 weights are provided; BF16 weights can be generated with the provided fp8_cast_bf16.py conversion script.
- Hugging Face Transformers is not directly supported yet.
- SGLang supports NVIDIA and AMD GPUs and multi-node tensor parallelism in BF16 and FP8 modes.
Before You Adopt
- License: MIT. Review its terms before using, modifying, or distributing the project.
- GitHub records the last push on 2025-08-28, at least 12 months before this listing was prepared; verify the current maintenance status before adoption.
- DeepSeek-Infer Demo has platform restrictions and supports only Linux with Python 3.10; Mac and Windows are not supported.
- Hugging Face Transformers is not directly supported yet, so users need another supported inference path.
- MTP support is under active development within the community, so that capability may require community progress or contributions.
- Only FP8 weights are provided, and BF16 weights require a conversion script for experimentation.