Skip to main content
AdOpenFree logoPromote your productReach more potential users and drive product growth and revenue.Advertise
Favicon of DeepSeek-V3

DeepSeek-V3

Free Listing

DeepSeek-V3 is a Mixture-of-Experts language model with 671B total parameters and 37B activated per token. It supports a 128K context and is available in Base and Chat variants with open weights for local deployment and API use.

Visit deepseek-ai/DeepSeek-V3

Overview

DeepSeek-V3 is a Mixture-of-Experts (MoE) language model with 671B total parameters and 37B activated per token. It uses Multi-head Latent Attention and DeepSeekMoE, an auxiliary-loss-free load balancing strategy, and a Multi-Token Prediction objective. It was pre-trained on 14.8 trillion tokens and then fine-tuned and reinforced, with evaluations reporting strong performance on math and code tasks and a 128K context window.

Key Features

  • Mixture-of-Experts architecture with 671B total parameters and 37B activated per token.
  • Multi-head Latent Attention (MLA) and DeepSeekMoE architectures.
  • Auxiliary-loss-free load balancing strategy for minimizing performance degradation.
  • Multi-Token Prediction (MTP) training objective that can also support speculative decoding.
  • FP8 mixed precision training framework validated on an extremely large-scale model.
  • Pre-trained on 14.8 trillion tokens with 128K context length and Base and Chat variants.

Use Cases

  • Chatting with DeepSeek-V3 through the official website or an OpenAI-compatible API.
  • Running local inference and deployment with DeepSeek-Infer Demo, SGLang, LMDeploy, TensorRT-LLM, vLLM, or LightLLM.
  • Handling code and math tasks where DeepSeek-V3 reports strong benchmark performance.
  • Researching MoE architectures, FP8 training, Multi-Token Prediction, and reasoning distillation from DeepSeek-R1.

Getting Started

  • Clone the DeepSeek-V3 GitHub repository and navigate to the inference folder.
  • Install dependencies from requirements.txt, preferably in a fresh conda or uv virtual environment.
  • Download model weights from Hugging Face and place them in a local folder.
  • Convert Hugging Face model weights using convert.py with the desired expert and model-parallel settings.
  • Run interactive or batch inference with torchrun and generate.py using the provided config.

Deployment & Requirements

  • DeepSeek-Infer Demo requires Linux with Python 3.10 only; Mac and Windows are not supported.
  • DeepSeek-Infer Demo dependencies include torch 2.4.1, triton 3.0.0, transformers 4.46.3, and safetensors 0.4.5.
  • FP8 weights are provided; BF16 weights can be generated with the provided fp8_cast_bf16.py conversion script.
  • Hugging Face Transformers is not directly supported yet.
  • SGLang supports NVIDIA and AMD GPUs and multi-node tensor parallelism in BF16 and FP8 modes.

Before You Adopt

  • License: MIT. Review its terms before using, modifying, or distributing the project.
  • GitHub records the last push on 2025-08-28, at least 12 months before this listing was prepared; verify the current maintenance status before adoption.
  • DeepSeek-Infer Demo has platform restrictions and supports only Linux with Python 3.10; Mac and Windows are not supported.
  • Hugging Face Transformers is not directly supported yet, so users need another supported inference path.
  • MTP support is under active development within the community, so that capability may require community progress or contributions.
  • Only FP8 weights are provided, and BF16 weights require a conversion script for experimentation.

Comments

Sign in to leave a comment.

More like DeepSeek-V3

Favicon of Qwen3-VL

Qwen3-VL

Free ListingStars: 20K

Multimodal large language model series by Qwen team

Language Models

Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud, offering dense and MoE architectures, advanced visual reasoning, long context, and expanded OCR.

Favicon of Qwen3

Qwen3

Free ListingStars: 27.7K

Large language model series by the Qwen team, Alibaba Cloud

Language Models

Qwen3 is the large language model series from Alibaba Cloud's Qwen team. It offers dense and Mixture-of-Experts models with thinking and non-thinking modes, multilingual, long-context, reasoning, coding, and agent capabilities.

Favicon of llama.cpp

llama.cpp

Free ListingStars: 129.5K

LLM inference in C/C++

Language Models

LLM inference in C/C++ with minimal setup and state-of-the-art performance on a wide range of hardware, locally and in the cloud.