vLLM x VastAI/Recipes
DeepSeek

deepseek-ai/DeepSeek-V3-0324

DeepSeek-V3-0324 is a 671B-parameter MoE model with 37B active parameters, supporting up to 128K context length

Open-weights MoE model with native FP8. Verified on 8x/16x VA16

moe671B / 37B131,072 ctxvLLM 0.17.0text
Guide

Overview

DeepSeek-V3-0324 is a 671B-parameter Mixture of Experts (MoE) model with 37B active parameters, supporting up to 128K context length. It features native FP8 support for efficient inference and is verified on 16x VA16 with tool calling capabilities enabled.

Prerequisites

  • vLLM version: 0.17.0
  • Config: GPUs = 8x VA16, Precision = FP8, TP = 32, max-model-len = 64K

Example config only. Refer to above for others.

Start Docker Container

docker run \
  --privileged=true \
  --name vllm_service \
  --shm-size=256g \
  --ipc=host \
  -p 8000:8000 \
  -it \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  harbor.vastaitech.com/ai_deliver/vllm_vacc:latest \
  bash

Launching the Server

FP8 on 8x VA16

vllm serve deepseek-ai/DeepSeek-V3-0324 \
  --trust-remote-code \
  --tensor-parallel-size 32 \
  --max-model-len 65536 \
  --enforce-eager

Client Usage

from openai import OpenAI

client = OpenAI(api_key="EMPTY", base_url="http://localhost:8000/v1")
resp = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V3-0324",
    messages=[{"role": "user", "content": "Explain gated delta networks in one paragraph."}],
    max_tokens=512,
)
print(resp.choices[0].message.content)

References

Updated 2026-06-01
deepseek-ai/DeepSeek-V3-0324 | vLLM Recipes