vLLM x VastAI/Recipes
Qwen

Qwen/Qwen3-235B-A22B-Instruct-2507

Flagship Qwen3 MoE instruct model with 235B total and 22B active parameters, tuned for high-quality text generation.

Verified on 8x/16x VA16 (FP8)

moe235B / 22B131,072 ctxvLLM 0.17.0text
Guide

Overview

Qwen3-235B-A22B-Instruct-2507 is the flagship instruct MoE model in the Qwen3 series, with 235B total parameters and 22B active parameters. This guide covers deploying the model efficiently with vLLM on Vastai GPUs.

Prerequisites

  • vLLM version: 0.17.0
  • Config: GPUs = 4x VA16, Precision = FP8, TP = 16, max-model-len = 128K

Example config only. Refer to above for others.

Start Docker Container

docker run \
  --privileged=true \
  --name vllm_service \
  --shm-size=256g \
  --ipc=host \
  -p 8000:8000 \
  -it \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  harbor.vastaitech.com/ai_deliver/vllm_vacc:latest \
  bash

Launching the Server

FP8 on 4x VA16

export LLM_MAX_PREFILL_SEQ_LEN="102400"
vllm serve Qwen/Qwen3-235B-A22B-Instruct-2507-FP8 \
  --trust-remote-code \
  --tensor-parallel-size 16 \
  --max-model-len 131072 \
  --enable-auto-tool-choice --tool-call-parser hermes \
  --enforce-eager 

Client Usage

from openai import OpenAI

client = OpenAI(api_key="EMPTY", base_url="http://localhost:8000/v1")
resp = client.chat.completions.create(
    model="Qwen/Qwen3-235B-A22B-Instruct-2507-FP8",
    messages=[{"role": "user", "content": "Explain gated delta networks in one paragraph."}],
    max_tokens=512,
)
print(resp.choices[0].message.content)

References

Updated 2026-04-17
Qwen/Qwen3-235B-A22B-Instruct-2507 | vLLM Recipes