Made Sglang + Dflash2 + Docker

#4
by AIReach - opened

Dflash2 has not yet installed in the latest Sglang official docker images. I made this repo to build and server Dflash2 embedded Sglang containers.
https://github.com/xycjscs/sglang-dflash2

On my dual 5090s machine, following speeds are achieved.

Performance (2× RTX 5090, TP=2)

Measured from production logs of this image and start.sh flags. Prefill for large requests uses whole-request wall time.

Prefill

tok/s
Large requests (≥2k tokens, n=93), mean 4391
Median 3736
p10 / p90 2257 / 7331
Token-weighted 4583
Longest request (108k tokens) 3731 (29 s)
Short turns (<2k tokens, median) ~470 (includes scheduling gaps)

The first chunk of each large request is slowed by Triton JIT (one logged chunk at 25.6 tok/s). Later chunks return to 4k–5k tok/s.

Decode (DFlash2)

Mean / median / p90 / peak 172 / 156 / 280 / 542 tok/s
Mean accept length 3.5 tokens / step (8 draft tokens)
Accept rate ~0.40 (p10 0.20 / p90 0.59)

Throughput barely drops with context length (hybrid GDN linear attention):

Context tok/s
0–20k 182
40–60k 162
120–140k 201
140–160k 151

Samples at 100–150 tok/s line up with accept rate 0.15–0.25 (draft misses, close to raw decode).

AIReach changed discussion status to closed
AIReach changed discussion status to open

Sign up or log in to comment

Free AI Image Generator No sign-up. Instant results. Open Now