MTPLX.COM: 2 to 3x speedup. The fastest way to run models on a Mac.

Qwen 3.8 27B Optimized Quality FP16

8-bit dynamic quant. Good coding speeds and perfect quality. This is the M1 and M2 build.

The FP16 precision sibling of Qwen 3.8 27B Optimized Quality. M1 and M2 Macs do not run bf16 well, so this build keeps every quantized weight byte-identical to the parent and stores the remaining floating tensors (scales, biases, norms, the GDN convolution and state parameters, and the MTP head) in fp16 instead of bf16. Same layout, same tuned depth and draft settings, same context window. On an M3 or newer Mac use the parent instead.

MTPLX picks the right one for you: the app and mtplx start route M1 and M2 Macs to the FP16 builds and everything newer to the parents. You want 36 GB of unified memory or more for this one.

Speeds

The numbers we publish for the parent were measured on an M5 Max: 48.3 tok/s on the coding task and 33.2 tok/s on long xhigh reasoning, official Qwen 3.8 sampling, generation running to the model's own stop. This FP16 build has the same weights and runs the same MTPLX turbo path, so the speculative math is identical; absolute tok/s on an M1 or M2 depends on that chip. We have not published M1 or M2 numbers for it yet.

Download 29.4 GB
Peak unified memory (parent, measured on M5 Max) 32.7 GB
Context window 262,144 tokens
MTP depth 3
Sampling temperature 1.0, top-p 0.95, top-k 20 (the official Qwen 3.8 contract)

MTPLX_FP16_CONVERSION_MANIFEST.json in the repo lists every tensor that was cast and every tensor that was preserved, with sha256 for each shard. Speculation in MTPLX is exact at any temperature: drafts are accepted with the probability-ratio rule plus residual resampling.

Use it

Mac app: download at mtplx.com, pick "Qwen 3.8 27B Optimized Quality FP16".

Command line:

pip install mtplx
mtplx serve --model Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality-FP16

The other two FP16 builds: Optimized Speed FP16, Bare Speed FP16.

Downloads last month
747
Safetensors
Model size
8B params
Tensor type
F16
·
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality-FP16

Base model

Qwen/Qwen3.8-27B
Quantized
(479)
this model
Free AI Image Generator No sign-up. Instant results. Open Now