MTPLX.COM: 2 to 3x speedup. The fastest way to run models on a Mac.

Qwen 3.8 27B Bare Speed FP16

Quickest burst chat speeds. Lower quality and slower on long coding tasks. This is the M1 and M2 build.

The FP16 precision sibling of Qwen 3.8 27B Bare Speed. M1 and M2 Macs do not run bf16 well, so this build keeps every quantized weight byte-identical to the parent and stores the remaining floating tensors (scales, biases, norms, the GDN convolution and state parameters, and the MTP head) in fp16 instead of bf16. Same layout, same tuned depth and draft settings, same context window. On an M3 or newer Mac use the parent instead.

MTPLX picks the right one for you: the app and mtplx start route M1 and M2 Macs to the FP16 builds and everything newer to the parents. Draft sampler temperature 0.6 is stamped in, the measured winner for this build.

Speeds

The numbers we publish for the parent were measured on an M5 Max: 65.2 tok/s on the coding task and 32.4 tok/s sustained over a single 52,740-token answer, official Qwen 3.8 sampling, generation running to the model's own stop. This FP16 build has the same weights and runs the same MTPLX turbo path, so the speculative math is identical; absolute tok/s on an M1 or M2 depends on that chip. We have not published M1 or M2 numbers for it yet.

Download 16.0 GB
Peak unified memory (parent, measured on M5 Max) 17.0 GB
Context window 262,144 tokens
MTP depth 3
Sampling temperature 1.0, top-p 0.95, top-k 20 (the official Qwen 3.8 contract)

MTPLX_FP16_CONVERSION_MANIFEST.json in the repo lists every tensor that was cast and every tensor that was preserved, with sha256 for each shard. Speculation in MTPLX is exact at any temperature: drafts are accepted with the probability-ratio rule plus residual resampling.

Use it

Mac app: download at mtplx.com, pick "Qwen 3.8 27B Bare Speed FP16".

Command line:

pip install mtplx
mtplx serve --model Youssofal/Qwen3.8-27B-MTPLX-Bare-Speed-FP16

The other two FP16 builds: Optimized Speed FP16, Optimized Quality FP16.

Downloads last month
3,243
Safetensors
Model size
5B params
Tensor type
F16
·
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Youssofal/Qwen3.8-27B-MTPLX-Bare-Speed-FP16

Base model

Qwen/Qwen3.8-27B
Quantized
(725)
this model
Free AI Image Generator No sign-up. Instant results. Open Now