MTP version request

#20
by anjeysapkovski - opened

The model is great according to tests. The only issue is that MTP version of original one already runs 2X faster, thus -33% tokens or 1.5X speedup cannot match MTP acceleration. I wonder if MTP layer from base model makes TC model more verbose.

BottleCapAI org

According to our tests MTP and even DFLash adapter trained for the original Qwen still work just fine, giving 2-5x speedup over plain Qwen. If you use it for speculative decoding it should be lossless as it's still verified by the TC so it should not make it more verbose, please let us know if you have any specific issues with it

BottleCapAI org

Let me borrow a table from our FP8 model page, where we show our brevity improvements compound cleanly with MTP and quantization-induced speedups. You can check the link for details about the experiment, as well as RealWorldQA results.

image

Adamiros changed discussion status to closed

Sign up or log in to comment

Free AI Image Generator No sign-up. Instant results. Open Now