inferencerlabs commited on
Commit
9466fbb
·
verified ·
1 Parent(s): b6dd3d4

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +1 -1
README.md CHANGED
@@ -15,7 +15,7 @@ See GLM-5.2 in action: [demonstration videos](https://youtube.com/xcreate)
15
 
16
  This draft model contains the **Multi-Token Prediction (MTP)** layers from **[zai-org/GLM-5.2](https://huggingface.co/zai-org/GLM-5.2)** for use alongside the [GLM-5.2-MLX](https://huggingface.co/models?search=inferencerlabs/glm-5.2) model as a speculative decoder.
17
 
18
- Q4 quant typically achieves higher throughput with less RAM usage at no loss in quality.
19
 
20
  #### Tested on a M3 Ultra 512GB RAM using [Inferencer app v2.0.6](https://inferencer.com)
21
  <table style="border-collapse: collapse; border: none; text-align:left; margin-top:10px; margin-bottom:0px;">
 
15
 
16
  This draft model contains the **Multi-Token Prediction (MTP)** layers from **[zai-org/GLM-5.2](https://huggingface.co/zai-org/GLM-5.2)** for use alongside the [GLM-5.2-MLX](https://huggingface.co/models?search=inferencerlabs/glm-5.2) model as a speculative decoder.
17
 
18
+ Q4 quant typically achieves higher throughput with less RAM usage (compared to base MTP) at no loss in quality.
19
 
20
  #### Tested on a M3 Ultra 512GB RAM using [Inferencer app v2.0.6](https://inferencer.com)
21
  <table style="border-collapse: collapse; border: none; text-align:left; margin-top:10px; margin-bottom:0px;">
Free AI Image Generator No sign-up. Instant results. Open Now