File size: 5,537 Bytes
ce4a9b9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
04e1ac9
 
 
bcacf45
04e1ac9
338277b
 
 
 
 
 
 
 
 
 
 
 
 
ce4a9b9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
fea50ed
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ce4a9b9
 
 
 
 
 
 
 
 
 
 
 
 
 
04e1ac9
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
---
language:
- en
- fr
- multilingual
license: apache-2.0
tags:
- gguf
- quantized
- mac
- apple-silicon
- local-inference
- worthdoing
base_model: Qwen/Qwen2.5-Coder-7B-Instruct
quantized_by: worthdoing
pipeline_tag: text-generation
---

<p align="center">
  <img src="https://raw.githubusercontent.com/Worth-Doing/brand-assets/main/png/variants/04-horizontal.png" alt="worthdoing" width="400"/>
</p>
<p align="center"><strong>Author: Simon-Pierre Boucher</strong></p>

<p align="center">
  <img src="https://img.shields.io/badge/Format-GGUF-blue?style=for-the-badge" alt="GGUF"/>
  <img src="https://img.shields.io/badge/Params-7B-orange?style=for-the-badge" alt="Parameters"/>
  <img src="https://img.shields.io/badge/Platform-Apple_Silicon-black?style=for-the-badge&logo=apple" alt="Apple Silicon"/>
  <img src="https://img.shields.io/badge/License-Apache_2.0-green?style=for-the-badge" alt="License"/>
  <img src="https://img.shields.io/badge/Quantized_by-worthdoing-purple?style=for-the-badge" alt="worthdoing"/>
</p>
<p align="center">
  <img src="https://img.shields.io/badge/Q4__K__M-3.7_GB-brightgreen?style=flat-square" alt="Q4_K_M"/>
  <img src="https://img.shields.io/badge/Q5__K__M-4.3_GB-yellow?style=flat-square" alt="Q5_K_M"/>
  <img src="https://img.shields.io/badge/Q8__0-6.5_GB-red?style=flat-square" alt="Q8_0"/>
</p>

# Qwen2.5-Coder-7B-Instruct - GGUF Quantized by worthdoing

> Quantized for local Mac inference (Apple Silicon / Metal) by **worthdoing**

## About

This is a GGUF quantized version of [Qwen2.5-Coder-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct), optimized for running locally on Apple Silicon Macs with `llama.cpp`, `Ollama`, or `LM Studio`.

- **Original model:** [Qwen/Qwen2.5-Coder-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct)
- **Parameters:** 7B
- **Quantized by:** worthdoing
- **Pipeline:** corelm-model v1.0

## Description

Qwen's dedicated coding model. Top-tier code generation and understanding.

## Available Quantizations

| File | Quant | BPW | Size | Use Case |
|------|-------|-----|------|----------|
| `qwen2.5-coder-7b-instruct-Q4_K_M-worthdoing.gguf` | Q4_K_M | 4.58 | ~3.7 GB | **Recommended** - Best quality/size ratio |
| `qwen2.5-coder-7b-instruct-Q5_K_M-worthdoing.gguf` | Q5_K_M | 5.33 | ~4.3 GB | Higher quality, still fast |
| `qwen2.5-coder-7b-instruct-Q8_0-worthdoing.gguf` | Q8_0 | 7.96 | ~6.5 GB | Near-original quality |

## How to Use

### With Ollama
```bash
# Create a Modelfile
cat > Modelfile <<'MODELEOF'
FROM ./qwen2.5-coder-7b-instruct-Q4_K_M-worthdoing.gguf
MODELEOF

ollama create qwen2.5-coder-7b-instruct -f Modelfile
ollama run qwen2.5-coder-7b-instruct
```

### With llama.cpp
```bash
llama-cli -m qwen2.5-coder-7b-instruct-Q4_K_M-worthdoing.gguf -p "Your prompt here" -ngl 99
```

### With LM Studio
1. Download the GGUF file
2. Open LM Studio -> My Models -> Import
3. Select the GGUF file and start chatting

## Quantization Method

Our quantization pipeline (**corelm-model v1.0**) follows a rigorous multi-step process to ensure maximum quality and compatibility:

### Step 1 — Download & Validation
- Model weights are downloaded from HuggingFace Hub in **SafeTensors** format (`.safetensors`)
- Legacy formats (`.bin`, `.pt`) are excluded to ensure clean, verified weights
- Tokenizer, configuration, and all metadata are preserved

### Step 2 — Conversion to GGUF F16 Baseline
- The original model is converted to **GGUF format at FP16 precision** using `convert_hf_to_gguf.py` from [llama.cpp](https://github.com/ggml-org/llama.cpp)
- This lossless baseline preserves the full original model quality
- Architecture-specific tensors (attention, FFN, embeddings, MoE routing) are mapped to their GGUF equivalents

### Step 3 — K-Quant Quantization
- The F16 baseline is quantized using `llama-quantize` with **k-quant methods**
- K-quants use a mixed-precision approach: more important layers (attention, output) retain higher precision, while less sensitive layers (FFN) are compressed more aggressively
- Each quantization level offers a different quality/size tradeoff:

| Method | Bits per Weight | Strategy |
|--------|----------------|----------|
| **Q4_K_M** | ~4.58 bpw | Mixed 4/5-bit. Attention & output layers use Q5_K, FFN layers use Q4_K. Best balance of quality and size. |
| **Q5_K_M** | ~5.33 bpw | Mixed 5/6-bit. Attention & output layers use Q6_K, FFN layers use Q5_K. Higher quality with moderate size increase. |
| **Q8_0** | ~7.96 bpw | Uniform 8-bit. All layers quantized to 8-bit. Near-lossless quality, largest file size. |

### Step 4 — Metadata Injection
- Custom metadata is embedded directly in each GGUF file:
  - `general.quantized_by`: worthdoing
  - `general.quantization_version`: corelm-1.0
- This ensures full traceability and provenance of every quantized file

### Tools & Environment
- **llama.cpp**: Used for both conversion and quantization — the industry-standard open-source LLM inference engine
- **Target platform**: Apple Silicon Macs (M1/M2/M3/M4) with Metal GPU acceleration
- **Inference runtimes**: Compatible with `llama.cpp`, `Ollama`, `LM Studio`, `koboldcpp`, and any GGUF-compatible runtime

## Recommended Hardware

| Quant | Min RAM | Recommended |
|-------|---------|-------------|
| Q4_K_M | 4 GB | Mac with 8 GB+ RAM |
| Q5_K_M | 5 GB | Mac with 8 GB+ RAM |
| Q8_0 | 8 GB | Mac with 12 GB+ RAM |

## Tags

`coding`, `code-generation`, `code-review`

---

*Quantized with corelm-model pipeline by **worthdoing** on 2026-04-17*
Free AI Image Generator No sign-up. Instant results. Open Now