amd
/

Llama-3.1-405B-Instruct-MXFP4-Preview

Safetensors

llama

8-bit precision

quark

Model card Files Files and versions Community

linzhao-amd commited on Jun 27

Commit

1a53b9a

verified ·

1 Parent(s): 0f7ea55

Create README.md

Browse files

Files changed (1) hide show

README.md +154 -0

README.md ADDED Viewed

	@@ -0,0 +1,154 @@

+---
+license: llama3.1
+base_model:
+- meta-llama/Llama-3.1-405B-Instruct
+---
+# Model Overview
+- **Model Architecture:** Meta-Llama-3.1
+  - **Input:** Text
+  - **Output:** Text
+- **Supported Hardware Microarchitecture:** AMD MI300/MI350
+- **Preferred Operating System(s):** Linux
+- **Inference Engine:** [vLLM](https://docs.vllm.ai/en/latest/)
+- **Model Optimizer:** [AMD-Quark](https://quark.docs.amd.com/latest/index.html)
+- **Calibration Dataset:** [Pile](https://huggingface.co/datasets/mit-han-lab/pile-val-backup)
+The model is the quantized version of the Meta-Llama 3.1-405B-Instruct model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check [here](https://huggingface.co/meta-llama/Llama-3.1-405B-Instruct). The MXFP4 model is quantized with [AMD-Quark](https://quark.docs.amd.com/latest/index.html).
+# Model Quantization
+This model was obtained by quantizing weights and activations of [Meta-Llama-3.1-405B-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3.1-405B-Instruct) to MXFP4 and KV cache to FP8, using AutoSmoothQuant algorithm in AMD-Quark.
+**Quantization scripts:**
+```
+cd Quark/examples/torch/language_modeling/llm_ptq/
+python3 quantize_quark.py --model_dir "meta-llama/Meta-Llama-3.1-405B-Instruct" \
+                          --model_attn_implementation "sdpa" \
+                          --quant_scheme w_mxfp4_a_mxfp4 \
+                          --kv_cache_dtype fp8 \
+                          --quant_algo autosmoothquant \
+                          --min_kv_scale 1.0 \
+                          --model_export hf_format \
+                          --output_dir $output_path \
+                          --multi_gpu
+```
+# Deployment
+## Use with vLLM
+This model can be deployed efficiently using the [vLLM](https://docs.vllm.ai/en/latest/) backend, as shown in the example below.
+```python
+from vllm import LLM, SamplingParams
+from transformers import AutoTokenizer
+model_id = "amd/Llama-3.1-405B-Instruct-MXFP4-Preview"
+number_gpus = 8
+sampling_params = SamplingParams(temperature=0.6, top_p=0.9, max_tokens=256)
+tokenizer = AutoTokenizer.from_pretrained(model_id)
+messages = [
+    {"role": "system", "content": "You are a pirate chatbot who always responds in pirate speak!"},
+    {"role": "user", "content": "Who are you?"},
+]
+prompts = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
+llm = LLM(model=model_id, tensor_parallel_size=number_gpus, max_model_len=4096)
+outputs = llm.generate(prompts, sampling_params)
+generated_text = outputs[0].outputs[0].text
+print(generated_text)
+```
+## Evaluation
+The model was evaluated on MMLU and GSM8K_COT.
+Evaluation was conducted using the framework [lm-evaluation-harness](https://github.com/EleutherAI/lm-evaluation-harness) and the vLLM engine.
+### Accuracy
+#### Open LLM Leaderboard evaluation scores
+<table>
+  <tr>
+   <td><strong>Benchmark</strong>
+   </td>
+   <td><strong>Meta-Llama-3.1-405B-Instruct </strong>
+   </td>
+   <td><strong>Meta-Llama-3.1-405B-Instruct-MXFP4(this model)</strong>
+   </td>
+   <td><strong>Recovery</strong>
+   </td>
+  </tr>
+  <tr>
+   <td>MMLU (5-shot)
+   </td>
+   <td>87.63
+   </td>
+   <td>86.62
+   </td>
+   <td>98.85%
+   </td>
+  </tr>
+  <tr>
+   <td>GSM-8K-cot (8-shot, strict-match)
+   </td>
+   <td>96.51
+   </td>
+   <td>96.06
+   </td>
+   <td>99.53%
+   </td>
+  </tr>
+</table>
+### Reproduction
+The results were obtained using the following commands:
+#### MMLU
+```
+lm_eval \
+    --model vllm \
+    --model_args pretrained="amd/Llama-3.1-405B-Instruct-MXFP4-Preview",gpu_memory_utilization=0.85,tensor_parallel_size=8,kv_cache_dtype='fp8' \
+    --tasks mmlu_llama \
+    --fewshot_as_multiturn \
+    --apply_chat_template \
+    --num_fewshot 5 \
+    --batch_size auto
+```
+#### GSM8K_COT
+```
+lm_eval \
+    --model vllm \
+    --model_args pretrained="amd/Llama-3.1-405B-Instruct-MXFP4-Preview",gpu_memory_utilization=0.85,tensor_parallel_size=8,kv_cache_dtype='fp8' \
+    --tasks gsm8k_llama \
+    --fewshot_as_multiturn \
+    --apply_chat_template \
+    --num_fewshot 8 \
+    --batch_size auto
+```
+#### License
+Modifications copyright(c) 2024 Advanced Micro Devices,Inc. All rights reserved.
+Licensed under the Apache License, Version 2.0 (the "License");
+you may not use this file except in compliance with the License.
+You may obtain a copy of the License at
+    http://www.apache.org/licenses/LICENSE-2.0
+Unless required by applicable law or agreed to in writing, software
+distributed under the License is distributed on an "AS IS" BASIS,
+WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+See the License for the specific language governing permissions and
+limitations under the License.