SahilCarterr
/

Qwen-Image-Distill-Full

Diffusers

Safetensors

Model card Files Files and versions

xet

Community

SahilCarterr commited on 15 days ago

Commit

db33f70

verified ·

1 Parent(s): 92f361c

Update README.md

Browse files

Files changed (1) hide show

README.md +25 -20

README.md CHANGED Viewed

@@ -5,54 +5,59 @@ tasks:
 - text-to-image-synthesis
 #model-type:
-##如 gpt、phi、llama、chatglm、baichuan 等
 #- gpt
 #domain:
-##如 nlp、cv、audio、multi-modal
 #- nlp
 #language:
-##语言代码列表 https://help.aliyun.com/document_detail/215387.html?spm=a2c4g.11186623.0.0.9f8d7467kni6Aa
 #- cn
 #metrics:
-##如 CIDEr、Blue、ROUGE 等
 #- CIDEr
 #tags:
-##各种自定义，包括 pretrained、fine-tuned、instruction-tuned、RL-tuned 等训练方法和其他
 #- pretrained
 #tools:
-##如 vllm、fastchat、llamacpp、AdaSeq 等
 #- vllm
 base_model_relation: finetune
 base_model:
   - Qwen/Qwen-Image
 ---
-# Qwen-Image 全量蒸馏加速模型
 ![](./assets/title.jpg)
-## 模型介绍
-本模型是 [Qwen-Image](https://www.modelscope.cn/models/Qwen/Qwen-Image) 的蒸馏加速版本。原版模型需要进行 40 步推理，且需要开启 classifier-free guidance (CFG)，总计需要 80 次模型前向推理。蒸馏加速模型仅需要进行 15 步推理，且无需开启 CFG，总计需要 15 次模型前向推理，**实现约 5 倍的加速**。当然，可根据需要进一步减少推理步数，但生成效果会有一定损失。
-训练框架基于 [DiffSynth-Studio](https://github.com/modelscope/DiffSynth-Studio) 构建，训练数据是由原模型根据 [DiffusionDB](https://www.modelscope.cn/datasets/AI-ModelScope/diffusiondb) 中随机抽取的提示词生成的 1.6 万张图，训练程序在 8 * MI308X GPU 上运行了约 1 天。
-## 效果展示
-||原版模型|原版模型|加速模型|
 |-|-|-|-|
-|推理步数|40|15|15|
-|CFG scale|4|1|1|
-|前向推理次数|80|15|15|
-|样例1|![](./assets/image_1_full.jpg)|![](./assets/image_1_original.jpg)|![](./assets/image_1_ours.jpg)|
-|样例2|![](./assets/image_2_full.jpg)|![](./assets/image_2_original.jpg)|![](./assets/image_2_ours.jpg)|
-|样例3|![](./assets/image_3_full.jpg)|![](./assets/image_3_original.jpg)|![](./assets/image_3_ours.jpg)|
-## 推理代码
 ```shell
 git clone https://github.com/modelscope/DiffSynth-Studio.git
@@ -75,7 +80,7 @@ pipe = QwenImagePipeline.from_pretrained(
     ],
     tokenizer_config=ModelConfig(model_id="Qwen/Qwen-Image", origin_file_pattern="tokenizer/"),
 )
-prompt = "精致肖像，水下少女，蓝裙飘逸，发丝轻扬，光影透澈，气泡环绕，面容恬静，细节精致，梦幻唯美。"
 image = pipe(prompt, seed=0, num_inference_steps=15, cfg_scale=1)
 image.save("image.jpg")
 ```

 - text-to-image-synthesis
 #model-type:
+## e.g., gpt, phi, llama, chatglm, baichuan, etc.
 #- gpt
 #domain:
+## e.g., nlp, cv, audio, multi-modal
 #- nlp
 #language:
+## Language code list: https://help.aliyun.com/document_detail/215387.html?spm=a2c4g.11186623.0.0.9f8d7467kni6Aa
 #- cn
 #metrics:
+## e.g., CIDEr, BLEU, ROUGE, etc.
 #- CIDEr
 #tags:
+## Various custom tags, including pretrained, fine-tuned, instruction-tuned, RL-tuned, etc.
 #- pretrained
 #tools:
+## e.g., vllm, fastchat, llamacpp, AdaSeq, etc.
 #- vllm
 base_model_relation: finetune
 base_model:
   - Qwen/Qwen-Image
 ---
+# Qwen-Image Full Distillation Accelerated Model
 ![](./assets/title.jpg)
+## Model Introduction
+This model is a distilled and accelerated version of [Qwen-Image](https://www.modelscope.cn/models/Qwen/Qwen-Image).
+The original model requires 40 inference steps and uses classifier-free guidance (CFG), resulting in a total of 80 forward passes.
+The distilled accelerated model only requires 15 inference steps and does not need CFG, resulting in only 15 forward passes — **achieving about 5× speed-up**.
+Of course, the number of inference steps can be further reduced if needed, but generation quality may decrease.
+The training framework is built using [DiffSynth-Studio](https://github.com/modelscope/DiffSynth-Studio).
+The training dataset consists of 16,000 images generated by the original model using randomly sampled prompts from [DiffusionDB](https://www.modelscope.cn/datasets/AI-ModelScope/diffusiondb).
+Training was conducted for about 1 day on 8 × MI308X GPUs.
+## Performance Comparison
+| | Original Model | Original Model | Accelerated Model |
 |-|-|-|-|
+| Inference Steps | 40 | 15 | 15 |
+| CFG Scale | 4 | 1 | 1 |
+| Forward Passes | 80 | 15 | 15 |
+| Example 1 | ![](./assets/image_1_full.jpg) | ![](./assets/image_1_original.jpg) | ![](./assets/image_1_ours.jpg) |
+| Example 2 | ![](./assets/image_2_full.jpg) | ![](./assets/image_2_original.jpg) | ![](./assets/image_2_ours.jpg) |
+| Example 3 | ![](./assets/image_3_full.jpg) | ![](./assets/image_3_original.jpg) | ![](./assets/image_3_ours.jpg) |
+## Inference Code
 ```shell
 git clone https://github.com/modelscope/DiffSynth-Studio.git
     ],
     tokenizer_config=ModelConfig(model_id="Qwen/Qwen-Image", origin_file_pattern="tokenizer/"),
 )
+prompt = "Delicate portrait, underwater girl, flowing blue dress, hair floating, clear light and shadows, bubbles surrounding, serene face, exquisite details, dreamy and beautiful."
 image = pipe(prompt, seed=0, num_inference_steps=15, cfg_scale=1)
 image.save("image.jpg")
 ```