ISL Conditional Diffusion — 64×64

A class-conditioned DDPM trained from scratch to generate 64×64 RGB images of Indian Sign Language hand gestures.

The model uses class conditioning on one of 35 ISL classes and supports classifier-free guidance during sampling.

The complete implementation and experiments are available in the GitHub repository.

Model

  • DDPM with UNet2D architecture
  • 64×64 resolution
  • RGB images
  • Class-conditional generation
  • 35 ISL classes
  • Classifier-free guidance (CFG)
  • EMA weights

Training

  • Dataset: 42,000 images (1,200 images per class × 35 classes)
  • Noise schedule: cosine
  • Batch size: 128
  • Learning rate: 1e-4
  • Mixed precision: fp16
  • Training steps: 65,000
  • EMA decay: 0.9999
  • CFG label dropout: 0.15
  • Data augmentation: enabled

Sampling

  • Default sampler: DDIM
  • Training diffusion timesteps: 1,000
  • Routine inference steps: 100
  • Evaluation inference steps: 50
  • Default guidance scale: 1.0
  • Random seed: 42

Results

At the FID-optimal guidance scale of 1.0:

  • FID: 57.24
  • Semantic accuracy: 99.2%

The model provides strong class control under the reported evaluation setting.

Downloads last month
22
Safetensors
Model size
29.8M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Free AI Image Generator No sign-up. Instant results. Open Now