ISL Conditional Diffusion โ€” 64ร—64

A class-conditioned DDPM trained from scratch to generate 64ร—64 RGB images of Indian Sign Language hand gestures.

The model uses class conditioning on one of 35 ISL classes and supports classifier-free guidance during sampling.

The complete implementation and experiments are available in the GitHub repository.

Model

  • DDPM with UNet2D architecture
  • 64ร—64 resolution
  • RGB images
  • Class-conditional generation
  • 35 ISL classes
  • Classifier-free guidance (CFG)
  • EMA weights

Training

  • Dataset: 42,000 images (1,200 images per class ร— 35 classes)
  • Noise schedule: cosine
  • Batch size: 128
  • Learning rate: 1e-4
  • Mixed precision: fp16
  • Training steps: 65,000
  • EMA decay: 0.9999
  • CFG label dropout: 0.15
  • Data augmentation: enabled

Sampling

  • Default sampler: DDIM
  • Training diffusion timesteps: 1,000
  • Routine inference steps: 100
  • Evaluation inference steps: 50
  • Default guidance scale: 1.0
  • Random seed: 42

Results

At the FID-optimal guidance scale of 1.0:

  • FID: 57.24
  • Semantic accuracy: 99.2%

The model provides strong class control under the reported evaluation setting.

Downloads last month
21
Safetensors
Model size
29.8M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support