ISL Conditional Diffusion

A class-conditioned DDPM trained from scratch to generate 128×128 RGB images of Indian Sign Language hand gestures.

The model uses class conditioning on one of 35 ISL classes and supports classifier-free guidance during sampling.

The complete implementation and experiments are available in the GitHub repository.

Model

  • DDPM with UNet2D architecture
  • 128×128 resolution
  • RGB images
  • Class-conditional generation
  • 35 ISL classes
  • Classifier-free guidance (CFG)
  • EMA weights

Training

  • Dataset: 42,000 images (1,200 images per class × 35 classes)
  • Noise schedule: cosine
  • Batch size: 64
  • Learning rate: 1e-4
  • Mixed precision: fp16
  • Training steps: 65,000
  • EMA decay: 0.9999
  • CFG label dropout: 0.15
  • Data augmentation: enabled

Sampling

  • Default sampler: DDIM
  • Training diffusion timesteps: 1,000
  • Routine inference steps: 100
  • Evaluation inference steps: 50
  • Default guidance scale: 3.0
  • Random seed: 42

Results

At the FID-optimal guidance scale of 3.0:

  • FID: 86.73
  • Semantic accuracy: 94.0%

The model provides reliable class control. Under the reported evaluation setting, its FID is comparable to the unconditional baseline.

Downloads last month
30
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support