ISL Unconditional Diffusion

An unconditional DDPM trained from scratch to generate 128×128 RGB images of Indian Sign Language hand gestures.

The model learns the overall distribution of ISL gesture images without receiving a class label during training or inference.

The complete implementation and experiments are available in the GitHub repository.

Model

  • DDPM with UNet2D architecture
  • 128×128 resolution
  • RGB images
  • Unconditional generation
  • EMA weights

Training

  • Dataset: 42,000 images (1,200 images per class × 35 classes)
  • Noise schedule: cosine
  • Batch size: 64
  • Learning rate: 1e-4
  • Mixed precision: fp16
  • Training steps: 65,000
  • EMA decay: 0.9999
  • Data augmentation: enabled

Sampling

  • Default sampler: DDIM
  • Training diffusion timesteps: 1,000
  • Routine inference steps: 100
  • Evaluation inference steps: 50
  • Random seed: 42

Results

The unconditional model achieves an FID of 86.06 under the reported evaluation setting.

Downloads last month
16
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support