text2asmr triggers -- v1 (frozen)

LoRA finetune of stabilityai/stable-audio-open-1.0 on ASMR trigger sounds (tapping, brushing, crinkling, fabric rustling, breathing, etc.), trained to step 2600.

Status: frozen as v1. Samples at this checkpoint do not reliably produce clean trigger sounds -- generations exhibit artifacts bleeding in from other regions of the training distribution (low-frequency hum, snoring-like FX, mouth clicks/whispered chuckles, broadband hiss) instead of isolated tapping/brushing. Training on this run has been stopped rather than pushed further; epoch=0-step=2600.ckpt is kept as the final v1 checkpoint and the starting point for diagnosing/re-approaching in v2 (likely dataset labeling or trigger-class separation, not just more steps).

Files:

  • epoch=0-step=2600.ckpt -- final v1 checkpoint
  • v1_sample_tapping.wav, v1_sample_brushing.wav -- raw single-trigger samples from this checkpoint (audible artifacts described above)
  • v1_composed_demo.wav -- full mixed speech+trigger ASMR composition (via scripts/compose_asmr.py), same underlying trigger artifacts present in the non-speech segments
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support