text2asmr triggers -- v1 (frozen)
LoRA finetune of stabilityai/stable-audio-open-1.0 on ASMR trigger sounds
(tapping, brushing, crinkling, fabric rustling, breathing, etc.), trained to
step 2600.
Status: frozen as v1. Samples at this checkpoint do not reliably produce
clean trigger sounds -- generations exhibit artifacts bleeding in from other
regions of the training distribution (low-frequency hum, snoring-like FX,
mouth clicks/whispered chuckles, broadband hiss) instead of isolated
tapping/brushing. Training on this run has been stopped rather than pushed
further; epoch=0-step=2600.ckpt is kept as the final v1 checkpoint and the
starting point for diagnosing/re-approaching in v2 (likely dataset labeling
or trigger-class separation, not just more steps).
Files:
epoch=0-step=2600.ckpt-- final v1 checkpointv1_sample_tapping.wav,v1_sample_brushing.wav-- raw single-trigger samples from this checkpoint (audible artifacts described above)v1_composed_demo.wav-- full mixed speech+trigger ASMR composition (viascripts/compose_asmr.py), same underlying trigger artifacts present in the non-speech segments