Details

Dune (doo-neh ุฏูˆู†ู‡ as I like to call it, or just the english dune) is a fast and compact 12.5hz speech tokenizer trained on tens of thousands of hours of multilingual data.

the encoder is based on nvidia's nano codec architecture that compresses your audio to 22khz FSQ tokens; the decoder, using a different design then reconstructs your input to high quality 44.1khz.

Batched Extraction / Inference

fill in the path to your data in dune_extraction.py, then run it.

~$ python dune_extraction.py

Important note

this is strictly a speech tokenizer, trained only on human speech; it won't do well with music.

Downloads last month
52
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using Respair/dune_codec 1