MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
Paper • 2602.10934 • Published • 50
Converted to the Vokra GGUF format for Vokra, a zero-dependency speech-AI inference runtime.
This is a conversion, not a new model. The weights are the upstream ones; Vokra re-packages them so its runtime can memory-map them directly. Credit for the model belongs upstream — see Source below.
| File | Size | SHA-256 |
|---|---|---|
moss-audio-tokenizer-full.gguf |
6769.6 MB | 11cfd3853ab7949188eee1621c1beb650c9863b7a28c1adfc56c5469ead74f0e |
# Download (any HTTP client works — the file is a plain GGUF)
curl -L -o moss-audio-tokenizer-full.gguf \
https://huggingface.co/vokra/moss-audio-tokenizer-full/resolve/main/moss-audio-tokenizer-full.gguf
vokra-cli run --model moss-audio-tokenizer-full.gguf --input input.wav
| Field | Value |
|---|---|
| Architecture | moss_audio_tokenizer |
| Tensors | 1600 |
| Upstream source | OpenMOSS-Team/MOSS-Audio-Tokenizer (MOSS-Audio-Tokenizer codec Full ~1.77B params F32, arXiv:2602.10934, apache-2.0) |
| Upstream licence | apache-2.0 |
| Licence class | permissive |
| Registry model id | moss-audio-tokenizer |
| Vokra GGUF schema | 1 |
| Converted by | vokra-core 0.1.0-alpha.0 |
Every row above is read out of this file's own vokra.* metadata, so the card cannot claim something the artifact does not carry.
The weights are distributed under apache-2.0, unchanged from upstream. Conversion does not alter the licence, and your obligations run to the upstream author.
shasum -a 256 moss-audio-tokenizer-full.gguf
# expect: 11cfd3853ab7949188eee1621c1beb650c9863b7a28c1adfc56c5469ead74f0e
We're not able to determine the quantization variants.