BuzzASR Padding-Tail Reproduction
This repository contains reproduction material from a small independent investigation of short-context inference in BuzzASR/Whisper-style speech recognition models.
The work started as a practical attempt to reduce latency for fully local Danish voice control in Home Assistant. During that work, an unexpected dependency on the padded encoder tail was observed.
The material is shared so that the observations can be independently tested, criticised and extended. The datasets used here are small, and the repository should not be read as establishing broad conclusions about BuzzASR, Whisper, languages, speakers or speech recognition in general.
Scope
Two parallel experiment branches are included:
| Danish | English | |
|---|---|---|
| Development set | ORIGINAL20 | ORIGINAL20 |
| Holdout set | HOLDOUT15 | HOLDOUT15 |
| Speech source | Human recording | Piper TTS |
| Canonical sample rate | 16 kHz | 16 kHz |
| BuzzASR model | Danish | English |
| Experiment series | V1-V23 | V1-V23 replication |
The English branch is a cross-model / cross-language / cross-speaker-source replication. It is not a controlled language-only comparison because the Danish speech is human-recorded while the English speech is synthetic.
English audio distribution
The historical English V1-V23 replication used locally generated Piper
en_US-lessac-medium audio. Those WAV files are intentionally not
redistributed in this public repository because we want to respect
third-party licensing and redistribution rights where those rights are not
sufficiently clear to us. Their omission is not intended to restrict access
to the experiment.
The English prompts, generation provenance, processing information and historical SHA-256 hashes are retained so that comparable synthetic input can be independently regenerated and compared with the documented historical experiment. Byte-identical regeneration of the historical WAV files is not guaranteed. See REPRODUCIBILITY.md and THIRD_PARTY.md for details.
Important limitations
- The datasets are small.
- A difference such as 13/15 versus 14/15 is a one-utterance difference and should not be over-interpreted.
- Some exact-output mismatches are punctuation-only.
- "Exact control" measures representation/output equivalence in these experiments; it is not equivalent to ASR quality.
- The Danish HOLDOUT15 set was used diagnostically in later experiments and should not be treated as a pristine final validation set.
- Results from these experiments should therefore be regarded as observations and hypotheses to test on larger and independent datasets.
Reproduction
See REPRODUCIBILITY.md.
For the V11 predictor dataset format and suggestions for independent large-scale follow-up experiments, see EXTENDING.md.
Licensing
Original project code and documentation are released under the MIT License.
The 35 Danish human-recorded WAV files under danish/audio/ are released
under the Creative Commons Attribution-NonCommercial 4.0 International
license (CC BY-NC 4.0).
Third-party models, voices, datasets and generated material may be subject to separate terms. See THIRD_PARTY.md for provenance and licensing boundaries.
Repository integrity
Canonical datasets contain per-file SHA-256 hashes. A repository-wide
SHA256SUMS file is generated before release.
Background and attribution
BuzzASR and the underlying model work belong to their respective authors and projects. This repository contains independent reproduction material and does not imply affiliation with BuzzASR, Lemn Lab, UC San Diego, or their collaborators.
The experiments were run and reviewed by Lars Folmann Christiansen as part of a personal Home Assistant / local speech-recognition project in Denmark.
ChatGPT was used as a technical and research assistant for scripting, experiment planning, analysis, bookkeeping and documentation. Experimental runs and resulting artefacts were produced on the described compute systems.
Status
This repository is being prepared as an open reproduction package for independent testing and extension.
The preserved experiment series is V1-V23. Larger-scale predictor experiments, additional languages and continuous evaluation metrics are intentionally left as follow-up work rather than presented here as completed results.