pianissimo-sv: transcribe.cpp GGUF

GGUF conversion of KlangAI/pianissimo-sv (Klang Pianissimo) for use with transcribe.cpp.

Ported from upstream commit 95c85c9, pinned 2026-09-28; the .nemo this was built from is byte-identical to the upstream LFS object (ca340b82โ€ฆ). Validated against the NeMo reference at transcribe.cpp commit c1fee50 on 2026-09-28.

Klang's Swedish fine-tune of NVIDIA's parakeet-tdt-0.6b-v3: a FastConformer encoder with a TDT/RNNT transducer decoder, taking a 16 kHz mono WAV and producing a punctuated, cased transcript with optional token-level timestamps. It keeps v3's geometry and 8,192-piece SentencePiece tokenizer but runs a local attention window of 256 encoder frames on each side at every layer, so attention cost grows linearly with recording length and long recordings stay practical. Fine-tuned on roughly 50,000 hours of Swedish speech. Not a streaming model and does not translate.

Downloads

Quantization Download Size WER (FLEURS sv_se test)
Q8_0 pianissimo-sv-Q8_0.gguf 705 MB 6.52%

WER on the full FLEURS sv_se test split (759 utterances), batch size 1, timestamps none, greedy transducer decoding with no external LM. SHA-256 of the Q8_0 file: 350bd38ea3849e0d6c05e680595ef81c185678e2d28eec0524ef704f4c345a4f.

Klang's published figure on the same split is 6.51%, and running Klang's own NeMo checkpoint over the identical manifest gives 6.53% WER (95% CI 6.02โ€“7.10). This export scores 6.52% (95% CI 6.00โ€“7.10) โ€” a 0.01pp delta against the reference, so Q8_0 quantization costs nothing measurable here, and the export reproduces the model card's quality. On a 64-utterance subset, 56 of 64 hypotheses were byte-identical between the NeMo fp32 reference and this file. Median per-clip latency was 276 ms on an Apple M2 (Metal) against 814 ms for NeMo fp32 on the same machine, and all 759 clips transcribed without error.

Usage

Build transcribe.cpp from source:

git clone git@github.com:handy-computer/transcribe.cpp.git
cd transcribe.cpp
cmake -B build && cmake --build build

Run on a 16 kHz mono WAV:

build/bin/transcribe-cli \
  -m pianissimo-sv-Q8_0.gguf \
  input.wav

If your audio isn't already 16 kHz mono WAV, convert it first:

ffmpeg -i input.mp3 -ar 16000 -ac 1 output.wav

-l sv is accepted and is the only language hint the model will honour, but it is not required. Token and word timestamps are supported.

Notes on this port

  • No runtime changes needed. This file loads in stock transcribe.cpp: no fork, patch, build flag, or catalog entry is required to run it. Verified against upstream main (c1fee50) with an unpatched build โ€” transcribe-cli loads the file and transcribes Swedish correctly, with or without -l sv. The model is a drop-in file: point -m at it instead of a v3 GGUF and the rest of the pipeline is unchanged.
  • Language claim. general.languages is ["sv"] with lang_detect disabled, rather than v3's 25-language list. Klang publishes Pianissimo as a Swedish model and states that performance on other languages and code-switching has not been established, so inheriting v3's list would be a capability claim this fine-tune has not been evaluated for. A hint naming any other language is refused (run: unsupported language, exit 1); passing no hint at all works.
  • Local attention. No loader work was needed for the 256/256 window: the C++ already runs the Regular-style local path for parakeet-tdt_ctc-1.1b ([128,128]), and the window travels as data in stt.parakeet.encoder.att_context_{left,right}. This is the family's first 0.6B local-attention variant, so the real-model loader test โ€” a hard-coded variant allow-list in tests/parakeet_real_smoke.cpp โ€” was extended to accept it. That gate is test-only and not part of the loader; the only cost of skipping it is that upstream's own test suite rejects the variant string.
  • Converter profile. transcribe.cpp's parakeet converter dispatches on the output slug, which must exist in VARIANT_PROFILES in scripts/convert-parakeet.py; a pianissimo-sv entry was added to carry the declared language, license, and size label. This matters only for re-creating the GGUF from the .nemo, not for running it.

License

Inherited from the base model: CC-BY-4.0. See the upstream model card for full terms.

Downloads last month
52
GGUF
Model size
0.6B params
Architecture
parakeet
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for nub235/pianissimo-sv-gguf

Quantized
(8)
this model