Klang Pianissimo MLX bf16

A version of Klang Pianissimo, our Swedish speech recognition model, for Macs with Apple silicon (M1 or later). It runs on the Mac's GPU using MLX and parakeet-mlx, without PyTorch or NeMo.

This repository has the 16-bit (bfloat16) version: the same accuracy as the full-size model at half the size.

Versions

Repository Weights Download CV test FLEURS test Klang Dialects Speed Memory, 5 min
pianissimo-sv Original fp32 2.51 GB 4.46 6.51 4.85
pianissimo-sv-mlx-8bit 8-bit 759 MB 4.51 6.45 4.83 151× 2.2 GB
pianissimo-sv-mlx-4bit 4-bit 528 MB 4.66 6.53 5.20 126× 2.2 GB
pianissimo-sv-mlx bf16 1.25 GB 4.45 6.43 4.85 140× 2.7 GB
pianissimo-sv-mlx-fp32 fp32 2.51 GB 4.46 6.49 4.90 173× 3.8 GB

Word error rate (WER, %); lower is better. Speed is Real Time Factor (RTFx) measured on a 30-second clip. Benchmarked on Apple M5 Pro GPU (24 GB). Memory is the most GPU memory used while transcribing a 5-minute recording.

For other computers, use the ONNX versions.

Usage

pip install parakeet-mlx==0.5.2 huggingface_hub
brew install ffmpeg   # to read audio files
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("KlangAI/pianissimo-sv-mlx")
sys.path.append(path)
import pianissimo_mlx

model = pianissimo_mlx.load(path)
result = pianissimo_mlx.transcribe(model, "audio.wav")
print(result.text)

Audio longer than 2 minutes is transcribed in 2-minute chunks with 15 s overlap. Change this with chunk_duration (seconds, or None for one pass). transcribe also accepts a 16 kHz mono numpy array. Timestamps are in result.sentences and sentence.tokens.

Use pianissimo_mlx.py from this repository rather than parakeet-mlx's own transcribe. It can load the 8-bit and 4-bit versions, which parakeet-mlx cannot do on its own, and it prepares the audio exactly as it was prepared during training. parakeet-mlx prepares it slightly differently, which adds about 0.2 percentage points of WER.

How they were made

All versions are converted from the released checkpoint and keep its local attention (256 frames on each side). In the 8-bit and 4-bit versions, the encoder's linear layers are quantized with MLX's built-in quantization, in groups of 64 weights for 8-bit and 32 for 4-bit. Everything else is stored in bfloat16.

The test sets and scoring are the same as on the main model card: 5,516 Common Voice v26 test clips, 758 FLEURS test clips, and the 1,804-recording clean set of Klang Dialects. WER is calculated over each whole set after lowercasing, replacing punctuation with spaces, and collapsing whitespace. The WER was measured with this code on a Linux computer; on the M5 Pro, the same versions gave identical transcripts for 198 to 200 of 200 FLEURS clips.

License and citation

Released under CC BY 4.0, like the original model.

@misc{klang2026pianissimo,
  title = {Klang Pianissimo},
  author = {{Klang}},
  year = {2026},
  howpublished = {Hugging Face model repository},
  url = {https://huggingface.co/KlangAI/pianissimo-sv}
}
Downloads last month
13
Safetensors
Model size
0.6B params
Tensor type
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for KlangAI/pianissimo-sv-mlx

Finetuned
(6)
this model