MOSS-Transcribe-Diarize — MLX 8-bit

8-bit MLX quantization of OpenMOSS-Team/MOSS-Transcribe-Diarize (0.9B end-to-end transcription + speaker diarization + timestamps).

  • Quantization: affine 8-bit, group size 64, text backbone only — the Whisper encoder and the VQ adaptor stay full precision (quantizing the encoder's positional embedding breaks the feature broadcast).
  • Quality: byte-identical transcripts to the bf16 reference across Korean, Korean↔English code-switching, and multi-speaker English in our checks.
  • Load: mlx-audio / mlx-audio-swift via fromModelDirectory.

Original model © the OpenMOSS / MOSI.AI team, Apache-2.0. This repository only re-hosts a quantized copy for reproducible on-device deployment.

Downloads last month
67
Safetensors
Model size
0.6B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kuotient/MOSS-Transcribe-Diarize-MLX-8bit

Quantized
(11)
this model