ONNX
Arabic
wav2vec2
ctc
forced-alignment
quantized
parseh

Parseh Arabic forced-alignment network

This CTC network aligns known Arabic caption text to 16 kHz mono audio in Parseh. It is not a general-purpose speech-recognition model. Parseh runs it locally with NumPy and ONNX Runtime, then uses CTC Viterbi alignment for word spans or character spans.

Provenance and conversion

Derived from jonatasgrosman/wav2vec2-large-xlsr-53-arabic at af46c2d8531b8dcbb5e23b952f739b372c2e5d2d, based on facebook/wav2vec2-large-xlsr-53. It was converted to ONNX opset 17 and dynamically quantized to int8 (QUInt8, per-channel MatMul weights); no retraining was performed. Exact source and output hashes are in meta.json and SHA256SUMS.

Verification

  • ONNX Runtime CPU versions: 1.23.2, 1.30.0; dynamic 0.5 s and 60 s inputs passed.
  • fp32 ONNX/PyTorch log-softmax max absolute difference: 0.00032806396484375; frame-argmax agreement: 1.0.
  • int8/fp32 frame-argmax agreement: 1.0; generated-speech span-start delta median/p95/max: 0.0 / 0.0 / 0.0 ms.
  • Generated 30-second CPU alignment cost: 4.811386 s; peak RSS: 2886012928 bytes.

espeak-ng generated audio was used only for pipeline monotonicity, not a real-speech accuracy claim. Unknown characters use a wildcard score to preserve a monotonic path.

Limitations

Input must be 16 kHz mono audio and known caption text. Noise, accents, text errors, vocabulary gaps, and quantization can affect timestamps. No accuracy guarantee is made.

Licence and attribution

The source tag is apache-2.0. LICENSE and NOTICE retain source, base-model, dataset attribution, and changes. No endorsement by original authors, Meta, Mozilla, or Parseh is implied.

Citation from the source card

@misc{grosman2021xlsr53-large-arabic,
  title={Fine-tuned {XLSR}-53 large model for speech recognition in {A}rabic},
  author={Grosman, Jonatas},
  howpublished={\url{https://huggingface.co/jonatasgrosman/wav2vec2-large-xlsr-53-arabic}},
  year={2021}
}

Contact

Contact: development@parseh.io.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for parseh/aligner-ar

Quantized
(5)
this model

Datasets used to train parseh/aligner-ar