wav2vec2-base-960h, ONNX, 8-bit (word alignment)

A converted copy of facebook/wav2vec2-base-960h (Apache-2.0), used by Transcut Studio to place each word of a transcript precisely in the audio.

  • Exported to ONNX (opset 17). Input input_values [1, samples]: 16 kHz mono, zero mean and unit variance. Output logits [1, frames, 32], one frame per 20 ms, the model's 32 symbols.
  • The transformer's matrix multiplications are quantized to 8 bits (per channel); the convolutional front-end is kept at 32 bits. Word times match the full model's.
  • SHA-256: 79cf3a25b88b84660deaf48f178077fda2de3b5d793d0b14986a74e95ac6b4c9

Licence: Apache-2.0 (see LICENSE and NOTICE). Original model © Facebook, Inc. and its affiliates.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Dikaa16/transcut-aligner

Quantized
(10)
this model