wav2vec2-base-960h, ONNX, 8-bit (word alignment)
A converted copy of facebook/wav2vec2-base-960h (Apache-2.0), used by Transcut Studio to place each word of a transcript precisely in the audio.
- Exported to ONNX (opset 17). Input
input_values[1, samples]: 16 kHz mono, zero mean and unit variance. Outputlogits[1, frames, 32], one frame per 20 ms, the model's 32 symbols. - The transformer's matrix multiplications are quantized to 8 bits (per channel); the convolutional front-end is kept at 32 bits. Word times match the full model's.
- SHA-256:
79cf3a25b88b84660deaf48f178077fda2de3b5d793d0b14986a74e95ac6b4c9
Licence: Apache-2.0 (see LICENSE and NOTICE). Original model © Facebook, Inc. and its affiliates.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for Dikaa16/transcut-aligner
Base model
facebook/wav2vec2-base-960h