w2v-bert-2.0-lingala-main-best-2

A Lingala automatic speech recognition (ASR) model, fine-tuned from facebook/w2v-bert-2.0.

This checkpoint follows the same recipe as keystats/w2v-bert-2.0-lingala-main-best, with one change to the data split: WAXAL's validation split was folded into the training pool, and WAXAL's test split was used as the held-out evaluation set instead.

Note on evaluation comparability: because the held-out set is now WAXAL test rather than WAXAL validation, WER/CER numbers from this checkpoint are not directly comparable to main-best or main-best-3, which were evaluated on validation. Only compare this checkpoint's scores against other checkpoints also evaluated on WAXAL test.

Model description

facebook/w2v-bert-2.0 — a large-scale, multilingual self-supervised speech encoder pretrained with a BERT-style masked prediction objective — is used as the backbone, with a from-scratch character-level CTC (Connectionist Temporal Classification) head fine-tuned specifically for Lingala.

Text casing note: training targets were kept in their raw, cased form — the vocabulary was built directly from the original transcriptions, punctuation and casing included, rather than lowercased first.

Training data

Source Role
google/WaxalNLP (lin_asr config) train and validation splits both pooled into training; test split held out untouched as the fixed evaluation benchmark
KasuleTrevor/Lingala_100hrs All splits pooled into training

Deduplication by audio hash and by transcription-within-source was applied across the pooled training data, as in the other checkpoints in this family. WAXAL's test split is the only data used for evaluation, and it was never included in training.

Training procedure

  • Base model: facebook/w2v-bert-2.0
  • Architecture: Wav2Vec2BertForCTC, add_adapter=True
  • Processor: Wav2Vec2BertProcessor — SeamlessM4TFeatureExtractor for audio features + a Wav2Vec2CTCTokenizer built from scratch on the combined training + evaluation transcriptions (character-level vocabulary, raw/cased text, | as the word delimiter, [PAD] doubling as the CTC blank token)
  • Sample rate: 16 kHz mono
  • Epochs: 4 (with early stopping, patience 5, on eval WER)
  • Effective batch size: 32 (per-device batch size 4 × gradient accumulation 8)
  • Learning rate: 3e-5, cosine schedule, 10% warmup
  • Precision: fp16, gradient checkpointing enabled
  • Regularization: attention/hidden/feature-projection dropout 0.05
  • Data filtering: clips whose transcript is too long for CTC to align within the available encoder output length ("CTC-impossible" clips) are dropped from both train and eval before training
  • Seed: 42 (deterministic — same seed for Python/NumPy/PyTorch/CUDA)

Evaluation results

Evaluated on the WAXAL Lingala test split (used here as the held-out benchmark, since validation was moved into training), greedy decoding vs. greedy + KenLM (keystats/waxal-kenlm-models-best). Adding the KLM gives a consistent, meaningful WER/CER improvement over greedy decoding alone — pair the two for the best results.

How to use

Option 1 — model alone (greedy decoding)

import torch
import librosa
from transformers import Wav2Vec2BertForCTC, Wav2Vec2BertProcessor

MODEL_ID = "keystats/w2v-bert-2.0-lingala-main-best-2"
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"

processor = Wav2Vec2BertProcessor.from_pretrained(MODEL_ID)
model = Wav2Vec2BertForCTC.from_pretrained(MODEL_ID).to(DEVICE).eval()

audio_array, sr = librosa.load("path/to/audio.wav", sr=16000, mono=True)
inputs = processor(audio_array, sampling_rate=16000, return_tensors="pt")
with torch.no_grad():
    logits = model(input_features=inputs.input_features.to(DEVICE)).logits

predicted_ids = torch.argmax(logits, dim=-1)
transcription = processor.batch_decode(predicted_ids)[0]

print(transcription)  # cased, punctuated Lingala text

Option 2 — model + KLM (recommended, higher accuracy)

# pip install pyctcdecode
# pip install https://github.com/kpu/kenlm/archive/master.zip

import torch
import librosa
from huggingface_hub import hf_hub_download
from transformers import Wav2Vec2BertForCTC, Wav2Vec2BertProcessor
from pyctcdecode import build_ctcdecoder

MODEL_ID = "keystats/w2v-bert-2.0-lingala-main-best-2"
KLM_REPO_ID = "keystats/waxal-kenlm-models-best"
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"

processor = Wav2Vec2BertProcessor.from_pretrained(MODEL_ID)
model = Wav2Vec2BertForCTC.from_pretrained(MODEL_ID).to(DEVICE).eval()

klm_path = hf_hub_download(repo_id=KLM_REPO_ID, repo_type="dataset",
                            filename="lingala/lingala_5gram_correct-best.arpa")

def build_vocab_list(tokenizer, vocab_size):
    vocab_dict = tokenizer.get_vocab()
    vocab_list = [None] * vocab_size
    for tok, idx in sorted(vocab_dict.items(), key=lambda kv: kv[1]):
        if idx < vocab_size:
            vocab_list[idx] = tok
    pad_id = tokenizer.pad_token_id
    if pad_id is not None and pad_id < len(vocab_list):
        vocab_list[pad_id] = ""
    word_delim = getattr(tokenizer, "word_delimiter_token", None)
    if word_delim:
        delim_id = vocab_dict.get(word_delim)
        if delim_id is not None:
            vocab_list[delim_id] = " "
    return vocab_list

vocab_list = build_vocab_list(processor.tokenizer, model.config.vocab_size)
decoder = build_ctcdecoder(
    vocab_list,
    kenlm_model_path=klm_path,
    alpha=0.5,
    beta=0.7,
)

audio_array, sr = librosa.load("path/to/audio.wav", sr=16000, mono=True)
inputs = processor(audio_array, sampling_rate=16000, return_tensors="pt")
with torch.no_grad():
    logits = model(input_features=inputs.input_features.to(DEVICE)).logits

transcription = decoder.decode(logits.cpu().numpy()[0], beam_width=100)
print(transcription)

Intended uses & limitations

  • Intended for transcribing spoken Lingala audio into cased, punctuated text.
  • As a CTC-based model, it assumes single-speaker, forward-only audio and has no mechanism for overlapping speech from multiple speakers.
  • Because WAXAL validation was used in training here (unlike other checkpoints in this family), don't use WAXAL validation to evaluate this specific checkpoint — it's no longer held out.

Citation

@misc{keystats_wav2vec2bert_lingala_best2,
  title={w2v-bert-2.0-lingala-main-best-2: A Lingala ASR model fine-tuned from facebook/w2v-bert-2.0},
  author={keystats},
  year={2026},
  howpublished={\url{https://huggingface.co/keystats/w2v-bert-2.0-lingala-main-best-2}}
}

@misc{waxal,
  title={WAXAL: A Multilingual African Speech Dataset},
  author={Google},
  howpublished={\url{https://huggingface.co/datasets/google/WaxalNLP}}
}

@misc{kasule_lingala_100hrs,
  title={Lingala\_100hrs},
  author={KasuleTrevor},
  howpublished={\url{https://huggingface.co/datasets/KasuleTrevor/Lingala_100hrs}}
}

@inproceedings{w2vbert2,
  title={Seamless: Multilingual Expressive and Streaming Speech Translation},
  author={Seamless Communication and others},
  year={2023},
  howpublished={\url{https://huggingface.co/facebook/w2v-bert-2.0}}
}
Downloads last month
31
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for keystats/w2v-bert-2.0-lingala-main-best-2

Finetuned
(534)
this model

Datasets used to train keystats/w2v-bert-2.0-lingala-main-best-2