Pearl Model (5M-causal-v3): Native Luganda Conversational Language Model

This is the Pearl Model (5M-causal-v3), a decoder-only GPT-2 style autoregressive language model trained from scratch on native Luganda text and conversational corpora. It contains approximately 4,300,000 parameters and is designed to generate coherent Luganda text and respond to conversational prompts, greetings, and basic English-to-Luganda queries.

This model was engineered entirely in pure JAX using the Flax NNX API as part of the technical workshop "Building Pearl-Chat: Engineering a Native Language Model from Scratch in Pure JAX" for PyCon Africa 2026.

Model overview

  • Model name: Pearl Model (5M-causal-v3)
  • Repository: kambale/pearl-chat
  • Architecture: Decoder-only Transformer (GPT-2 style)
  • Framework: Pure JAX and Flax NNX
  • Target language: Luganda ('lg')
  • Total parameters: ~4,300,000
  • Context length: 256 tokens
  • Layers: 6 transformer decoder blocks
  • Attention heads: 8
  • Embedding dimension: 256
  • Feed-forward dimension: 1024 (4x embed_dim)
  • Tokenizer: Byte-level Byte-Pair Encoding (BPE), vocabulary size 8,192
  • Weight tying: Output projection tied to token embedding table
  • Training schedule: AdamW with warmup cosine decay

Intended use

The model is designed for:

  • Generative text continuation in Luganda.
  • Conversational interaction and question answering on basic topics (geography, greetings, education, culture in Uganda).
  • Research in low-resource African language modeling with pure functional JAX.
  • Educational demonstration of explicit state management with Flax NNX.

Out-of-scope:

  • High-stakes factual retrieval or medical and legal counsel.
  • Highly specialized technical jargon absent from the training corpus.

Training details

Datasets

The model is trained on a combination of:

  1. wikimedia/wikipedia (20231101.lg split): Encyclopedic articles covering geography, history, and society in Uganda and Buganda.
  2. Sunbird/salt (text-all subset): Parallel English-Luganda sentences and translation pairs.
  3. Curated conversational dialogues: Native Luganda question-answer pairs and greetings.

Infrastructure and hyperparameters

  • Hardware: Compatible with 1x NVIDIA T4 / A100 or multi-core CPU.
  • Batch size: 16
  • Sequence length: 256 tokens
  • Optimizer: Optax AdamW (weight decay 0.01, gradient clipping 1.0)
  • Peak learning rate: 5e-4 with 100-step linear warmup and cosine decay to 1e-5
  • Loss function: Softmax cross-entropy over discrete tokens

How to use

Load the model using native Pearl-Chat library or direct Flax NNX definitions:

from huggingface_hub import snapshot_download
from flax import nnx
from pearlchat.tokenizer import LugandaTokenizer
from pearlchat.model import PearlChatModel
from pearlchat.config import ModelConfig
from pearlchat.checkpointing import CheckpointManager
from pearlchat.generate import generate_text

# Download model repository
repo_dir = snapshot_download(repo_id="kambale/pearl-chat")

# Load tokenizer and configuration
tokenizer = LugandaTokenizer.load(repo_dir)
config = ModelConfig(
    vocab_size=tokenizer.vocab_size,
    context_length=256,
    embed_dim=256,
    num_heads=8,
    num_layers=6,
    feed_forward_dim=1024,
    dropout_rate=0.0,
)

# Initialize model and restore weights
rngs = nnx.Rngs(params=0, dropout=0)
model = PearlChatModel(config, rngs)
manager = CheckpointManager(repo_dir)
model, _ = manager.restore_latest_checkpoint(model)

# Generate response to a Luganda prompt
prompt = "Oli otya?"
response = generate_text(model, tokenizer, prompt=prompt, max_new_tokens=48, temperature=0.65)
print(f"Prompt: {prompt}")
print(f"Response: {response}")

Limitations and bias

  • Corpus volume: Luganda is an under-resourced language. While the Wikipedia and SALT datasets provide foundational coverage, overall vocabulary and domain diversity remain limited compared to high-resource languages.
  • Model capacity: At 4.3M parameters, the model is built for efficient live workshop execution and demonstration. It may occasionally generate repetitive phrasing on out-of-domain inputs.

Ethical considerations

Generative models may reflect biases present in the training sources. Users should exercise critical review of generated outputs in public applications.

Citation

If you use this model or code in your research or educational work, please cite:

@misc{kambale2026pearlchat,
  author = {Wesley Kambale},
  title = {Pearl-Chat: Engineering a Native Language Model from Scratch in Pure JAX},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {https://huggingface.co/kambale/pearl-chat},
  note = {PyCon Africa 2026 Workshop}
}

License

Apache 2.0

Downloads last month
19
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train kambale/pearl-chat