Safetensors
English
byte-level
hierarchical
tokenizer-free

Hierarchical byte baseline 3x (pretrained)

Hierarchical byte model without the multi-byte prediction head, segmenting at ~3.3 bytes per latent token. The no-MBP baseline.

From Dynamic Multi-Byte Prediction With Hierarchical Language Models. Code: skai-research/lca-multibyte.

Parameters 374M
Vocabulary 261 (256 bytes + <pad>/</s>/<unk> + <en>, <eot>)
Context 4096 bytes
model_config [2, (16,), 4, 0]
attn_type None
Precision fp32

Results

Metric Value
Byte-level BPC (validation) 0.903
Bytes per latent token 3.35
Training steps 48,342

Usage

This is not a transformers architecture — load it with the code from the paper repo:

git clone https://github.com/skai-research/lca-multibyte && cd lca-multibyte
uv sync && export PYTHONPATH=$(pwd)
from huggingface_hub import snapshot_download
from src.eval.model_loader import load_fxt_model

path = snapshot_download("skai-research/fxt-3x-base")
model, tokenizer, config = load_fxt_model(path, device="cuda")

Or generate directly:

python src/eval/generate.py --model_path $(python -c \
  "from huggingface_hub import snapshot_download; print(snapshot_download('skai-research/fxt-3x-base'))") \
  --prompt "The capital of France is" --mode cached

--show_tokenization prints the learned segment boundaries.

Training

Pretrained on FineWeb-Edu sample-100BT. Config: configs/train/modern_fxt_priors_0.3_en_lambda_3_fxt_vanilla_256_scale_bp.yaml.

Citation

@article{owodunni2026lca,
  title   = {Dynamic Multi-Byte Prediction With Hierarchical Language Models},
  author  = {Owodunni, Abraham Toluwase and Okocha, Chibuzor and Grant, Christan
             and Limisiewicz, Tomasz and Kumar, Sachin},
  year    = {2026},
  journal = {arXiv preprint arXiv:2608.15454},
  url     = {https://arxiv.org/abs/2608.15454}
}
Downloads last month
4
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for skai-research/fxt-3x-base

Finetunes
1 model

Dataset used to train skai-research/fxt-3x-base

Collection including skai-research/fxt-3x-base

Paper for skai-research/fxt-3x-base