rmala-gla-100m-3b

Final 100,465,280-parameter Turkish base language-model research checkpoint, trained from scratch on exactly 3,000,000,000 target tokens. Variant: gla. This is a base model, not an instruction-tuned or chat model.

Mirrors: Ethosoft/rmala-gla-100m-3b and MercanAI/rmala-gla-100m-3b.

Results and scope

All five variants used the same tokenizer, target tokens in the same order, and held-out validation/test sets. One training seed (41001) was used. Lower PPL/BPB is better.

Variant Test PPL Test BPB Total training algorithmic FLOP estimate
gla 25.561380 1.321054 1.866978e+18
v14_full 25.591686 1.321246 1.886712e+18
v14_half 25.482146 1.319796 1.886712e+18
full 22.187501 1.262025 2.173686e+18
hola 20.419358 1.228817 1.963008e+18

PPL includes terminal EOS; BPB uses content NLL divided by original UTF-8 byte count and excludes terminal EOS. Complete-document evaluation uses 2048-token chunks. Test: 11,352,596 tokens / 39,843,759 original UTF-8 bytes. Full validation/test metrics are in evaluation.json. Training and inference FLOP estimates are not full hardware FLOP measurements: excluded operations and profiler coverage are documented in LM100_PROTOCOL.md.

The V14-half gain over GLA is small and does not establish robust superiority. V14-full did not improve test PPL. No general harmlessness or production-readiness claim is made. Long-range retrieval/reasoning diagnostics are not included as completed results in this release. HoLA uses the official pinned GatedDeltaNet/cache implementation; it is a different backbone from normalized GLA, so this is not a cache-only ablation.

Use (native PyTorch, Linux NVIDIA CUDA)

This architecture is not registered with Transformers AutoModel. Use the bundled loader. FP32 checkpoint values are preserved exactly; the tested compute mode uses BF16 autocast and disables TF32. The bundled tokenizer native binary targets Linux x86_64 / CPython 3.11+.

hf download MercanAI/rmala-gla-100m-3b --local-dir rmala-gla-100m-3b
cd rmala-gla-100m-3b
pip install -r requirements.txt
python inference.py --prompt "Bilim ve teknoloji" --max-new-tokens 64

The reference decoder recomputes the full prefix and stops at the 2048-token context limit; it is not an optimized KV-cache decoder. First execution compiles Triton kernels. The original training model sources and pinned HOLA/FLA sources are bundled.

Training and limitations

16 layers, width 640, tied 32K input/output embeddings, RMSNorm and SwiGLU. GLA/V14: 10 heads of size 64. Full attention: causal SDPA and RoPE. HoLA: 5 heads of size 128, official betae cache, window 64 and chunk size 256. Context: 2048. Dataset: a fixed 3B-token subset of the MercanSet V11 / MercanPretraining pretokenized collection, with shard-disjoint validation/test. Cross-collection text deduplication was not performed; absolute contamination-freedom is not claimed. Dataset files are not redistributed. See training_config.json and LM100_PROTOCOL.md for optimization and V14-LM adaptation details.

V14-LM is an adaptation of the earlier synthetic V14 gate: contextual keys, int8 values, per-head 2048-byte bank budget and 5% read/write admission limits. A learned straight-through gate uses a hard 0.99 forward threshold, with accepted memory alpha=1 or 0.5. This does not imply a 95% reduction in whole-model FLOPs or that accepted reads are correct.

Only final model tensors are distributed; optimizer states, credentials and training text are not included. The tied head is stored once and restored by inference.py. Weights/code licensing is not newly assigned by this upload; see THIRD_PARTY_NOTICES.md.

Downloads last month
13
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support