rmala-gla-100m-3b
Final 100,465,280-parameter Turkish base language-model research checkpoint, trained from scratch on exactly 3,000,000,000 target tokens. Variant: gla. This is a base model, not an instruction-tuned or chat model.
Mirrors: Ethosoft/rmala-gla-100m-3b and MercanAI/rmala-gla-100m-3b.
Results and scope
All five variants used the same tokenizer, target tokens in the same order, and held-out validation/test sets. One training seed (41001) was used. Lower PPL/BPB is better.
| Variant | Test PPL | Test BPB | Total training algorithmic FLOP estimate |
|---|---|---|---|
| gla | 25.561380 | 1.321054 | 1.866978e+18 |
| v14_full | 25.591686 | 1.321246 | 1.886712e+18 |
| v14_half | 25.482146 | 1.319796 | 1.886712e+18 |
| full | 22.187501 | 1.262025 | 2.173686e+18 |
| hola | 20.419358 | 1.228817 | 1.963008e+18 |
PPL includes terminal EOS; BPB uses content NLL divided by original UTF-8 byte count and excludes terminal EOS. Complete-document evaluation uses 2048-token chunks. Test: 11,352,596 tokens / 39,843,759 original UTF-8 bytes. Full validation/test metrics are in evaluation.json. Training and inference FLOP estimates are not full hardware FLOP measurements: excluded operations and profiler coverage are documented in LM100_PROTOCOL.md.
The V14-half gain over GLA is small and does not establish robust superiority. V14-full did not improve test PPL. No general harmlessness or production-readiness claim is made. Long-range retrieval/reasoning diagnostics are not included as completed results in this release. HoLA uses the official pinned GatedDeltaNet/cache implementation; it is a different backbone from normalized GLA, so this is not a cache-only ablation.
Use (native PyTorch, Linux NVIDIA CUDA)
This architecture is not registered with Transformers AutoModel. Use the bundled loader. FP32 checkpoint values are preserved exactly; the tested compute mode uses BF16 autocast and disables TF32. The bundled tokenizer native binary targets Linux x86_64 / CPython 3.11+.
hf download MercanAI/rmala-gla-100m-3b --local-dir rmala-gla-100m-3b
cd rmala-gla-100m-3b
pip install -r requirements.txt
python inference.py --prompt "Bilim ve teknoloji" --max-new-tokens 64
The reference decoder recomputes the full prefix and stops at the 2048-token context limit; it is not an optimized KV-cache decoder. First execution compiles Triton kernels. The original training model sources and pinned HOLA/FLA sources are bundled.
Training and limitations
16 layers, width 640, tied 32K input/output embeddings, RMSNorm and SwiGLU. GLA/V14: 10 heads of size 64. Full attention: causal SDPA and RoPE. HoLA: 5 heads of size 128, official betae cache, window 64 and chunk size 256. Context: 2048. Dataset: a fixed 3B-token subset of the MercanSet V11 / MercanPretraining pretokenized collection, with shard-disjoint validation/test. Cross-collection text deduplication was not performed; absolute contamination-freedom is not claimed. Dataset files are not redistributed. See training_config.json and LM100_PROTOCOL.md for optimization and V14-LM adaptation details.
V14-LM is an adaptation of the earlier synthetic V14 gate: contextual keys, int8 values, per-head 2048-byte bank budget and 5% read/write admission limits. A learned straight-through gate uses a hard 0.99 forward threshold, with accepted memory alpha=1 or 0.5. This does not imply a 95% reduction in whole-model FLOPs or that accepted reads are correct.
Only final model tensors are distributed; optimizer states, credentials and training text are not included. The tied head is stored once and restored by inference.py. Weights/code licensing is not newly assigned by this upload; see THIRD_PARTY_NOTICES.md.
- Downloads last month
- 13