YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
# UltraSparse
UltraSparse — CPU adaptive causal language model
A small, reproducible PyTorch language model designed for CPU experimentation. The model combines:
- a ~10M-parameter Transformer,
- SentencePiece BPE tokenization,
- adaptive depth at inference,
- a current-token difficulty head,
- a short-horizon future-demand head,
- a persistent computation ledger across autoregressive decoding,
- strict zero-debt compute control,
- matched-budget benchmarking and quality gates.
Important positioning
This is an experimental/educational language model, not a state-of-the-art LLM and not instruction tuned. The persistent computation ledger is a project design feature; do not describe it as novel without a formal literature review.
Build
python -m pip install -r requirements.txt
python train_tokenizer.py `
--text data\corpus.txt `
--out artifacts\tokenizer `
--vocab-size 1024
python train.py `
--text data\corpus.txt `
--tokenizer artifacts\tokenizer.model `
--steps 20000 `
--batch-size 2 `
--seq-len 256 `
--threads 8 `
--out results\ultrasparse_v4_10m.pt
First do a smoke test
python train.py --steps 20 --batch-size 1 --seq-len 64 --threads 4 `
--out results\smoke_v4.pt
python -m pytest -q
Do not start a 20k-step CPU run until the smoke test passes.
Evaluate
python evaluate.py `
--checkpoint results\ultrasparse_v4_10m.pt `
--text data\corpus.txt
python quality_gate.py `
--checkpoint results\ultrasparse_v4_10m.pt
Generate
python generate.py `
--checkpoint results\ultrasparse_v4_10m.pt `
--prompt "It was a beautiful morning" `
--tokens 150 `
--policy predictive_ledger `
--temperature 0.8 `
--top-k 40 `
--top-p 0.9
Benchmark
python benchmark.py `
--checkpoint results\ultrasparse_v4_10m.pt `
--tokens 128 `
--threads 8
Benchmark every policy at the same budget. Report quality and compute together; never claim a speed improvement from a single timing run.
Release checklist
- tokenizer artifact committed
- checkpoint loads from a clean environment
- tests pass
- quality gate passes
- train/validation/test split documented
- benchmark JSON committed
- parameter count measured
- CPU hardware and thread count recorded
- training corpus provenance recorded
- model card states limitations
- no API keys or secrets in repository
\n## Senior-engineering design decisions\n\n
- RoPE instead of a learned absolute position table: avoids the old positional-index failure mode when context is clipped and keeps positional parameters out of the model.
- Pre-normalized residual blocks: stable training and clean layer-prefix execution.
- Tied input/output embeddings: reduces parameters and keeps the small model focused on transformer capacity.
- First-block probe + continuation: the adaptive inference path does not intentionally execute the first block twice.
- Zero-debt ledger: a policy cannot borrow future compute; overspend attempts are counted and forced down to the highest legal depth.
- Separate v3 baseline: the old byte-token model remains useful as a historical baseline; v4 is trained from scratch because changing tokenization changes the embedding/output shapes.
- Reproducibility first: fixed seeds, saved tokenizer, config, optimizer/scheduler state, benchmark JSON, and quality gate.