BananaMind 2.1 Lite

BananaMind 2.1 Lite is a 24,949,999-parameter causal language model built for the under-25M class.

  • 19,950,029 parameters in the tied-token Transformer core
  • 4,999,970 parameters in a two-hash causal trigram module
  • BananaMind 2 Nano's 8,192-token tokenizer
  • 13 physical layers and 15 effective layer passes
  • execution: L1 β†’ L2 β†’ L3 β†’ L4 β†’ L5 β†’ L5 β†’ L6 β†’ L7 β†’ L8 β†’ L9 β†’ L9 β†’ L10 β†’ L11 β†’ L12 β†’ L13
  • both visits to L5 share one physical layer; both visits to L9 share another
  • 4,096-token context

The trigram module uses two independent 51,699-entry tables with 48 features per hash. Their 96 concatenated features are projected to the 384-wide residual stream. The trigram representation is computed once and reinjected before every visit to L5 and L9. L5 and L9 have separate learned injection scales, both initialized to 0.5.

75B-token streamed curriculum

Source Tokens Share
FineWeb-Edu 37.50B 50%
DCLM Baseline 16.50B 22%
Cosmopedia v2 7.50B 10%
FineMath 4+ 6.75B 9%
FinePhrase 4.50B 6%
NPset-2 Python-Edu 2.25B 3%

The NPset allocation is 75% normalized code and 25% raw original_code. Curriculum weights move from web-heavy foundations to a balanced phase and then greater math, synthetic-text, phrase, and code coverage. A token-credit scheduler keeps the complete run within one global batch of every target allocation.

Training uses Muon for matrix parameters, AdamW for embeddings and control parameters, and an independent AdamW learning rate for the n-gram module. Checkpoints are uploaded every 5% with safetensors, tokenizer files, pinned dataset revisions, metrics, exact source-token accounting, and complete state for resuming.

Launch on Hugging Face Jobs

# Start a new 4 x H200 run (32 sequences/GPU, global batch 128).
./launch_lite_25m_hf_job.sh 4 fresh

# Resume it later.
./launch_lite_25m_hf_job.sh 4 resume

# The same global batch on 8 x H200 (16 sequences/GPU).
./launch_lite_25m_hf_job.sh 8 resume

The launcher expects HF_TOKEN to be available to the Hugging Face CLI and as a Hugging Face Jobs secret.

Downloads last month
-
Safetensors
Model size
28.4M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support