BananaMind 2.1 MiniCoder

BananaMind 2.1 MiniCoder is an under-25M-parameter code language model.

  • 24,949,999 total parameters
  • 19,950,029 Transformer parameters
  • 4,999,970 parameters in causal bigram/four-gram memory
  • BananaMind 2 Nano's 8,192-token tokenizer
  • 13 physical layers and 18 effective passes
  • execution: L1 β†’ L2 β†’ L3 β†’ L4 β†’ L5 β†’ L6 β†’ L7 β†’ L8 β†’ L9 β†’ L5 β†’ L6 β†’ L7 β†’ L8 β†’ L9 β†’ L10 β†’ L11 β†’ L12 β†’ L13
  • the complete L5-L9 middle stack shares weights across its two passes
  • 4,096-token context

The n-gram module has independent 51,699-entry bigram and four-gram hash tables, each with 48-dimensional values. Their concatenated representation is projected to the 384-wide residual stream. It is injected through separate learned gates at the beginning of both middle-stack passes.

Training data

The exact 30B-token streamed mixture is:

Source Tokens Share
The Stack v3 train 22.5B 75%
FineWeb-Edu 7.5B 25%

Stack v3 is streamed as repository-ordered source files. Vendored files are skipped, while repository path, file path, and detected language are included in the training text. FineWeb-Edu supplies prose, naming, comments, and general language knowledge.

Checkpoints are uploaded every 5% with safetensors, tokenizer files, metrics, pinned dataset revisions, exact source-token accounting, and optimizer state.

Launch

./launch_training_hf_job.sh 4 fresh
./launch_training_hf_job.sh 4 resume
./launch_training_hf_job.sh 8 resume
Downloads last month
-
Safetensors
Model size
28.4M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support