Text Generation
Transformers
Safetensors
English
mume_gpt
causal-lm
custom_code
math
base-model
Eval Results (legacy)

Mume Math 125M

Mume Math 125M is a 134M-parameter language model trained from scratch on 1.33B tokens of mathematical web text (OpenWebMath), with the recipe, budget and tokenizer of Mume English 125M v0.2.0. It reads math far better than the English model of the same size (bits per byte on MATH test, clean subset: -62.6%), but solves almost no word problems (GSM8K 0.91%).

From Muse Mesh (Hugging Face). Part of the Muse Mesh English and Math models (collections). Every training run, including the ones not released, is logged at mume.ai/sansar/runs (moving to mume.ai/lab).

Model

Architecture decoder-only transformer, pre-LayerNorm (LayerNorm without bias), no biases anywhere; GPT-2-small shape
Layers / heads / width 12 / 12 / 768 (head dim 64); MLP 4x = 3,072
Positions rotary (base 10,000; half-split pairing), no position table
Attention causal SDPA; queries and keys RMS-normalised per head (QK-norm)
MLP activation squared ReLU
Output head separate (untied), zero-initialised
Vocabulary 32,000 (MuseMesh/mume-tokenizer-32k v0.1.0, bundled)
Context 1,024 tokens
Parameters, total 134,105,856
Parameters, non-embedding 84,953,856
Token embedding 24,576,000
Output head 24,576,000
Weights in this repo bfloat16 safetensors (trained as fp32 master weights under bf16 autocast)

Why this run: The English winner's recipe and budget, unchanged, on mathematical web text.

Evaluation

Metric: bits per UTF-8 byte (lower is better): the model's negative log-likelihood of a text divided by the text's UTF-8 byte count, so models with different tokenizers share a denominator. Each set's records are joined with </s> into one stream and scored teacher-forced in 1,024-token windows with stride 512 (every scored token after the first window has at least 512 tokens of context); </s> carries loss but no bytes. Scored with scripts/train/eval_bpb.py on a GPU in bf16 autocast. The sets are frozen and published as MuseMesh/mume-eval-suites.

The English model is MuseMesh/mume-english-125m v0.2.0 (EN1-full): the same architecture, recipe, tokenizer and token budget, trained on FineWeb instead of OpenWebMath.

set what this model English model change
gsm8k_test GSM8K test, 1,319 problems: question, blank line, worked answer (calculator annotations removed) 1.07016 1.20665 -11.3%
owm_val OpenWebMath shard 113 (never trained on), the first 2,704 documents up to 20 MB 0.99258 1.54392 -35.7%
owm_val_clean owm_val minus the 1,578 documents that share a copied passage with the training data (1,126 documents) 1.06595 1.57023 -32.1%
math_test MATH test (Hendrycks), 5,000 problems: problem, blank line, solution 0.84625 2.36162 -64.2%
math_test_clean math_test minus the 1,246 problems that share a copied passage with the training data (3,754 problems) 0.86027 2.30167 -62.6%

There is no gsm8k_clean: the contamination check found no GSM8K test problem with a copied passage in the training data (0.13% of its 8-grams occur there, scattered), so gsm8k_test is already clean.

With 512-token windows (stride 256): owm_val 1.0198, gsm8k_test 1.0736, math_test 0.8650 (eval/bpb_ctx512.json).

Contamination, measured before training (lowercase word 8-grams of each set looked up in the whole training slice; a record has a copied passage when 5+ consecutive 8-grams hit): owm_val 1,578 of 2,704 documents (58%; the web repeats pages across shards), MATH test 1,246 of 5,000 problems (25%; solutions copied onto web pages), GSM8K test 0 of 1,319. The _clean sets drop every such record. On MATH the clean set moves this model by +1.7% (0.84625 -> 0.86027), so the MATH number is not memorisation. On OpenWebMath the clean subset reads +7.4% higher for this model and +1.7% for the English model, so part of the owm_val number does come from copies; quote the _clean columns. Against the English model on the clean sets: owm_val_clean -32.1%, math_test_clean -62.6%.

GSM8K exact match

3-shot, greedy: three worked GSM8K train examples ("Question: ...\nAnswer: ... #### N"), then the test question and "Answer:"; up to 200 new tokens, stopping at the next "Question:" or after the #### line. The prediction is the number after #### (else the last number); it counts when it equals the gold number. All 1,319 test problems (scripts/experts/math/gsm8k_exact.py).

model right exact match
English model (EN1-full) 21 / 1,319 1.59%
this model (MATH0) 12 / 1,319 0.91%

These runs used the fp32 training weights (GPU, bf16 autocast). This repo's bf16 weights, same protocol and GPU: 13 / 1,319 = 0.99% (greedy decoding of a model this small flips on near-ties, so single problems change).

Low single digits is what a 125M base model scores here; the number is the family's first capability metric. A fine-tuned version on the sft branch (tag v0.1.0-sft) reaches 2.43% and shows why the score stays low (see its card).

Checked before release

  • modeling_mume.py against the training code (scripts/train/model.py), fp32 on CPU, 6 inputs (one evaluation text per set, cropped to 1,024 tokens, and 1,024 random ids): with the same bf16 weights the logits are identical (max |difference| 0); against the fp32 training weights the bf16 storage moves logits by at most 0.205.
  • Bits per byte with eval_bpb.py's scoring on the first 20 records of each set (each cut to 20,000 characters): 1.06580 (training checkpoint, fp32) vs 1.06587 (this repo, bf16 weights) = +0.0063%.
  • The whole gsm8k_test set through this repo's model (bf16 weights, fp32 on CPU): 1.07000, against 1.07016 in the table (training checkpoint, GPU bf16 autocast), -0.015%.
  • Tokenizer: AutoTokenizer(..., trust_remote_code=True) gives the training ids on 100/100 sample texts and decodes them back exactly on 100/100; on all 58,338 texts of the tokenizer's own check the same holds (see MuseMesh/mume-tokenizer-32k).
  • Left-padded batches give the same logits as single sequences (max |difference| 3.1e-05).

Quick start

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "MuseMesh/mume-math-125m"
tok = AutoTokenizer.from_pretrained(repo, revision="v0.1.0", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, revision="v0.1.0", trust_remote_code=True)

inputs = tok("Let $f(x) = x^2 + 1$. Then", return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=60, do_sample=True, temperature=0.8, top_k=100)
print(tok.decode(out[0], skip_special_tokens=True))

Needs transformers (tested with 5.18.0), torch and sentencepiece. trust_remote_code=True runs two small files from this repository, modeling_mume.py (the network; the Sansar model code with the classes renamed) and tokenization_mume.py (SentencePiece, so the ids are exactly the training ids); read them before you run them. Without trust_remote_code the tokenizer falls back to tokenizer.json, which agrees on short texts but splits a few near-tie words differently in long documents. The weights are stored in bfloat16; pass dtype=torch.float32 for fp32, usually the faster choice on a CPU. The tokenizer adds no BOS/EOS: training documents were separated by </s> only.

What this checkpoint wrote for that prompt (CPU, torch.manual_seed(0), 60 new tokens):

Let $f(x) = x^2 + 1$. Then $|f(x)| = |x|^2$. If $f(x) = x^2 + 1$, then $|f'(x)| = |x|^2 = |x|^2$. Thus $|f(x)| \leq |x^2 + 

and with greedy decoding:

Let $f(x) = x^2 + 1$. Then $f(x) = x^2 + 1$ and $f(x) = x^2 + 1$ for all $x \in \mathbb{R}$.

Now, let $f(x) = x^2 + 1$ and $f(

To score text instead of generating it, call the model with labels=input_ids; loss is the mean negative log-likelihood per token in nats.

Training data

OpenWebMath shards 4-15 of 114 (12 parquet files): 662,940 web pages with mathematical content, 816M words; 1,824 pages whose URL also occurs in shard 113 (the held-out shard) were skipped. Shards 0-3 were the tokenizer's math sample and are not in the training data.

Preparation (scripts/train/prep_data.py): Unicode NFC; documents longer than 20,000 characters cut into parts at whitespace (774,551 records); tokenized with the shared 32k tokenizer; one </s> after every record. A 0.5% validation stream by hash of each record's text (3,769 records); the rest is training: 770,782 records, 1,555,523,031 tokens, 5.12 GB (3.29 bytes per token).

Every training document is listed (source file, row, id; no text) in MuseMesh/mume-eval-suites, config training_manifests, split math_pretrain, with the parts that went to validation, so the slice can be rebuilt from the public source.

Training

Optimizer, block matrices Muon (84,934,656 params): momentum 0.95 (Nesterov), 5 Newton-Schulz steps in bf16, peak lr 0.02, update scaled by sqrt(max(1, rows/cols)), no weight decay
Optimizer, everything else AdamW: embedding + head with weight decay 0.1, LayerNorm gains without; peak lr 0.001, betas (0.9, 0.95), eps 1e-8
Learning-rate schedule linear warm-up over 500 steps, then cosine decay to 0.1 x peak at step 27,106
Batch 12 sequences x 4 accumulation steps x 1,024 tokens = 49,152 tokens per step
Steps 27,106
Tokens seen 1,332,314,112 (0.86 passes over 1,555.5M training tokens)
Gradient clipping global norm 1.0
Precision bf16 autocast, fp32 master weights, torch.compile
Seed 1337
Data sampling 1,024-token windows at uniformly random offsets of the token stream (documents separated by </s>)
Hardware 1x NVIDIA RTX 3060 12 GB, power-capped at 150 W
Wall time 889 min (14.8 h), 25.1k tokens/s; one power cut, resumed from a checkpoint
Compute cost own hardware, no cloud cost
Final validation 1.0307 bits per byte on the run's own validation stream (0.5% of documents, non-overlapping windows)

Limitations

  • Base model. It continues text. It has not been instruction-tuned or aligned: it does not follow instructions, answer questions or chat, and it will not refuse anything.
  • It makes things up. Fluent text with invented facts, names, numbers, citations and proofs. Do not use it as a source of facts or of correct mathematics.
  • Small and short. 134M parameters and a 1,024-token context. This implementation has no key/value cache, so long generations are slow (each new token re-reads the context).
  • Web math. OpenWebMath is forum posts, lecture notes, blogs and Q&A pages with LaTeX; the model writes in that register, can switch into English prose, and has little arithmetic ability.
  • Contamination. Some evaluation text occurs in the training data (measured above); prefer the _clean sets.
  • Tokenizer. 32k pieces shared with Sanskrit (in SLP1 transliteration) and math; FineWeb text takes 13% more tokens than with GPT-2's tokenizer (1.52 vs 1.35 per word).

Versions

Each version is a git tag on this repository; main is version v0.1.0. Pin one with revision=. A version is one training run, named by its run id in the training log.

version where date training run key numbers
v0.1.0 tag v0.1.0, main 2026-09-29 math0_shared32k_d125m_ctx1024 (MATH0) owm_val_clean 1.0659, GSM8K 0.91%
v0.1.0-sft branch sft, tag v0.1.0-sft 2026-10-01 math0_sft (MATH0-SFT) GSM8K 2.43%

The fine-tuned model lives on its own branch, sft, so main stays the base model that the bits-per-byte numbers describe; the tag v0.1.0-sft pins it. A later fine-tune of this base will move the sft branch and get its own tag.

Files

file what
model.safetensors the weights, bfloat16 (no optimizer state)
config.json, generation_config.json architecture and default generation settings
configuration_mume.py, modeling_mume.py the network for transformers (auto_map, trust_remote_code)
tokenizer.model, tokenizer.json, tokenization_mume.py, tokenizer_config.json, special_tokens_map.json the tokenizer (MuseMesh/mume-tokenizer-32k v0.1.0)
eval/bpb_clean_ctx1024.json, eval/bpb_ctx1024.json, eval/bpb_ctx512.json, eval/contamination_overlap.json, eval/english_model_bpb_clean_ctx1024.json, eval/english_model_bpb_ctx1024.json, eval/english_model_gsm8k_exact_3shot.json, eval/gsm8k_exact_3shot.json, eval/gsm8k_exact_3shot_released_weights.json, eval/heldout_meta.json, eval/smoke_test.json, eval/verification.json the evaluation outputs quoted above, the release checks and the sample generations
training/data_meta.json, training/slice_summary.json, training/summary.json every training argument, the validation points and the data preparation record
LICENSE, LICENSE-CODE, CHANGELOG.md licences and version history

Attribution

data Hugging Face reference licence
OpenWebMath open-web-math/open-web-math Paster et al. 2023, OpenWebMath, arXiv:2310.06786 ODC-By 1.0; use is also subject to the Common Crawl Terms of Use
GSM8K openai/gsm8k Cobbe et al. 2021, Training Verifiers to Solve Math Word Problems, arXiv:2110.14168 MIT
MATH EleutherAI/hendrycks_math Hendrycks et al. 2021, Measuring Mathematical Problem Solving With the MATH Dataset, arXiv:2103.03874 MIT

Licence

This release is for research and non-commercial use.

  • Weights (model.safetensors), tokenizer.model and tokenizer.json: CC BY-NC 4.0 (LICENSE). Non-commercial use (research, teaching, non-profit work) with attribution to "Muse Mesh Private Limited".
  • Code (configuration_mume.py, modeling_mume.py, tokenization_mume.py): Apache-2.0 (LICENSE-CODE).
  • Commercial use: contact kushal@muse-mesh.com.

Why non-commercial: The weights saw only permissively licensed text: OpenWebMath (ODC-By 1.0). The tokenizer is the issue: the shared 32k tokenizer was trained on a 1.2 GB sample of which 400 MB is Sanskrit from the Sansar corpus, and part of that Sanskrit is licensed for non-commercial use only (GRETIL, CC BY-NC-SA 4.0; Muktabodha, CC BY-NC 4.0) or carries no licence statement. The tokenizer, and the weights that only work with it, are therefore released under CC BY-NC 4.0, the same terms as the Sansar Sanskrit tokenizer. An Apache-2.0 line is planned as separate repositories: a new tokenizer trained without those sources, and the models retrained on it.

Citation

@misc{mume_math_125m_2026,
  title  = {Mume Math 125M},
  author = {Muse Mesh},
  year   = {2026},
  note   = {v0.1.0},
  url    = {https://huggingface.co/MuseMesh/mume-math-125m}
}

Contact: kushal@muse-mesh.com

Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train MuseMesh/mume-math-125m

Collection including MuseMesh/mume-math-125m

Papers for MuseMesh/mume-math-125m

Evaluation results