Instructions to use MuseMesh/mume-math-125m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MuseMesh/mume-math-125m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="MuseMesh/mume-math-125m", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("MuseMesh/mume-math-125m", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use MuseMesh/mume-math-125m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "MuseMesh/mume-math-125m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuseMesh/mume-math-125m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/MuseMesh/mume-math-125m
- SGLang
How to use MuseMesh/mume-math-125m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "MuseMesh/mume-math-125m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuseMesh/mume-math-125m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "MuseMesh/mume-math-125m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuseMesh/mume-math-125m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use MuseMesh/mume-math-125m with Docker Model Runner:
docker model run hf.co/MuseMesh/mume-math-125m
Mume Math 125M
Mume Math 125M is a 134M-parameter language model trained from scratch on 1.33B tokens of mathematical web text (OpenWebMath), with the recipe, budget and tokenizer of Mume English 125M v0.2.0. It reads math far better than the English model of the same size (bits per byte on MATH test, clean subset: -62.6%), but solves almost no word problems (GSM8K 0.91%).
From Muse Mesh (Hugging Face). Part of the Muse Mesh English and Math models (collections). Every training run, including the ones not released, is logged at mume.ai/sansar/runs (moving to mume.ai/lab).
Model
| Architecture | decoder-only transformer, pre-LayerNorm (LayerNorm without bias), no biases anywhere; GPT-2-small shape |
| Layers / heads / width | 12 / 12 / 768 (head dim 64); MLP 4x = 3,072 |
| Positions | rotary (base 10,000; half-split pairing), no position table |
| Attention | causal SDPA; queries and keys RMS-normalised per head (QK-norm) |
| MLP activation | squared ReLU |
| Output head | separate (untied), zero-initialised |
| Vocabulary | 32,000 (MuseMesh/mume-tokenizer-32k v0.1.0, bundled) |
| Context | 1,024 tokens |
| Parameters, total | 134,105,856 |
| Parameters, non-embedding | 84,953,856 |
| Token embedding | 24,576,000 |
| Output head | 24,576,000 |
| Weights in this repo | bfloat16 safetensors (trained as fp32 master weights under bf16 autocast) |
Why this run: The English winner's recipe and budget, unchanged, on mathematical web text.
Evaluation
Metric: bits per UTF-8 byte (lower is better): the model's negative log-likelihood of a text divided by the text's UTF-8 byte count, so models with different tokenizers share a denominator. Each set's records are joined with </s> into one stream and scored teacher-forced in 1,024-token windows with stride 512 (every scored token after the first window has at least 512 tokens of context); </s> carries loss but no bytes. Scored with scripts/train/eval_bpb.py on a GPU in bf16 autocast. The sets are frozen and published as MuseMesh/mume-eval-suites.
The English model is MuseMesh/mume-english-125m v0.2.0 (EN1-full): the same architecture, recipe, tokenizer and token budget, trained on FineWeb instead of OpenWebMath.
| set | what | this model | English model | change |
|---|---|---|---|---|
gsm8k_test |
GSM8K test, 1,319 problems: question, blank line, worked answer (calculator annotations removed) | 1.07016 | 1.20665 | -11.3% |
owm_val |
OpenWebMath shard 113 (never trained on), the first 2,704 documents up to 20 MB | 0.99258 | 1.54392 | -35.7% |
owm_val_clean |
owm_val minus the 1,578 documents that share a copied passage with the training data (1,126 documents) | 1.06595 | 1.57023 | -32.1% |
math_test |
MATH test (Hendrycks), 5,000 problems: problem, blank line, solution | 0.84625 | 2.36162 | -64.2% |
math_test_clean |
math_test minus the 1,246 problems that share a copied passage with the training data (3,754 problems) | 0.86027 | 2.30167 | -62.6% |
There is no gsm8k_clean: the contamination check found no GSM8K test problem with a copied passage in the training data (0.13% of its 8-grams occur there, scattered), so gsm8k_test is already clean.
With 512-token windows (stride 256): owm_val 1.0198, gsm8k_test 1.0736, math_test 0.8650 (eval/bpb_ctx512.json).
Contamination, measured before training (lowercase word 8-grams of each set looked up in the whole training slice; a record has a copied passage when 5+ consecutive 8-grams hit): owm_val 1,578 of 2,704 documents (58%; the web repeats pages across shards), MATH test 1,246 of 5,000 problems (25%; solutions copied onto web pages), GSM8K test 0 of 1,319. The _clean sets drop every such record. On MATH the clean set moves this model by +1.7% (0.84625 -> 0.86027), so the MATH number is not memorisation. On OpenWebMath the clean subset reads +7.4% higher for this model and +1.7% for the English model, so part of the owm_val number does come from copies; quote the _clean columns. Against the English model on the clean sets: owm_val_clean -32.1%, math_test_clean -62.6%.
GSM8K exact match
3-shot, greedy: three worked GSM8K train examples ("Question: ...\nAnswer: ... #### N"), then the test question and "Answer:"; up to 200 new tokens, stopping at the next "Question:" or after the #### line. The prediction is the number after #### (else the last number); it counts when it equals the gold number. All 1,319 test problems (scripts/experts/math/gsm8k_exact.py).
| model | right | exact match |
|---|---|---|
| English model (EN1-full) | 21 / 1,319 | 1.59% |
| this model (MATH0) | 12 / 1,319 | 0.91% |
These runs used the fp32 training weights (GPU, bf16 autocast). This repo's bf16 weights, same protocol and GPU: 13 / 1,319 = 0.99% (greedy decoding of a model this small flips on near-ties, so single problems change).
Low single digits is what a 125M base model scores here; the number is the family's first capability metric. A fine-tuned version on the sft branch (tag v0.1.0-sft) reaches 2.43% and shows why the score stays low (see its card).
Checked before release
modeling_mume.pyagainst the training code (scripts/train/model.py), fp32 on CPU, 6 inputs (one evaluation text per set, cropped to 1,024 tokens, and 1,024 random ids): with the same bf16 weights the logits are identical (max |difference| 0); against the fp32 training weights the bf16 storage moves logits by at most 0.205.- Bits per byte with
eval_bpb.py's scoring on the first 20 records of each set (each cut to 20,000 characters): 1.06580 (training checkpoint, fp32) vs 1.06587 (this repo, bf16 weights) = +0.0063%. - The whole
gsm8k_testset through this repo's model (bf16 weights, fp32 on CPU): 1.07000, against 1.07016 in the table (training checkpoint, GPU bf16 autocast), -0.015%. - Tokenizer:
AutoTokenizer(..., trust_remote_code=True)gives the training ids on 100/100 sample texts and decodes them back exactly on 100/100; on all 58,338 texts of the tokenizer's own check the same holds (see MuseMesh/mume-tokenizer-32k). - Left-padded batches give the same logits as single sequences (max |difference| 3.1e-05).
Quick start
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "MuseMesh/mume-math-125m"
tok = AutoTokenizer.from_pretrained(repo, revision="v0.1.0", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, revision="v0.1.0", trust_remote_code=True)
inputs = tok("Let $f(x) = x^2 + 1$. Then", return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=60, do_sample=True, temperature=0.8, top_k=100)
print(tok.decode(out[0], skip_special_tokens=True))
Needs transformers (tested with 5.18.0), torch and sentencepiece. trust_remote_code=True runs two small files from this repository, modeling_mume.py (the network; the Sansar model code with the classes renamed) and tokenization_mume.py (SentencePiece, so the ids are exactly the training ids); read them before you run them. Without trust_remote_code the tokenizer falls back to tokenizer.json, which agrees on short texts but splits a few near-tie words differently in long documents. The weights are stored in bfloat16; pass dtype=torch.float32 for fp32, usually the faster choice on a CPU. The tokenizer adds no BOS/EOS: training documents were separated by </s> only.
What this checkpoint wrote for that prompt (CPU, torch.manual_seed(0), 60 new tokens):
Let $f(x) = x^2 + 1$. Then $|f(x)| = |x|^2$. If $f(x) = x^2 + 1$, then $|f'(x)| = |x|^2 = |x|^2$. Thus $|f(x)| \leq |x^2 +
and with greedy decoding:
Let $f(x) = x^2 + 1$. Then $f(x) = x^2 + 1$ and $f(x) = x^2 + 1$ for all $x \in \mathbb{R}$.
Now, let $f(x) = x^2 + 1$ and $f(
To score text instead of generating it, call the model with labels=input_ids; loss is the mean negative log-likelihood per token in nats.
Training data
OpenWebMath shards 4-15 of 114 (12 parquet files): 662,940 web pages with mathematical content, 816M words; 1,824 pages whose URL also occurs in shard 113 (the held-out shard) were skipped. Shards 0-3 were the tokenizer's math sample and are not in the training data.
Preparation (scripts/train/prep_data.py): Unicode NFC; documents longer than 20,000 characters cut into parts at whitespace (774,551 records); tokenized with the shared 32k tokenizer; one </s> after every record. A 0.5% validation stream by hash of each record's text (3,769 records); the rest is training: 770,782 records, 1,555,523,031 tokens, 5.12 GB (3.29 bytes per token).
Every training document is listed (source file, row, id; no text) in MuseMesh/mume-eval-suites, config training_manifests, split math_pretrain, with the parts that went to validation, so the slice can be rebuilt from the public source.
Training
| Optimizer, block matrices | Muon (84,934,656 params): momentum 0.95 (Nesterov), 5 Newton-Schulz steps in bf16, peak lr 0.02, update scaled by sqrt(max(1, rows/cols)), no weight decay |
| Optimizer, everything else | AdamW: embedding + head with weight decay 0.1, LayerNorm gains without; peak lr 0.001, betas (0.9, 0.95), eps 1e-8 |
| Learning-rate schedule | linear warm-up over 500 steps, then cosine decay to 0.1 x peak at step 27,106 |
| Batch | 12 sequences x 4 accumulation steps x 1,024 tokens = 49,152 tokens per step |
| Steps | 27,106 |
| Tokens seen | 1,332,314,112 (0.86 passes over 1,555.5M training tokens) |
| Gradient clipping | global norm 1.0 |
| Precision | bf16 autocast, fp32 master weights, torch.compile |
| Seed | 1337 |
| Data sampling | 1,024-token windows at uniformly random offsets of the token stream (documents separated by </s>) |
| Hardware | 1x NVIDIA RTX 3060 12 GB, power-capped at 150 W |
| Wall time | 889 min (14.8 h), 25.1k tokens/s; one power cut, resumed from a checkpoint |
| Compute cost | own hardware, no cloud cost |
| Final validation | 1.0307 bits per byte on the run's own validation stream (0.5% of documents, non-overlapping windows) |
Limitations
- Base model. It continues text. It has not been instruction-tuned or aligned: it does not follow instructions, answer questions or chat, and it will not refuse anything.
- It makes things up. Fluent text with invented facts, names, numbers, citations and proofs. Do not use it as a source of facts or of correct mathematics.
- Small and short. 134M parameters and a 1,024-token context. This implementation has no key/value cache, so long generations are slow (each new token re-reads the context).
- Web math. OpenWebMath is forum posts, lecture notes, blogs and Q&A pages with LaTeX; the model writes in that register, can switch into English prose, and has little arithmetic ability.
- Contamination. Some evaluation text occurs in the training data (measured above); prefer the
_cleansets. - Tokenizer. 32k pieces shared with Sanskrit (in SLP1 transliteration) and math; FineWeb text takes 13% more tokens than with GPT-2's tokenizer (1.52 vs 1.35 per word).
Versions
Each version is a git tag on this repository; main is version v0.1.0. Pin one with revision=. A version is one training run, named by its run id in the training log.
| version | where | date | training run | key numbers |
|---|---|---|---|---|
| v0.1.0 | tag v0.1.0, main |
2026-09-29 | math0_shared32k_d125m_ctx1024 (MATH0) |
owm_val_clean 1.0659, GSM8K 0.91% |
| v0.1.0-sft | branch sft, tag v0.1.0-sft |
2026-10-01 | math0_sft (MATH0-SFT) |
GSM8K 2.43% |
The fine-tuned model lives on its own branch, sft, so main stays the base model that the bits-per-byte numbers describe; the tag v0.1.0-sft pins it. A later fine-tune of this base will move the sft branch and get its own tag.
Files
| file | what |
|---|---|
model.safetensors |
the weights, bfloat16 (no optimizer state) |
config.json, generation_config.json |
architecture and default generation settings |
configuration_mume.py, modeling_mume.py |
the network for transformers (auto_map, trust_remote_code) |
tokenizer.model, tokenizer.json, tokenization_mume.py, tokenizer_config.json, special_tokens_map.json |
the tokenizer (MuseMesh/mume-tokenizer-32k v0.1.0) |
eval/bpb_clean_ctx1024.json, eval/bpb_ctx1024.json, eval/bpb_ctx512.json, eval/contamination_overlap.json, eval/english_model_bpb_clean_ctx1024.json, eval/english_model_bpb_ctx1024.json, eval/english_model_gsm8k_exact_3shot.json, eval/gsm8k_exact_3shot.json, eval/gsm8k_exact_3shot_released_weights.json, eval/heldout_meta.json, eval/smoke_test.json, eval/verification.json |
the evaluation outputs quoted above, the release checks and the sample generations |
training/data_meta.json, training/slice_summary.json, training/summary.json |
every training argument, the validation points and the data preparation record |
LICENSE, LICENSE-CODE, CHANGELOG.md |
licences and version history |
Attribution
| data | Hugging Face | reference | licence |
|---|---|---|---|
| OpenWebMath | open-web-math/open-web-math | Paster et al. 2023, OpenWebMath, arXiv:2310.06786 | ODC-By 1.0; use is also subject to the Common Crawl Terms of Use |
| GSM8K | openai/gsm8k | Cobbe et al. 2021, Training Verifiers to Solve Math Word Problems, arXiv:2110.14168 | MIT |
| MATH | EleutherAI/hendrycks_math | Hendrycks et al. 2021, Measuring Mathematical Problem Solving With the MATH Dataset, arXiv:2103.03874 | MIT |
Licence
This release is for research and non-commercial use.
- Weights (
model.safetensors),tokenizer.modelandtokenizer.json: CC BY-NC 4.0 (LICENSE). Non-commercial use (research, teaching, non-profit work) with attribution to "Muse Mesh Private Limited". - Code (
configuration_mume.py,modeling_mume.py,tokenization_mume.py): Apache-2.0 (LICENSE-CODE). - Commercial use: contact kushal@muse-mesh.com.
Why non-commercial: The weights saw only permissively licensed text: OpenWebMath (ODC-By 1.0). The tokenizer is the issue: the shared 32k tokenizer was trained on a 1.2 GB sample of which 400 MB is Sanskrit from the Sansar corpus, and part of that Sanskrit is licensed for non-commercial use only (GRETIL, CC BY-NC-SA 4.0; Muktabodha, CC BY-NC 4.0) or carries no licence statement. The tokenizer, and the weights that only work with it, are therefore released under CC BY-NC 4.0, the same terms as the Sansar Sanskrit tokenizer. An Apache-2.0 line is planned as separate repositories: a new tokenizer trained without those sources, and the models retrained on it.
Citation
@misc{mume_math_125m_2026,
title = {Mume Math 125M},
author = {Muse Mesh},
year = {2026},
note = {v0.1.0},
url = {https://huggingface.co/MuseMesh/mume-math-125m}
}
Contact: kushal@muse-mesh.com
- Downloads last month
- -
Datasets used to train MuseMesh/mume-math-125m
MuseMesh/mume-eval-suites
Collection including MuseMesh/mume-math-125m
Papers for MuseMesh/mume-math-125m
OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text
Training Verifiers to Solve Math Word Problems
Measuring Mathematical Problem Solving With the MATH Dataset
Evaluation results
- bits per UTF-8 byte (1,024-token windows, stride 512) on owm_valtest set self-reported0.993
- bits per UTF-8 byte (1,024-token windows, stride 512) on gsm8k_testtest set self-reported1.070
- bits per UTF-8 byte (1,024-token windows, stride 512) on math_testtest set self-reported0.846
- bits per UTF-8 byte (1,024-token windows, stride 512) on owm_val_cleantest set self-reported1.066
- bits per UTF-8 byte (1,024-token windows, stride 512) on math_test_cleantest set self-reported0.860
- exact match, 3-shot, greedy on GSM8Ktest set self-reported0.910