Text Generation
Transformers
Safetensors
Sanskrit
sansar
sanskrit
devanagari
slp1
causal-lm
base-model
custom_code
Eval Results (legacy)

Sansar 125M

A 97.2M-parameter Sanskrit language model trained from scratch, from the Sansar project at Muse Mesh (Hugging Face): Sanskrit language models, their corpus and their tokenizer. Size class: the d125m preset, GPT-2-small shape: 12 layers, width 768 (85.0M parameters outside the embeddings; 97.2M in total with the 8k vocabulary and untied head).

  • This version: v0.1.0 = training run f5_slp1_uni8k_d125m_plus_clean_2x, finished 2026-10-02 (experiment F6-clean-2x: F6-clean data and recipe at two passes).
  • Held-out score: 0.6039 bits per Devanagari byte (pooled over four held-out sets, excluding the Bhagavad-gītā; lower is better); 0.6254 on the stricter clean_v1 sets.
  • Base model: it continues Devanagari Sanskrit text; it is not instruction-tuned.
  • Why this checkpoint: The best 125M-class checkpoint on the held-out ex-Gītā number.

Quick start

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "MuseMesh/sansar-125m"
tok = AutoTokenizer.from_pretrained(repo, revision="v0.1.0", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, revision="v0.1.0", trust_remote_code=True)

inputs = tok("संस्कृतं नाम दैवी वाक्", return_tensors="pt")      # Devanagari in
out = model.generate(**inputs, max_new_tokens=50, do_sample=True, temperature=0.8, top_k=100)
print(tok.decode(out[0], skip_special_tokens=True))              # Devanagari out

Needs transformers (tested with 5.18.0), torch and sentencepiece. trust_remote_code=True runs three small files from this repository: modeling_sansar.py (the network), tokenization_sansar.py and translit.py (Devanagari <-> SLP1); read them before you run them. The weights are stored in bfloat16 (transformers 5 loads them as such); pass dtype=torch.float32 for fp32, usually the faster choice on a CPU. The tokenizer adds no BOS/EOS: training records were separated by </s> only.

What this checkpoint wrote for that prompt (CPU, torch.manual_seed(0), 50 new tokens):

संस्कृतं नाम दैवी वाक् । तस्मिस्तु - - असौ मम पतिर्भां लक्ष्मी रमयतु । कीदृशः - असौ लक्ष्मीपतिः कामः तस्य वल्लभः । एण्विति

and with greedy decoding:

संस्कृतं नाम दैवी वाक् । असंस्कृतं वाक् । असंस्कृतं वाक् । असंस्कृतं वाक् । असंस्कृतं वाक् । असंस्कृतं वाक् । असंस्कृतं वाक् ।

To score text instead of generating it, call the model with labels=input_ids; loss is the mean negative log-likelihood per token in nats.

Model details

Architecture decoder-only transformer, pre-LayerNorm (LayerNorm without bias), no biases anywhere
Layers / heads / width 12 / 12 / 768 (head dim 64); MLP 4x = 3,072
Positions rotary (base 10,000; half-split pairing), no position table
Attention causal SDPA; queries and keys RMS-normalised per head (QK-norm)
MLP activation squared ReLU
Output head separate (untied), zero-initialised
Vocabulary 8,000 (SentencePiece unigram over SLP1, MuseMesh/sansar-sanskrit-tokenizer v0.1.0)
Context 512 tokens (about 3.6 kB of Devanagari text at the training data's 6.93 Devanagari bytes per token)
Parameters, total 97,241,856
Parameters, non-embedding 84,953,856
Token embedding 6,144,000
Output head 6,144,000
Weights in this repo bfloat16 safetensors (trained as fp32 master weights under bf16 autocast)

Training

Optimizer, block matrices Muon (84,934,656 params): momentum 0.95 (Nesterov), 5 Newton-Schulz steps in bf16, peak lr 0.02, update scaled by sqrt(max(1, rows/cols)), no weight decay
Optimizer, everything else AdamW: embedding + head (12,288,000 params) with weight decay 0.1, LayerNorm gains without; peak lr 0.001, betas (0.9, 0.95), eps 1e-8
Learning-rate schedule linear warm-up over 500 steps, then cosine decay to 0.1 x peak at step 54,212 (Muon and AdamW share it)
Batch 24 sequences x 4 accumulation steps x 512 tokens = 49,152 tokens per step
Steps 54,212
Tokens seen 2,664,628,224 (1.84 passes over 1,444.7M training tokens)
Gradient clipping global norm 1.0
Dropout 0.0
Precision bf16 autocast, fp32 master weights, torch.compile
Seed 1337
Data sampling 512-token windows at uniformly random offsets of the token stream (records separated by </s>), so passes are counted in expectation
Hardware 1x NVIDIA RTX 3060 12 GB, power-capped at 150 W
Wall time 24.7 h (89,076 s of training), 30.3k tokens/s
Compute cost own hardware, no cloud cost
Final validation loss 2.7557 nats/token = 0.5730 bits per Devanagari byte on the run's own validation split (0.5% of records by near-duplicate cluster; not comparable across data slices)

Training code: scripts/train/train.py, model.py and muon.py of the Sansar repository; the exact arguments are in training/summary.json.

Training data

train_slice_plus_clean (448.9M words, 15.1M records): corpus rebuild #5 without the archive.org OCR tier, plus TITUS (1.02M words) and grantha (0.99M words) with a held-out filter. After held-out masking: 15,012,243 training records, 1,444.7M tokens, 10.02 GB of Devanagari text. The run saw 2,664.6M tokens = 1.84 passes.

Held-out texts excluded by dedup key and masked inside training records (28-character windows, stride 4: 153,903 records masked, 16.3M characters removed, 22,736 records dropped). The Gītā is still partly memorised through near-copies in commentaries, so it stays out of the headline number.

The text comes from the Sansar corpus, which collects Sanskrit in Devanagari from public sources: classical e-text collections (GRETIL, SARIT, Muktabodha, the Digital Corpus of Sanskrit, DharmaNexus and others), Sanskrit Wikisource and Wikipedia, dictionaries, and the Sanskrit parts of web-crawl datasets (AI4Bharat Sangraha, IndicCorp, MADLAD-400, the sanskrit-monolingual-pretraining collection). Every record keeps its provenance and a licence tier (T0 permissive, T1 share-alike, T2 non-commercial, T3 no licence statement or all rights reserved). The training slice mixes all tiers: a large share is licensed for non-commercial use only and some sources state no licence, which is why the weights are released under CC BY-NC 4.0 (see Licence). Old archive.org OCR of printed books is left out (it measurably hurt the models).

Preparation: Unicode NFC; standalone / and // read as daṇḍa । and double daṇḍa ॥; machine reference markers (verse numbers of digital editions) stripped; transliterated to SLP1 and tokenized; one </s> after every record; a 0.5% validation split by near-duplicate cluster.

Evaluation

Metric: bits per Devanagari byte (lower is better): the model's negative log-likelihood of a text divided by the UTF-8 byte count of the same text in Devanagari, so models with different tokenizers are measured against the same denominator. Scored teacher-forced with scripts/train/eval_bpb.py: each set's records joined with </s> into one stream, 512-token windows with stride 256 (every scored token after the first window has at least 256 tokens of context), text cleaned exactly as the training text was.

Sets (E0, frozen before any model was trained and excluded from training): dcs_gold 3,000 sentences of the Digital Corpus of Sanskrit gold standard (classical), prose 2,470 prose passages, ood 2,000 web, Wikipedia and other out-of-domain texts, vedic 1,000 accented Ṛgveda pādas, gita all 700 verses of the Bhagavad-gītā. The headline is pooled excluding the Gītā (byte-weighted over the other four): the Gītā is quoted inside commentaries and epics throughout the training text, so its column measures memorisation.

ex-Gītā (headline) pooled, all five dcs_gold prose ood vedic gita (memorisation)
E0 sets (9,170 items) 0.6039 0.5920 0.6052 0.5859 0.6052 0.7833 0.2790
clean_v1 (8,270 items) 0.6254 0.6089 0.6079 0.5999 0.6347 0.7847 0.2860

Verse completion (600 verses: 200 each from the Bhagavad-gītā, Mahābhārata and Rāmāyaṇa; the model gets the first half-verse and greedily writes the second, stopping at ॥ or a newline): chrF 0.118, exact match 0/600. chrF credits shared character n-grams, so it rewards plausible vocabulary even when the half-verse is not the right one.

clean_v1 (added 2026-10-04): the same five sets minus every item that could overlap the training text in a way the held-out masking could not see. Some training text types the visarga as an ASCII colon (":" for "ः"); both the held-out mask and the leak gate ignored ":", so a 16-character window of a held-out item could survive inside such a record. clean_v1 removes every item with any such window in either training slice (900 of 9,170 items; mostly a shared stock phrase, rarely a real near-copy). It keeps 8,270 items; ood loses half its bytes and reads harder, so compare clean_v1 numbers only with clean_v1 numbers.

Reference points

Same metric and sets. External models were scored zero-shot with their own tokenizers and a 2,048-token window (stride 1,024), which gives them more context than our 512-token window; their training data may contain the public ood and prose texts.

model parameters ex-Gītā clean_v1 ex-Gītā notes
Sansar 20M v0.1.0 27.1M 0.7177 n/a TOK-v3 screen, arm v0.1, 0.34B tokens
Sansar 60M v0.1.0 63.2M 0.6434 n/a F0, 1.34B tokens
Sansar 125M v0.1.0 (this model) 97.2M 0.6039 0.6254 F6-clean-2x, 2.66B tokens
Sansar 350M v0.1.0 318.4M 0.5720 0.5937 F7, 2.66B tokens
Sansar 350M v0.2.0 318.4M 0.5547 0.5773 F9, 5.09B tokens
Krutrim-2-instruct (zero-shot, nf4) 12B 0.512 n/a general model, 2,048-token window
Sarvam-1 (zero-shot, fp16) 2.5B 0.750 n/a general model, 2,048-token window
Gemma 4 E2B (zero-shot, fp16) 4B 0.892 n/a general model, 2,048-token window

Checked before release

  • modeling_sansar.py against the training code (scripts/train/model.py), fp32 on CPU, 6 inputs (one held-out text per set, cropped to 512 tokens, and 512 random ids): with the same bf16 weights the logits are identical (max |difference| 0.0e+00); against the fp32 training weights the bf16 storage moves logits by at most 0.185.
  • Bits per byte on the first 20 records of each set with eval_bpb.py's own scoring: 0.55835 (training checkpoint, fp32) vs 0.55835 (this repo, bf16 weights) = -0.0015%.
  • Tokenizer: the same ids as the evaluation pipeline on 100/100 sample texts with fence_latin=False (95/100 with the default fence; the rest contain English words); Devanagari round trip exact on 100/100.
  • Left-padded batches give the same logits as single sequences (max |difference| 2.3e-05).

Limitations

  • Base model. It continues Sanskrit text. It has not been instruction-tuned or aligned: it does not follow instructions, answer questions or chat, and it will not refuse anything.
  • It makes things up. Output is fluent-looking Sanskrit that can be ungrammatical, mix registers and invent verses, authors and works. Do not use it as a source of quotations or facts; exact verse recall is close to zero (see verse completion).
  • Other languages in the data. The training text still holds some Hindi, Marathi and Pali lines that passed the Sanskrit filters of the time (a stricter filter came after this run), so the model can drift into Hindi.
  • Small and short. 97M parameters and a 512-token context. This implementation has no key/value cache, so long generations are slow (each new token re-reads the last 512).
  • Vedic is the weakest register: accent marks fragment the tokenization and Vedic text is a small share of the training data.
  • Script. Devanagari in and out. English words inside the text are fenced by the tokenizer so they come back unchanged; the model never saw the fence marks in training, so its predictions around them are weaker (AutoTokenizer.from_pretrained(..., fence_latin=False) reproduces the training encoding, but then Latin letters decode as Devanagari). IAST and other Indic scripts are not transliterated for you.
  • What the corpus says, the model says. Most of the text is religious, philosophical and classical literature, plus modern web and news Sanskrit; the model reproduces its views and its errors, including OCR errors that survived in web-crawl sources.
  • The Gītā is memorised in part (see the gita column), so do not read its score as generalisation.

Versions

Each version is a git tag on this repository; main is the newest. Pin one with revision="v0.1.0". A version is one training run, named by its run id in the Sansar experiment log.

version date training run tokens seen ex-Gītā clean_v1 ex-Gītā
v0.1.0 2026-10-02 f5_slp1_uni8k_d125m_plus_clean_2x (F6-clean-2x) 2.66B 0.6039 0.6254

All runs of this size (ex-Gītā on the same E0 sets):

run date recipe ex-Gītā status
F1 2026-09-14 AdamW, rebuild #3, whole-verse masking 0.6273* not staged
F1b (+ seed 2) 2026-09-14/15 AdamW, rebuild #3, window masking 0.6372 / 0.6410 not staged
F2 / F2-noocr / F3-mix12 2026-09-19/20 AdamW, rebuild #5 with 45% / 0% / 12.5% archive OCR 0.6399 / 0.6306 / 0.6336 not staged
F4 (+ seed 2) 2026-09-24/30 Muon + rotary/QK-norm/untied head/ReLU^2, noocr slice, 1.33B tokens 0.6189 / 0.6190 not staged
F5 curated / filtered / filtered v2 2026-09-26/28 F4 recipe on cleaned sub-slices (173M / 355M words) 0.6575 / 0.6238 / 0.6256 not staged
F6-clean 2026-10-01 F4 recipe, plus_clean slice, 1.33B tokens 0.6145 not staged
F8-small 2026-10-03 F4 recipe, plus_clean + tier 10 (509M words), 1.33B tokens 0.6135 not staged
F6-clean-2x 2026-10-02 F4 recipe, plus_clean slice, 2.66B tokens 0.6039 v0.1.0
* F1's numbers are flattered by a held-out leak that F1b's window masking removed.

Files

file what
model.safetensors the weights, bfloat16 (no optimizer state)
config.json, generation_config.json architecture and default sampling settings
configuration_sansar.py, modeling_sansar.py the network for transformers (auto_map, trust_remote_code)
tokenizer.model, tokenization_sansar.py, translit.py, tokenizer_config.json, special_tokens_map.json the tokenizer (MuseMesh/sansar-sanskrit-tokenizer v0.1.0) with its Devanagari <-> SLP1 wrapper
eval/bpb.json, eval/bpb_clean_v1.json, eval/verse_scores.json the evaluation outputs quoted above
eval/verification.json, eval/smoke_test.json the release checks and the sample generations
training/summary.json, training/data_meta.json every training argument, the loss curve's evaluation points, and the data preparation record
LICENSE, LICENSE-CODE, CHANGELOG.md licences and version history

Licence

This release is for research and non-commercial use. The training data includes texts licensed for non-commercial use only and texts without a licence statement, used here for research. A commercially licensed model, trained only on permissively licensed text, is planned as a separate release with its own version line.

  • Weights (model.safetensors) and tokenizer.model: CC BY-NC 4.0 (LICENSE). Non-commercial use (research, teaching, non-profit work) with attribution to "Sansar, Muse Mesh Private Limited".
  • Code (configuration_sansar.py, modeling_sansar.py, tokenization_sansar.py, translit.py): Apache-2.0 (LICENSE-CODE).
  • Commercial use: contact kushal@muse-mesh.com.

Citation

@misc{sansar_125m_2026,
  title  = {Sansar 125M: a Sanskrit language model},
  author = {Muse Mesh},
  year   = {2026},
  note   = {v0.1.0},
  url    = {https://huggingface.co/MuseMesh/sansar-125m}
}

Contact: kushal@muse-mesh.com

Downloads last month
243
Safetensors
Model size
97.2M params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train MuseMesh/sansar-125m

Collection including MuseMesh/sansar-125m

Evaluation results

  • bits per Devanagari byte, pooled excluding the Bhagavad-gītā (E0 held-out sets) on Sansar E0 held-out Sanskrit sets
    self-reported
    0.604
  • bits per Devanagari byte, pooled over all five E0 sets on Sansar E0 held-out Sanskrit sets
    self-reported
    0.592
  • bits per Devanagari byte, dcs_gold on Sansar E0 held-out Sanskrit sets
    self-reported
    0.605
  • bits per Devanagari byte, prose on Sansar E0 held-out Sanskrit sets
    self-reported
    0.586
  • bits per Devanagari byte, ood on Sansar E0 held-out Sanskrit sets
    self-reported
    0.605
  • bits per Devanagari byte, vedic on Sansar E0 held-out Sanskrit sets
    self-reported
    0.783
  • bits per Devanagari byte, gita on Sansar E0 held-out Sanskrit sets
    self-reported
    0.279
  • bits per Devanagari byte, pooled excluding the Bhagavad-gītā (clean_v1 sets) on Sansar E0 held-out Sanskrit sets
    self-reported
    0.625