Instructions to use MuseMesh/sansar-125m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MuseMesh/sansar-125m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="MuseMesh/sansar-125m", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("MuseMesh/sansar-125m", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use MuseMesh/sansar-125m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "MuseMesh/sansar-125m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuseMesh/sansar-125m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/MuseMesh/sansar-125m
- SGLang
How to use MuseMesh/sansar-125m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "MuseMesh/sansar-125m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuseMesh/sansar-125m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "MuseMesh/sansar-125m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuseMesh/sansar-125m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use MuseMesh/sansar-125m with Docker Model Runner:
docker model run hf.co/MuseMesh/sansar-125m
Sansar 125M
A 97.2M-parameter Sanskrit language model trained from scratch, from the Sansar project at Muse Mesh (Hugging Face): Sanskrit language models, their corpus and their tokenizer. Size class: the d125m preset, GPT-2-small shape: 12 layers, width 768 (85.0M parameters outside the embeddings; 97.2M in total with the 8k vocabulary and untied head).
- This version: v0.1.0 = training run
f5_slp1_uni8k_d125m_plus_clean_2x, finished 2026-10-02 (experiment F6-clean-2x: F6-clean data and recipe at two passes). - Held-out score: 0.6039 bits per Devanagari byte (pooled over four held-out sets, excluding the Bhagavad-gītā; lower is better); 0.6254 on the stricter clean_v1 sets.
- Base model: it continues Devanagari Sanskrit text; it is not instruction-tuned.
- Why this checkpoint: The best 125M-class checkpoint on the held-out ex-Gītā number.
Quick start
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "MuseMesh/sansar-125m"
tok = AutoTokenizer.from_pretrained(repo, revision="v0.1.0", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, revision="v0.1.0", trust_remote_code=True)
inputs = tok("संस्कृतं नाम दैवी वाक्", return_tensors="pt") # Devanagari in
out = model.generate(**inputs, max_new_tokens=50, do_sample=True, temperature=0.8, top_k=100)
print(tok.decode(out[0], skip_special_tokens=True)) # Devanagari out
Needs transformers (tested with 5.18.0), torch and sentencepiece. trust_remote_code=True runs three small files from this repository: modeling_sansar.py (the network), tokenization_sansar.py and translit.py (Devanagari <-> SLP1); read them before you run them. The weights are stored in bfloat16 (transformers 5 loads them as such); pass dtype=torch.float32 for fp32, usually the faster choice on a CPU. The tokenizer adds no BOS/EOS: training records were separated by </s> only.
What this checkpoint wrote for that prompt (CPU, torch.manual_seed(0), 50 new tokens):
संस्कृतं नाम दैवी वाक् । तस्मिस्तु - - असौ मम पतिर्भां लक्ष्मी रमयतु । कीदृशः - असौ लक्ष्मीपतिः कामः तस्य वल्लभः । एण्विति
and with greedy decoding:
संस्कृतं नाम दैवी वाक् । असंस्कृतं वाक् । असंस्कृतं वाक् । असंस्कृतं वाक् । असंस्कृतं वाक् । असंस्कृतं वाक् । असंस्कृतं वाक् ।
To score text instead of generating it, call the model with labels=input_ids; loss is the mean negative log-likelihood per token in nats.
Model details
| Architecture | decoder-only transformer, pre-LayerNorm (LayerNorm without bias), no biases anywhere |
| Layers / heads / width | 12 / 12 / 768 (head dim 64); MLP 4x = 3,072 |
| Positions | rotary (base 10,000; half-split pairing), no position table |
| Attention | causal SDPA; queries and keys RMS-normalised per head (QK-norm) |
| MLP activation | squared ReLU |
| Output head | separate (untied), zero-initialised |
| Vocabulary | 8,000 (SentencePiece unigram over SLP1, MuseMesh/sansar-sanskrit-tokenizer v0.1.0) |
| Context | 512 tokens (about 3.6 kB of Devanagari text at the training data's 6.93 Devanagari bytes per token) |
| Parameters, total | 97,241,856 |
| Parameters, non-embedding | 84,953,856 |
| Token embedding | 6,144,000 |
| Output head | 6,144,000 |
| Weights in this repo | bfloat16 safetensors (trained as fp32 master weights under bf16 autocast) |
Training
| Optimizer, block matrices | Muon (84,934,656 params): momentum 0.95 (Nesterov), 5 Newton-Schulz steps in bf16, peak lr 0.02, update scaled by sqrt(max(1, rows/cols)), no weight decay |
| Optimizer, everything else | AdamW: embedding + head (12,288,000 params) with weight decay 0.1, LayerNorm gains without; peak lr 0.001, betas (0.9, 0.95), eps 1e-8 |
| Learning-rate schedule | linear warm-up over 500 steps, then cosine decay to 0.1 x peak at step 54,212 (Muon and AdamW share it) |
| Batch | 24 sequences x 4 accumulation steps x 512 tokens = 49,152 tokens per step |
| Steps | 54,212 |
| Tokens seen | 2,664,628,224 (1.84 passes over 1,444.7M training tokens) |
| Gradient clipping | global norm 1.0 |
| Dropout | 0.0 |
| Precision | bf16 autocast, fp32 master weights, torch.compile |
| Seed | 1337 |
| Data sampling | 512-token windows at uniformly random offsets of the token stream (records separated by </s>), so passes are counted in expectation |
| Hardware | 1x NVIDIA RTX 3060 12 GB, power-capped at 150 W |
| Wall time | 24.7 h (89,076 s of training), 30.3k tokens/s |
| Compute cost | own hardware, no cloud cost |
| Final validation | loss 2.7557 nats/token = 0.5730 bits per Devanagari byte on the run's own validation split (0.5% of records by near-duplicate cluster; not comparable across data slices) |
Training code: scripts/train/train.py, model.py and muon.py of the Sansar repository; the exact arguments are in training/summary.json.
Training data
train_slice_plus_clean (448.9M words, 15.1M records): corpus rebuild #5 without the archive.org OCR tier, plus TITUS (1.02M words) and grantha (0.99M words) with a held-out filter. After held-out masking: 15,012,243 training records, 1,444.7M tokens, 10.02 GB of Devanagari text. The run saw 2,664.6M tokens = 1.84 passes.
Held-out texts excluded by dedup key and masked inside training records (28-character windows, stride 4: 153,903 records masked, 16.3M characters removed, 22,736 records dropped). The Gītā is still partly memorised through near-copies in commentaries, so it stays out of the headline number.
The text comes from the Sansar corpus, which collects Sanskrit in Devanagari from public sources: classical e-text collections (GRETIL, SARIT, Muktabodha, the Digital Corpus of Sanskrit, DharmaNexus and others), Sanskrit Wikisource and Wikipedia, dictionaries, and the Sanskrit parts of web-crawl datasets (AI4Bharat Sangraha, IndicCorp, MADLAD-400, the sanskrit-monolingual-pretraining collection). Every record keeps its provenance and a licence tier (T0 permissive, T1 share-alike, T2 non-commercial, T3 no licence statement or all rights reserved). The training slice mixes all tiers: a large share is licensed for non-commercial use only and some sources state no licence, which is why the weights are released under CC BY-NC 4.0 (see Licence). Old archive.org OCR of printed books is left out (it measurably hurt the models).
Preparation: Unicode NFC; standalone / and // read as daṇḍa । and double daṇḍa ॥; machine reference markers (verse numbers of digital editions) stripped; transliterated to SLP1 and tokenized; one </s> after every record; a 0.5% validation split by near-duplicate cluster.
Evaluation
Metric: bits per Devanagari byte (lower is better): the model's negative log-likelihood of a text divided by the UTF-8 byte count of the same text in Devanagari, so models with different tokenizers are measured against the same denominator. Scored teacher-forced with scripts/train/eval_bpb.py: each set's records joined with </s> into one stream, 512-token windows with stride 256 (every scored token after the first window has at least 256 tokens of context), text cleaned exactly as the training text was.
Sets (E0, frozen before any model was trained and excluded from training): dcs_gold 3,000 sentences of the Digital Corpus of Sanskrit gold standard (classical), prose 2,470 prose passages, ood 2,000 web, Wikipedia and other out-of-domain texts, vedic 1,000 accented Ṛgveda pādas, gita all 700 verses of the Bhagavad-gītā. The headline is pooled excluding the Gītā (byte-weighted over the other four): the Gītā is quoted inside commentaries and epics throughout the training text, so its column measures memorisation.
| ex-Gītā (headline) | pooled, all five | dcs_gold | prose | ood | vedic | gita (memorisation) | |
|---|---|---|---|---|---|---|---|
| E0 sets (9,170 items) | 0.6039 | 0.5920 | 0.6052 | 0.5859 | 0.6052 | 0.7833 | 0.2790 |
| clean_v1 (8,270 items) | 0.6254 | 0.6089 | 0.6079 | 0.5999 | 0.6347 | 0.7847 | 0.2860 |
Verse completion (600 verses: 200 each from the Bhagavad-gītā, Mahābhārata and Rāmāyaṇa; the model gets the first half-verse and greedily writes the second, stopping at ॥ or a newline): chrF 0.118, exact match 0/600. chrF credits shared character n-grams, so it rewards plausible vocabulary even when the half-verse is not the right one.
clean_v1 (added 2026-10-04): the same five sets minus every item that could overlap the training text in a way the held-out masking could not see. Some training text types the visarga as an ASCII colon (":" for "ः"); both the held-out mask and the leak gate ignored ":", so a 16-character window of a held-out item could survive inside such a record. clean_v1 removes every item with any such window in either training slice (900 of 9,170 items; mostly a shared stock phrase, rarely a real near-copy). It keeps 8,270 items; ood loses half its bytes and reads harder, so compare clean_v1 numbers only with clean_v1 numbers.
Reference points
Same metric and sets. External models were scored zero-shot with their own tokenizers and a 2,048-token window (stride 1,024), which gives them more context than our 512-token window; their training data may contain the public ood and prose texts.
| model | parameters | ex-Gītā | clean_v1 ex-Gītā | notes |
|---|---|---|---|---|
| Sansar 20M v0.1.0 | 27.1M | 0.7177 | n/a | TOK-v3 screen, arm v0.1, 0.34B tokens |
| Sansar 60M v0.1.0 | 63.2M | 0.6434 | n/a | F0, 1.34B tokens |
| Sansar 125M v0.1.0 (this model) | 97.2M | 0.6039 | 0.6254 | F6-clean-2x, 2.66B tokens |
| Sansar 350M v0.1.0 | 318.4M | 0.5720 | 0.5937 | F7, 2.66B tokens |
| Sansar 350M v0.2.0 | 318.4M | 0.5547 | 0.5773 | F9, 5.09B tokens |
| Krutrim-2-instruct (zero-shot, nf4) | 12B | 0.512 | n/a | general model, 2,048-token window |
| Sarvam-1 (zero-shot, fp16) | 2.5B | 0.750 | n/a | general model, 2,048-token window |
| Gemma 4 E2B (zero-shot, fp16) | 4B | 0.892 | n/a | general model, 2,048-token window |
Checked before release
modeling_sansar.pyagainst the training code (scripts/train/model.py), fp32 on CPU, 6 inputs (one held-out text per set, cropped to 512 tokens, and 512 random ids): with the same bf16 weights the logits are identical (max |difference| 0.0e+00); against the fp32 training weights the bf16 storage moves logits by at most 0.185.- Bits per byte on the first 20 records of each set with
eval_bpb.py's own scoring: 0.55835 (training checkpoint, fp32) vs 0.55835 (this repo, bf16 weights) = -0.0015%. - Tokenizer: the same ids as the evaluation pipeline on 100/100 sample texts with
fence_latin=False(95/100 with the default fence; the rest contain English words); Devanagari round trip exact on 100/100. - Left-padded batches give the same logits as single sequences (max |difference| 2.3e-05).
Limitations
- Base model. It continues Sanskrit text. It has not been instruction-tuned or aligned: it does not follow instructions, answer questions or chat, and it will not refuse anything.
- It makes things up. Output is fluent-looking Sanskrit that can be ungrammatical, mix registers and invent verses, authors and works. Do not use it as a source of quotations or facts; exact verse recall is close to zero (see verse completion).
- Other languages in the data. The training text still holds some Hindi, Marathi and Pali lines that passed the Sanskrit filters of the time (a stricter filter came after this run), so the model can drift into Hindi.
- Small and short. 97M parameters and a 512-token context. This implementation has no key/value cache, so long generations are slow (each new token re-reads the last 512).
- Vedic is the weakest register: accent marks fragment the tokenization and Vedic text is a small share of the training data.
- Script. Devanagari in and out. English words inside the text are fenced by the tokenizer so they come back unchanged; the model never saw the fence marks in training, so its predictions around them are weaker (
AutoTokenizer.from_pretrained(..., fence_latin=False)reproduces the training encoding, but then Latin letters decode as Devanagari). IAST and other Indic scripts are not transliterated for you. - What the corpus says, the model says. Most of the text is religious, philosophical and classical literature, plus modern web and news Sanskrit; the model reproduces its views and its errors, including OCR errors that survived in web-crawl sources.
- The Gītā is memorised in part (see the
gitacolumn), so do not read its score as generalisation.
Versions
Each version is a git tag on this repository; main is the newest. Pin one with revision="v0.1.0". A version is one training run, named by its run id in the Sansar experiment log.
| version | date | training run | tokens seen | ex-Gītā | clean_v1 ex-Gītā |
|---|---|---|---|---|---|
| v0.1.0 | 2026-10-02 | f5_slp1_uni8k_d125m_plus_clean_2x (F6-clean-2x) |
2.66B | 0.6039 | 0.6254 |
All runs of this size (ex-Gītā on the same E0 sets):
| run | date | recipe | ex-Gītā | status |
|---|---|---|---|---|
| F1 | 2026-09-14 | AdamW, rebuild #3, whole-verse masking | 0.6273* | not staged |
| F1b (+ seed 2) | 2026-09-14/15 | AdamW, rebuild #3, window masking | 0.6372 / 0.6410 | not staged |
| F2 / F2-noocr / F3-mix12 | 2026-09-19/20 | AdamW, rebuild #5 with 45% / 0% / 12.5% archive OCR | 0.6399 / 0.6306 / 0.6336 | not staged |
| F4 (+ seed 2) | 2026-09-24/30 | Muon + rotary/QK-norm/untied head/ReLU^2, noocr slice, 1.33B tokens | 0.6189 / 0.6190 | not staged |
| F5 curated / filtered / filtered v2 | 2026-09-26/28 | F4 recipe on cleaned sub-slices (173M / 355M words) | 0.6575 / 0.6238 / 0.6256 | not staged |
| F6-clean | 2026-10-01 | F4 recipe, plus_clean slice, 1.33B tokens | 0.6145 | not staged |
| F8-small | 2026-10-03 | F4 recipe, plus_clean + tier 10 (509M words), 1.33B tokens | 0.6135 | not staged |
| F6-clean-2x | 2026-10-02 | F4 recipe, plus_clean slice, 2.66B tokens | 0.6039 | v0.1.0 |
| * F1's numbers are flattered by a held-out leak that F1b's window masking removed. |
Files
| file | what |
|---|---|
model.safetensors |
the weights, bfloat16 (no optimizer state) |
config.json, generation_config.json |
architecture and default sampling settings |
configuration_sansar.py, modeling_sansar.py |
the network for transformers (auto_map, trust_remote_code) |
tokenizer.model, tokenization_sansar.py, translit.py, tokenizer_config.json, special_tokens_map.json |
the tokenizer (MuseMesh/sansar-sanskrit-tokenizer v0.1.0) with its Devanagari <-> SLP1 wrapper |
eval/bpb.json, eval/bpb_clean_v1.json, eval/verse_scores.json |
the evaluation outputs quoted above |
eval/verification.json, eval/smoke_test.json |
the release checks and the sample generations |
training/summary.json, training/data_meta.json |
every training argument, the loss curve's evaluation points, and the data preparation record |
LICENSE, LICENSE-CODE, CHANGELOG.md |
licences and version history |
Licence
This release is for research and non-commercial use. The training data includes texts licensed for non-commercial use only and texts without a licence statement, used here for research. A commercially licensed model, trained only on permissively licensed text, is planned as a separate release with its own version line.
- Weights (
model.safetensors) andtokenizer.model: CC BY-NC 4.0 (LICENSE). Non-commercial use (research, teaching, non-profit work) with attribution to "Sansar, Muse Mesh Private Limited". - Code (
configuration_sansar.py,modeling_sansar.py,tokenization_sansar.py,translit.py): Apache-2.0 (LICENSE-CODE). - Commercial use: contact kushal@muse-mesh.com.
Citation
@misc{sansar_125m_2026,
title = {Sansar 125M: a Sanskrit language model},
author = {Muse Mesh},
year = {2026},
note = {v0.1.0},
url = {https://huggingface.co/MuseMesh/sansar-125m}
}
Contact: kushal@muse-mesh.com
- Downloads last month
- 243
Datasets used to train MuseMesh/sansar-125m
chronbmm/sanskrit-monolingual-pretraining
MuseMesh/sansar-sanskrit-corpus
Collection including MuseMesh/sansar-125m
Evaluation results
- bits per Devanagari byte, pooled excluding the Bhagavad-gītā (E0 held-out sets) on Sansar E0 held-out Sanskrit setsself-reported0.604
- bits per Devanagari byte, pooled over all five E0 sets on Sansar E0 held-out Sanskrit setsself-reported0.592
- bits per Devanagari byte, dcs_gold on Sansar E0 held-out Sanskrit setsself-reported0.605
- bits per Devanagari byte, prose on Sansar E0 held-out Sanskrit setsself-reported0.586
- bits per Devanagari byte, ood on Sansar E0 held-out Sanskrit setsself-reported0.605
- bits per Devanagari byte, vedic on Sansar E0 held-out Sanskrit setsself-reported0.783
- bits per Devanagari byte, gita on Sansar E0 held-out Sanskrit setsself-reported0.279
- bits per Devanagari byte, pooled excluding the Bhagavad-gītā (clean_v1 sets) on Sansar E0 held-out Sanskrit setsself-reported0.625