Chandohasam pilot120m: a 120M masked diffusion language model for Telugu

pilot120m is a 120M-parameter masked diffusion language model (MDLM; Sahoo et al., 2024) for Telugu, trained from scratch on 4.98B tokens of web, news and encyclopedic text. Instead of writing left to right, it fills a canvas of 512 <mask> tokens in any order, seeing the whole canvas at every step. Its tokenizer gives one token per akshara (orthographic syllable), so every token is a unit of Telugu metre.

It is stage 1 of Chandohasam, a project on generating Telugu metrical poetry (padyams), whose metres fix the weight of every syllable, a rhyme in the second syllable of every line (prāsa) and a caesura agreement (yati). This stage teaches general Telugu; a later stage specialises the model to verse, decoded under exact metrical constraints. It has seen no poetry: documents that overlap a large collection of known Telugu poems were removed from its training data.

Model type Masked (absorbing-state) discrete diffusion LM; bidirectional transformer denoiser, no timestep input
Language Telugu (te)
Parameters 120,048,384 (84,953,856 outside the tied embedding)
Canvas 512 tokens, about 110 words of prose
Tokenizer 45,591 tokens: one per akshara, byte-level BPE fallback, no out-of-vocabulary input
Training 76,000 steps × 65,536 tokens = 4.98B tokens, on one 8 GB laptop GPU in 2 days 19 hours
Weights The EMA (decay 0.9999) at step 76,000, float32 (model.safetensors)
Held-out test NELBO 1.783 nats per token, weighted like the training mix
License MIT (weights and code)

Contents: Quick start · Intended use and limitations · Architecture · Tokenizer · Diffusion objective · Sampling · Training data · Training procedure · Evaluation · Growth compatibility · Files · License · Citation · References

Quick start

The model is a custom architecture; the code to load and sample it is in this repository (chandohasam_mdlm/). It needs Python ≥ 3.10, PyTorch ≥ 2.4 and safetensors.

pip install torch safetensors huggingface_hub
import sys
from huggingface_hub import snapshot_download

repo = snapshot_download("samvaran/chandohasam")
sys.path.insert(0, repo)
from chandohasam_mdlm import load, generate

model, tok = load(repo, device="cuda")           # or device="cpu"
for text in generate(model, tok, n=2, steps=512, temperature=0.9, seed=0):
    print(text, "\n---")

# fix the start of every canvas
print(generate(model, tok, n=1, prefix="తెలుగు భాష", seed=1)[0])

Or from a clone of the repository: python generate.py --n 4 --steps 512 --temperature 0.9.

generate starts from 512 <mask> tokens and unmasks them with the ancestral sampler. 512 steps at temperature 0.9 gave the most fluent samples in the evaluation (Sampling). On the training GPU (RTX 5050 Laptop), 64 canvases at 512 steps take about 12 minutes.

The denoiser can also be used directly: model(x) maps token ids x (batch × ≤ 512, with tok.special_to_id["<mask>"] at the positions to predict) to logits over the padded vocabulary of 45,696. Before taking a softmax, add invalid_bias(...) from chandohasam_mdlm, which sets <mask>, <pad>, <unk> and the 105 padding rows to −∞, as in training.

Intended use and limitations

Intended use. Research on masked diffusion language models for Telugu, and a starting point for continued pretraining or fine-tuning, in particular for constrained generation where each position is an akshara: metrical verse, infilling, and decoding steered by a rule engine.

It is a base model. It is not instruction-tuned, chat-tuned or safety-tuned, and it is not a source of facts. Its unconditional samples are fluent phrases in a web, news or blog register without a coherent thread across sentences. At the recommended setting, an independent judge (Gemma-3-1B) rates them at a perplexity of 45.8, against 16.9 for real Telugu text, and 86.4% of their words are real words, against 88.7% in real text.

Known limitations.

  • The heavily masked regime is weak. With 85% of the tokens masked, the model predicts 31–32% of them correctly. Generation starts from a fully masked canvas, so its first steps run in this regime, and more denoising steps barely help (Evaluation).
  • Confidence-ordered unmasking collapses into repeated tokens for unconditional generation. Use the ancestral sampler.
  • The data is web and news text. Generated text can mention real people, places and events and say false things about them, and it can reproduce biases and offensive language present in web text. Nothing beyond the cleaning described below was done to filter content.
  • Telugu only. Other scripts are kept losslessly as bytes by the tokenizer but are not modelled.
  • No poetry. Poems were deliberately removed from the training data, so the model knows the language of verse only as far as it appears in prose.
  • The canvas is fixed at 512 tokens. Training windows are 512 tokens cut from concatenated documents, so a canvas may contain the end of one document and the start of the next, separated by <eos> <bos>.

Architecture

pilot120m architecture

Hyperparameters

Field Value Meaning
vocab_size 45,591 Tokenizer vocabulary: 5 specials, 256 bytes, 485 BPE pieces, 10 Telugu digits and 44,835 whole aksharas. Padded to 45,696 (a multiple of 128) for the GPU; the 105 padding rows are never predicted
d_model 768 Width of the residual stream, the token embeddings and every block's input and output
n_layers 12 Transformer blocks
n_heads 12 Attention heads per block; head size 768 / 12 = 64
mlp_hidden 2,048 SwiGLU hidden width per block (8/3 × d_model)
max_len 512 Canvas length in tokens; the rotary tables are built for 512 positions
rope_base 10,000 Base frequency of the rotary position embeddings
norm_eps 1e-6 RMSNorm epsilon
init_std 0.02 Standard deviation of the normal initialisation of embeddings and linear layers; the two residual output projections of each block use 0.02 / √(2 × 12) ≈ 0.00408

Design choices

Choice Setting
Block Pre-norm: x + Attn(RMSNorm(x)), then x + SwiGLU(RMSNorm(x)); a final RMSNorm before the output layer
Attention Bidirectional (no causal mask); torch.nn.functional.scaled_dot_product_attention, softmax scale 1/√64; one fused query/key/value projection
Positions Rotary embeddings (RoPE) on queries and keys, rotating interleaved channel pairs; no learned position embedding
MLP SwiGLU: down(SiLU(gate(x)) ⊙ up(x)), with the gate and up projections fused
Biases, dropout None
Time conditioning None. Under masking noise the optimal denoiser does not depend on t (Ou et al., 2025; Zheng et al., 2025), and LLaDA (Nie et al., 2025) also drops it
Embeddings Input and output tied: one matrix E, and logits = h · Eᵀ
Precision in training float32 master weights and optimizer state, bf16 autocast for the forward and backward passes, float32 cross-entropy

Weight tensors

All 74 tensors (i = 0…11):

Tensor Shape Parameters Initialisation Weight decay Role
embed.weight 45,696 × 768 35,094,528 N(0, 0.02) 0.1 Token embedding; tied, so also the output layer
blocks.i.norm1.weight 768 768 ones 0 RMSNorm gain before attention
blocks.i.qkv.weight 2,304 × 768 1,769,472 N(0, 0.02) 0.1 Fused query/key/value projection
blocks.i.proj.weight 768 × 768 589,824 N(0, 0.00408) 0.1 Attention output projection (residual branch)
blocks.i.norm2.weight 768 768 ones 0 RMSNorm gain before the MLP
blocks.i.gate_up.weight 4,096 × 768 3,145,728 N(0, 0.02) 0.1 Fused SwiGLU gate and up projections
blocks.i.down.weight 768 × 2,048 1,572,864 N(0, 0.00408) 0.1 SwiGLU down projection (residual branch)
norm.weight 768 768 ones 0 Final RMSNorm gain
Total 120,048,384

Each block holds 7,079,424 parameters, so the 12 blocks hold 84,953,088; the final norm (768) and the tied embedding (35,094,528) make up the rest.

Tokenizer

  • Vocabulary. telugu_alldomain.tokenizer.json (sha256 200c814d84cc1b88bccdea3e4131a4c2fee296d050587e5d3da99f22d73d30fc), 45,591 tokens: 44,835 whole aksharas, 485 byte-level BPE pieces, 256 bytes, 10 Telugu digits and 5 specials: <pad> = 0, <unk> = 1, <bos> = 2, <eos> = 3, <mask> = 4.
  • One token per akshara. Text is split into aksharas by aksharanusarika (included, MIT), the same splitter as the Chandohasam metre engine, so token boundaries are metrical boundaries. The arasunna ఁ stays on its akshara and zero-width non-joiners are removed.
  • No out-of-vocabulary input. An akshara without its own token falls back to byte-level BPE. On unseen Sangraha news, 99.94% of syllables are single tokens.
  • Lossless. decode(encode(text)) returns the normalised text.

Diffusion objective

Item Setting
Forward (noising) process Each token is replaced by <mask> independently with probability t: the absorbing-state process with a log-linear schedule, α_t = 1 − t
Time sampling t ~ U(ε, 1) with ε = 0.001, spread evenly across the batch: one random offset u, and t_i = ε + (1 − ε)·((u + i/B) mod 1)
Parameterisation SUBS: the logits of <mask>, <pad>, <unk> and the padding rows are −∞, and unmasked tokens are copied through, with no loss on them
Loss NELBO per token: the sum over masked positions of cross-entropy / t, divided by batch × length. In nats per token; exp(NELBO) bounds the perplexity
Memory The output layer runs only at masked positions, in chunks of 2,048 whose logits are recomputed in the backward pass, so the 45,696-wide output layer fits in 8 GB

Sampling

  • Ancestral sampler (chandohasam_mdlm/sampling.py), MDLM's: going from t to s < t, each still-masked position is revealed with probability (t − s)/t, its token drawn from the model. Because the model has no time input, a step that reveals nothing reuses the previous forward pass.
  • Confidence order (MaskGIT/LLaDA style) is also implemented: each step reveals the masked positions whose sampled token is most probable. For unconditional generation it collapses into repetition.
  • Categorical draws use float64 Gumbel noise. Float32 noise quietly lowers the sampling temperature and flatters sample quality (Zheng et al., 2025).
  • Recommended: ancestral, 512 steps, temperature 0.9.

Training data

Sources

Source Hugging Face dataset Files (pinned revision) License Download
Sangraha, verified split, Telugu (Khan et al., 2024) ai4bharat/sangraha verified/tel/*.parquet @ 8b813c3f CC BY 4.0 15.1 GB
IndicCorp v2, Telugu (Doddapaneni et al., 2023) ai4bharat/IndicCorpV2 data/te.txt @ 2d7285e6 not stated on the dataset card 15.8 GB
Telugu Wikipedia, dump of 2023-11-01 wikimedia/wikipedia 20231101.te/*.parquet @ b04c8d1c CC BY-SA 3.0, GFDL 0.2 GB

Cleaning, deduplication and decontamination

  • Normalisation. NFC, the tokenizer's normalisation, control characters removed, whitespace collapsed, literal \n sequences turned into newlines.
  • Exact duplicates and boilerplate. Exact-duplicate documents are dropped, and so are boilerplate lines (any line found in 10 or more distinct documents; 90,446 such lines). IndicCorp paragraphs that already occur inside a Wikipedia or Sangraha document (5,518,447) are dropped.
  • Cleaning. URLs and e-mail addresses are cut out of their lines. A line is dropped if under 50% of its letters are Telugu. A document is dropped if it has fewer than 40 Telugu letters, if under 80% of its letters are Telugu, or if more than 30% of its lines are repeats. After tokenization, a document is dropped if more than 2% of its Telugu pieces have no whole-akshara token (garbled text), or if under half of its tokens carry text rather than whitespace, ASCII punctuation or digits (tables and number lists).
  • Near duplicates. MinHash (64 permutations over 8-token shingles), clustered at an estimated Jaccard similarity ≥ 0.8 (8 bands × 8 rows), one document kept per cluster; for Wikipedia and Sangraha.
  • Decontamination. Any document that shares a 20-token window with a large collection of known Telugu poems is dropped, so that the later poetry stage is not contaminated.
  • Splits. Documents go to train, validation (0.5%) and test (0.5%) by a stable hash. Each document is stored as <bos> … <eos>.

Documents at each stage:

Source Raw Duplicate Cleaning rejects Token-level rejects Near duplicate Poem overlap Kept
Wikipedia 87,854 100 1,950 561 4,963 205 80,075
Sangraha 7,081,734 77 77,947 6,977 115,540 68,888 6,812,305
IndicCorp (paragraphs) 21,458,261 7,546,348 5,642,341 56,809 0 573 8,212,190

For IndicCorp, "duplicate" includes the paragraphs already present in another source, and most cleaning rejects are paragraphs too short to keep.

Tokens, including <bos> and <eos>:

Source Train Validation Test
Wikipedia 102,634,991 560,104 497,202
Sangraha 7,327,831,173 38,164,489 37,278,696
IndicCorp 1,981,820,075 9,970,079 10,195,016
Total 9,412,286,239 48,694,672 47,970,914

Mixture

Source Sampling weight Tokens seen Passes over its train split
Sangraha 75% 3.74B 0.51
IndicCorp 20% 1.00B 0.50
Wikipedia 5% 0.25B 2.43

Each training example is 512 consecutive tokens from a random offset in the chosen source's train split, so a window may span several documents (MDLM's "wrapped" setting).

Training procedure

Setting Value
Optimizer AdamW (fused), β₁ = 0.9, β₂ = 0.98, ε = 1e-8
Learning rate Linear warm-up over 2,000 steps to 3e-4, then cosine decay to 3e-5 at step 76,000
Weight decay 0.1 on all matrices including the tied embedding; none on RMSNorm gains
Gradient clipping 1.0 (global norm)
Batch 128 sequences × 512 tokens = 65,536 tokens per step, as 16 accumulated micro-batches of 8
Steps, tokens 76,000 steps, 4,980,736,000 tokens: 41 tokens per parameter
EMA Decay 0.9999, warmed up as min(0.9999, (1 + step)/(10 + step)); the EMA weights are the ones released and evaluated
Seed 0; batch composition, mask rates and masks are a function of (seed, step)
Compilation torch.compile on the transformer; the output layer runs eagerly
Hardware One NVIDIA GeForce RTX 5050 Laptop GPU (8 GB), 16 CPU threads, 23 GB RAM
Software Python 3.12, PyTorch 2.14.0 with CUDA 13.0, TF32 matmuls
Throughput Median 20.8k tokens/s (3.15 s per step), about 1.8B tokens per day; 5.7 GB of GPU memory, 6.7 GB peak
Wall clock 2 days 19 hours (26 to 29 September 2026)

learning-rate schedule and gradient norm

The run never diverged, never ran out of memory and never skipped an update. It rode through two mains-power cuts on battery without losing work.

Evaluation

Likelihood

NELBO in nats per token (lower is better) of the released weights, on 512 fixed canvases per source with fixed mask rates and masks. exp(NELBO) is an upper bound on the per-token perplexity.

Split Sangraha IndicCorp Wikipedia Mean of the three Weighted like the training mix
Validation 1.899 1.719 0.994 1.537 1.818
Test 1.846 1.729 1.053 1.543 1.783

The test split was untouched until this evaluation and is no harder than validation. Wikipedia scores low because, beyond its first 128 canvases, its held-out text is dominated by templated village articles, which are nearly predictable; the per-source and mix-weighted values are the ones to compare.

loss curves

During training, the validation NELBO on 128 fixed canvases per source fell from 2.921 at step 1,000 to 1.747 at step 76,000, and was still falling slowly at the end.

Masked-token accuracy

Top-1 accuracy on masked tokens (test split): 87.9% with 15% of the tokens masked, 69.1% with 50% and 31.4% with 85% (validation: 89.4%, 68.8%, 32.3%).

masked-token accuracy by mask rate

Sample quality and sampler settings

64 canvases of 512 tokens per setting, generated from an all-mask canvas with the released weights. Judge perplexity is the perplexity of the samples under an independent model, google/gemma-3-1b-pt (Gemma Team, 2025); on real validation text it is 16.9, and 64.9 when the words of that text are shuffled. Real words: the share of generated words that occur in validation text (88.7% for real text).

Unmasking order Steps Temperature Judge perplexity Real words Token entropy Distinct-1 Repeated 4-grams
ancestral 256 1.0 71.5 77.5% 4.35 0.734 0.0%
ancestral 512 1.0 65.8 77.8% 4.32 0.729 0.0%
ancestral 1,024 1.0 64.9 78.8% 4.33 0.725 0.1%
confidence 256 1.0 1.7 99.7% 0.75 0.006 94.0%
confidence 512 1.0 1.4 99.7% 0.72 0.004 95.9%
ancestral 512 0.9 45.8 86.4% 4.18 0.629 0.0%
real Telugu text 16.9 88.7% 4.32 0.73

sampling study

  • The model, not the sampler, is the limit. Four times as many steps (256 → 1,024) bring judge perplexity only from 71.5 to 64.9, the level of word-shuffled real text.
  • Confidence order collapses. It fills the canvas with a few repeated tokens; the judge scores that as nearly perfect, which token entropy and repeated 4-grams expose.
  • Temperature 0.9 gives the most fluent samples (45.8, 86.4% real words), at some cost in diversity.

An unedited excerpt from a sample at step 76,000 (training sampler: ancestral, 256 steps, temperature 1.0):

వంతెన, వేలం సీజన్ లో వారి నిర్ణయాలు తీసుకుంటుంటారు. వారు పర్యావరణ నిర్మాణం లేదని చేయాలి, ఇది ప్రైవేట్ గదులు, చాలా గది వాతావరణం ఒక సౌకర్యం ఆస్వాదించడం

sample metrics during training

Growth compatibility

The model is shaped so that larger models can be initialised from it function-preservingly: the grown model starts out computing exactly what this one does. Every stage keeps the head size (64), the RoPE base, the tokenizer and the 512-token canvas, and keeps the SwiGLU width at 8/3 × d_model; each width is a whole multiple of the previous one. Width grows by cloning (HyperCloning; Samragh et al., 2024), depth by inserting blocks whose output projections start at zero.

Preset d_model × layers Heads × 64 SwiGLU width Parameters
small-768 (this model) 768 × 12 12 2,048 120,048,384
grow-1536x16 1,536 × 16 24 4,096 523,224,576
grow-1536x32 1,536 × 32 24 4,096 976,258,560
grow-2304x16 2,304 × 16 36 6,144 1,124,575,488

Files in this repository

File What it is
model.safetensors The weights: EMA at step 76,000, float32, 74 tensors
config.json Architecture, tokenizer, diffusion and training settings
telugu_alldomain.tokenizer.json The tokenizer vocabulary
chandohasam_mdlm/model.py The denoiser (Denoiser, ModelConfig), as trained
chandohasam_mdlm/sampling.py The ancestral and confidence samplers, as evaluated
chandohasam_mdlm/tokenizer/ The akshara tokenizer, with aksharanusarika v0.0.7a (MIT) for segmentation
chandohasam_mdlm/__init__.py load, load_model, load_tokenizer, generate, invalid_bias
generate.py Command-line sampling
figures/ Figures of this card
LICENSE MIT

License

The weights and code in this repository are released under the MIT License. aksharanusarika.py is © 2025 Aksharanusarika Contributors, also MIT (chandohasam_mdlm/tokenizer/LICENSE.aksharanusarika). The training data has its own terms: Sangraha (CC BY 4.0), Telugu Wikipedia (CC BY-SA 3.0 and GFDL) and IndicCorp v2 (no license stated on its dataset card).

Citation

@misc{chandohasam_pilot120m_2026,
  title        = {Chandohasam pilot120m: a 120M-parameter masked diffusion language model for Telugu},
  author       = {Rallabandi, Samvaran Kashyap and Emani, Mahesh and Salopanthula, Radhe Shyam},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/samvaran/chandohasam}}
}

References

  • Sahoo, S. S., Arriola, M., Schiff, Y., Gokaslan, A., Marroquin, E., Chiu, J. T., Rush, A. and Kuleshov, V. (2024). Simple and Effective Masked Diffusion Language Models. Advances in Neural Information Processing Systems 37.
  • Ou, J., Nie, S., Xue, K., Zhu, F., Sun, J., Li, Z. and Li, C. (2025). Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data. International Conference on Learning Representations.
  • Zheng, K., Chen, Y., Mao, H., Liu, M.-Y., Zhu, J. and Zhang, Q. (2025). Masked Diffusion Models are Secretly Time-Agnostic Masked Models and Exploit Inaccurate Categorical Sampling. International Conference on Learning Representations.
  • Nie, S., Zhu, F., You, Z., Zhang, X., Ou, J., Hu, J., Zhou, J., Lin, Y., Wen, J.-R. and Li, C. (2025). Large Language Diffusion Models. Advances in Neural Information Processing Systems 38.
  • Samragh, M., Mirzadeh, I., Alizadeh Vahid, K., Faghri, F., Cho, M., Nabi, M., Naik, D. and Farajtabar, M. (2024). Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization. Proceedings of the 4th NeurIPS Efficient Natural Language and Speech Processing Workshop, PMLR 262.
  • Khan, M. S. U. R., Mehta, P., Sankar, A., Kumaravelan, U., Doddapaneni, S., B, S., G, V., Jain, S., Kunchukuttan, A., Kumar, P., Dabre, R. and Khapra, M. M. (2024). IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages. Proceedings of ACL 2024.
  • Doddapaneni, S., Aralikatte, R., Ramesh, G., Goyal, S., Khapra, M. M., Kunchukuttan, A. and Kumar, P. (2023). Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages. Proceedings of ACL 2023.
  • Gemma Team (2025). Gemma 3 Technical Report. arXiv:2503.19786.
Downloads last month
284
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train samvaran/chandohasam

Paper for samvaran/chandohasam