Chandohasam pilot120m: a 120M masked diffusion language model for Telugu
pilot120m is a 120M-parameter masked diffusion language model (MDLM; Sahoo et al., 2024) for
Telugu, trained from scratch on 4.98B tokens of web, news and encyclopedic text. Instead of writing
left to right, it fills a canvas of 512 <mask> tokens in any order, seeing the whole canvas at every
step. Its tokenizer gives one token per akshara (orthographic syllable), so every token is a unit of
Telugu metre.
It is stage 1 of Chandohasam, a project on generating Telugu metrical poetry (padyams), whose metres fix the weight of every syllable, a rhyme in the second syllable of every line (prāsa) and a caesura agreement (yati). This stage teaches general Telugu; a later stage specialises the model to verse, decoded under exact metrical constraints. It has seen no poetry: documents that overlap a large collection of known Telugu poems were removed from its training data.
| Model type | Masked (absorbing-state) discrete diffusion LM; bidirectional transformer denoiser, no timestep input |
| Language | Telugu (te) |
| Parameters | 120,048,384 (84,953,856 outside the tied embedding) |
| Canvas | 512 tokens, about 110 words of prose |
| Tokenizer | 45,591 tokens: one per akshara, byte-level BPE fallback, no out-of-vocabulary input |
| Training | 76,000 steps × 65,536 tokens = 4.98B tokens, on one 8 GB laptop GPU in 2 days 19 hours |
| Weights | The EMA (decay 0.9999) at step 76,000, float32 (model.safetensors) |
| Held-out test NELBO | 1.783 nats per token, weighted like the training mix |
| License | MIT (weights and code) |
Contents: Quick start · Intended use and limitations · Architecture · Tokenizer · Diffusion objective · Sampling · Training data · Training procedure · Evaluation · Growth compatibility · Files · License · Citation · References
Quick start
The model is a custom architecture; the code to load and sample it is in this repository
(chandohasam_mdlm/). It needs Python ≥ 3.10, PyTorch ≥ 2.4 and safetensors.
pip install torch safetensors huggingface_hub
import sys
from huggingface_hub import snapshot_download
repo = snapshot_download("samvaran/chandohasam")
sys.path.insert(0, repo)
from chandohasam_mdlm import load, generate
model, tok = load(repo, device="cuda") # or device="cpu"
for text in generate(model, tok, n=2, steps=512, temperature=0.9, seed=0):
print(text, "\n---")
# fix the start of every canvas
print(generate(model, tok, n=1, prefix="తెలుగు భాష", seed=1)[0])
Or from a clone of the repository: python generate.py --n 4 --steps 512 --temperature 0.9.
generate starts from 512 <mask> tokens and unmasks them with the ancestral sampler. 512 steps at
temperature 0.9 gave the most fluent samples in the evaluation (Sampling). On the training
GPU (RTX 5050 Laptop), 64 canvases at 512 steps take about 12 minutes.
The denoiser can also be used directly: model(x) maps token ids x (batch × ≤ 512, with
tok.special_to_id["<mask>"] at the positions to predict) to logits over the padded vocabulary of 45,696.
Before taking a softmax, add invalid_bias(...) from chandohasam_mdlm, which sets <mask>, <pad>,
<unk> and the 105 padding rows to −∞, as in training.
Intended use and limitations
Intended use. Research on masked diffusion language models for Telugu, and a starting point for continued pretraining or fine-tuning, in particular for constrained generation where each position is an akshara: metrical verse, infilling, and decoding steered by a rule engine.
It is a base model. It is not instruction-tuned, chat-tuned or safety-tuned, and it is not a source of facts. Its unconditional samples are fluent phrases in a web, news or blog register without a coherent thread across sentences. At the recommended setting, an independent judge (Gemma-3-1B) rates them at a perplexity of 45.8, against 16.9 for real Telugu text, and 86.4% of their words are real words, against 88.7% in real text.
Known limitations.
- The heavily masked regime is weak. With 85% of the tokens masked, the model predicts 31–32% of them correctly. Generation starts from a fully masked canvas, so its first steps run in this regime, and more denoising steps barely help (Evaluation).
- Confidence-ordered unmasking collapses into repeated tokens for unconditional generation. Use the ancestral sampler.
- The data is web and news text. Generated text can mention real people, places and events and say false things about them, and it can reproduce biases and offensive language present in web text. Nothing beyond the cleaning described below was done to filter content.
- Telugu only. Other scripts are kept losslessly as bytes by the tokenizer but are not modelled.
- No poetry. Poems were deliberately removed from the training data, so the model knows the language of verse only as far as it appears in prose.
- The canvas is fixed at 512 tokens. Training windows are 512 tokens cut from concatenated documents,
so a canvas may contain the end of one document and the start of the next, separated by
<eos><bos>.
Architecture
Hyperparameters
| Field | Value | Meaning |
|---|---|---|
vocab_size |
45,591 | Tokenizer vocabulary: 5 specials, 256 bytes, 485 BPE pieces, 10 Telugu digits and 44,835 whole aksharas. Padded to 45,696 (a multiple of 128) for the GPU; the 105 padding rows are never predicted |
d_model |
768 | Width of the residual stream, the token embeddings and every block's input and output |
n_layers |
12 | Transformer blocks |
n_heads |
12 | Attention heads per block; head size 768 / 12 = 64 |
mlp_hidden |
2,048 | SwiGLU hidden width per block (8/3 × d_model) |
max_len |
512 | Canvas length in tokens; the rotary tables are built for 512 positions |
rope_base |
10,000 | Base frequency of the rotary position embeddings |
norm_eps |
1e-6 | RMSNorm epsilon |
init_std |
0.02 | Standard deviation of the normal initialisation of embeddings and linear layers; the two residual output projections of each block use 0.02 / √(2 × 12) ≈ 0.00408 |
Design choices
| Choice | Setting |
|---|---|
| Block | Pre-norm: x + Attn(RMSNorm(x)), then x + SwiGLU(RMSNorm(x)); a final RMSNorm before the output layer |
| Attention | Bidirectional (no causal mask); torch.nn.functional.scaled_dot_product_attention, softmax scale 1/√64; one fused query/key/value projection |
| Positions | Rotary embeddings (RoPE) on queries and keys, rotating interleaved channel pairs; no learned position embedding |
| MLP | SwiGLU: down(SiLU(gate(x)) ⊙ up(x)), with the gate and up projections fused |
| Biases, dropout | None |
| Time conditioning | None. Under masking noise the optimal denoiser does not depend on t (Ou et al., 2025; Zheng et al., 2025), and LLaDA (Nie et al., 2025) also drops it |
| Embeddings | Input and output tied: one matrix E, and logits = h · Eᵀ |
| Precision in training | float32 master weights and optimizer state, bf16 autocast for the forward and backward passes, float32 cross-entropy |
Weight tensors
All 74 tensors (i = 0…11):
| Tensor | Shape | Parameters | Initialisation | Weight decay | Role |
|---|---|---|---|---|---|
embed.weight |
45,696 × 768 | 35,094,528 | N(0, 0.02) | 0.1 | Token embedding; tied, so also the output layer |
blocks.i.norm1.weight |
768 | 768 | ones | 0 | RMSNorm gain before attention |
blocks.i.qkv.weight |
2,304 × 768 | 1,769,472 | N(0, 0.02) | 0.1 | Fused query/key/value projection |
blocks.i.proj.weight |
768 × 768 | 589,824 | N(0, 0.00408) | 0.1 | Attention output projection (residual branch) |
blocks.i.norm2.weight |
768 | 768 | ones | 0 | RMSNorm gain before the MLP |
blocks.i.gate_up.weight |
4,096 × 768 | 3,145,728 | N(0, 0.02) | 0.1 | Fused SwiGLU gate and up projections |
blocks.i.down.weight |
768 × 2,048 | 1,572,864 | N(0, 0.00408) | 0.1 | SwiGLU down projection (residual branch) |
norm.weight |
768 | 768 | ones | 0 | Final RMSNorm gain |
| Total | 120,048,384 |
Each block holds 7,079,424 parameters, so the 12 blocks hold 84,953,088; the final norm (768) and the tied embedding (35,094,528) make up the rest.
Tokenizer
- Vocabulary.
telugu_alldomain.tokenizer.json(sha256200c814d84cc1b88bccdea3e4131a4c2fee296d050587e5d3da99f22d73d30fc), 45,591 tokens: 44,835 whole aksharas, 485 byte-level BPE pieces, 256 bytes, 10 Telugu digits and 5 specials:<pad>= 0,<unk>= 1,<bos>= 2,<eos>= 3,<mask>= 4. - One token per akshara. Text is split into aksharas by aksharanusarika (included, MIT), the same splitter as the Chandohasam metre engine, so token boundaries are metrical boundaries. The arasunna ఁ stays on its akshara and zero-width non-joiners are removed.
- No out-of-vocabulary input. An akshara without its own token falls back to byte-level BPE. On unseen Sangraha news, 99.94% of syllables are single tokens.
- Lossless.
decode(encode(text))returns the normalised text.
Diffusion objective
| Item | Setting |
|---|---|
| Forward (noising) process | Each token is replaced by <mask> independently with probability t: the absorbing-state process with a log-linear schedule, α_t = 1 − t |
| Time sampling | t ~ U(ε, 1) with ε = 0.001, spread evenly across the batch: one random offset u, and t_i = ε + (1 − ε)·((u + i/B) mod 1) |
| Parameterisation | SUBS: the logits of <mask>, <pad>, <unk> and the padding rows are −∞, and unmasked tokens are copied through, with no loss on them |
| Loss | NELBO per token: the sum over masked positions of cross-entropy / t, divided by batch × length. In nats per token; exp(NELBO) bounds the perplexity |
| Memory | The output layer runs only at masked positions, in chunks of 2,048 whose logits are recomputed in the backward pass, so the 45,696-wide output layer fits in 8 GB |
Sampling
- Ancestral sampler (
chandohasam_mdlm/sampling.py), MDLM's: going from t to s < t, each still-masked position is revealed with probability (t − s)/t, its token drawn from the model. Because the model has no time input, a step that reveals nothing reuses the previous forward pass. - Confidence order (MaskGIT/LLaDA style) is also implemented: each step reveals the masked positions whose sampled token is most probable. For unconditional generation it collapses into repetition.
- Categorical draws use float64 Gumbel noise. Float32 noise quietly lowers the sampling temperature and flatters sample quality (Zheng et al., 2025).
- Recommended: ancestral, 512 steps, temperature 0.9.
Training data
Sources
| Source | Hugging Face dataset | Files (pinned revision) | License | Download |
|---|---|---|---|---|
| Sangraha, verified split, Telugu (Khan et al., 2024) | ai4bharat/sangraha | verified/tel/*.parquet @ 8b813c3f |
CC BY 4.0 | 15.1 GB |
| IndicCorp v2, Telugu (Doddapaneni et al., 2023) | ai4bharat/IndicCorpV2 | data/te.txt @ 2d7285e6 |
not stated on the dataset card | 15.8 GB |
| Telugu Wikipedia, dump of 2023-11-01 | wikimedia/wikipedia | 20231101.te/*.parquet @ b04c8d1c |
CC BY-SA 3.0, GFDL | 0.2 GB |
Cleaning, deduplication and decontamination
- Normalisation. NFC, the tokenizer's normalisation, control characters removed, whitespace collapsed,
literal
\nsequences turned into newlines. - Exact duplicates and boilerplate. Exact-duplicate documents are dropped, and so are boilerplate lines (any line found in 10 or more distinct documents; 90,446 such lines). IndicCorp paragraphs that already occur inside a Wikipedia or Sangraha document (5,518,447) are dropped.
- Cleaning. URLs and e-mail addresses are cut out of their lines. A line is dropped if under 50% of its letters are Telugu. A document is dropped if it has fewer than 40 Telugu letters, if under 80% of its letters are Telugu, or if more than 30% of its lines are repeats. After tokenization, a document is dropped if more than 2% of its Telugu pieces have no whole-akshara token (garbled text), or if under half of its tokens carry text rather than whitespace, ASCII punctuation or digits (tables and number lists).
- Near duplicates. MinHash (64 permutations over 8-token shingles), clustered at an estimated Jaccard similarity ≥ 0.8 (8 bands × 8 rows), one document kept per cluster; for Wikipedia and Sangraha.
- Decontamination. Any document that shares a 20-token window with a large collection of known Telugu poems is dropped, so that the later poetry stage is not contaminated.
- Splits. Documents go to train, validation (0.5%) and test (0.5%) by a stable hash. Each document is
stored as
<bos> … <eos>.
Documents at each stage:
| Source | Raw | Duplicate | Cleaning rejects | Token-level rejects | Near duplicate | Poem overlap | Kept |
|---|---|---|---|---|---|---|---|
| Wikipedia | 87,854 | 100 | 1,950 | 561 | 4,963 | 205 | 80,075 |
| Sangraha | 7,081,734 | 77 | 77,947 | 6,977 | 115,540 | 68,888 | 6,812,305 |
| IndicCorp (paragraphs) | 21,458,261 | 7,546,348 | 5,642,341 | 56,809 | 0 | 573 | 8,212,190 |
For IndicCorp, "duplicate" includes the paragraphs already present in another source, and most cleaning rejects are paragraphs too short to keep.
Tokens, including <bos> and <eos>:
| Source | Train | Validation | Test |
|---|---|---|---|
| Wikipedia | 102,634,991 | 560,104 | 497,202 |
| Sangraha | 7,327,831,173 | 38,164,489 | 37,278,696 |
| IndicCorp | 1,981,820,075 | 9,970,079 | 10,195,016 |
| Total | 9,412,286,239 | 48,694,672 | 47,970,914 |
Mixture
| Source | Sampling weight | Tokens seen | Passes over its train split |
|---|---|---|---|
| Sangraha | 75% | 3.74B | 0.51 |
| IndicCorp | 20% | 1.00B | 0.50 |
| Wikipedia | 5% | 0.25B | 2.43 |
Each training example is 512 consecutive tokens from a random offset in the chosen source's train split, so a window may span several documents (MDLM's "wrapped" setting).
Training procedure
| Setting | Value |
|---|---|
| Optimizer | AdamW (fused), β₁ = 0.9, β₂ = 0.98, ε = 1e-8 |
| Learning rate | Linear warm-up over 2,000 steps to 3e-4, then cosine decay to 3e-5 at step 76,000 |
| Weight decay | 0.1 on all matrices including the tied embedding; none on RMSNorm gains |
| Gradient clipping | 1.0 (global norm) |
| Batch | 128 sequences × 512 tokens = 65,536 tokens per step, as 16 accumulated micro-batches of 8 |
| Steps, tokens | 76,000 steps, 4,980,736,000 tokens: 41 tokens per parameter |
| EMA | Decay 0.9999, warmed up as min(0.9999, (1 + step)/(10 + step)); the EMA weights are the ones released and evaluated |
| Seed | 0; batch composition, mask rates and masks are a function of (seed, step) |
| Compilation | torch.compile on the transformer; the output layer runs eagerly |
| Hardware | One NVIDIA GeForce RTX 5050 Laptop GPU (8 GB), 16 CPU threads, 23 GB RAM |
| Software | Python 3.12, PyTorch 2.14.0 with CUDA 13.0, TF32 matmuls |
| Throughput | Median 20.8k tokens/s (3.15 s per step), about 1.8B tokens per day; 5.7 GB of GPU memory, 6.7 GB peak |
| Wall clock | 2 days 19 hours (26 to 29 September 2026) |
The run never diverged, never ran out of memory and never skipped an update. It rode through two mains-power cuts on battery without losing work.
Evaluation
Likelihood
NELBO in nats per token (lower is better) of the released weights, on 512 fixed canvases per source with fixed mask rates and masks. exp(NELBO) is an upper bound on the per-token perplexity.
| Split | Sangraha | IndicCorp | Wikipedia | Mean of the three | Weighted like the training mix |
|---|---|---|---|---|---|
| Validation | 1.899 | 1.719 | 0.994 | 1.537 | 1.818 |
| Test | 1.846 | 1.729 | 1.053 | 1.543 | 1.783 |
The test split was untouched until this evaluation and is no harder than validation. Wikipedia scores low because, beyond its first 128 canvases, its held-out text is dominated by templated village articles, which are nearly predictable; the per-source and mix-weighted values are the ones to compare.
During training, the validation NELBO on 128 fixed canvases per source fell from 2.921 at step 1,000 to 1.747 at step 76,000, and was still falling slowly at the end.
Masked-token accuracy
Top-1 accuracy on masked tokens (test split): 87.9% with 15% of the tokens masked, 69.1% with 50% and 31.4% with 85% (validation: 89.4%, 68.8%, 32.3%).
Sample quality and sampler settings
64 canvases of 512 tokens per setting, generated from an all-mask canvas with the released weights. Judge perplexity is the perplexity of the samples under an independent model, google/gemma-3-1b-pt (Gemma Team, 2025); on real validation text it is 16.9, and 64.9 when the words of that text are shuffled. Real words: the share of generated words that occur in validation text (88.7% for real text).
| Unmasking order | Steps | Temperature | Judge perplexity | Real words | Token entropy | Distinct-1 | Repeated 4-grams |
|---|---|---|---|---|---|---|---|
| ancestral | 256 | 1.0 | 71.5 | 77.5% | 4.35 | 0.734 | 0.0% |
| ancestral | 512 | 1.0 | 65.8 | 77.8% | 4.32 | 0.729 | 0.0% |
| ancestral | 1,024 | 1.0 | 64.9 | 78.8% | 4.33 | 0.725 | 0.1% |
| confidence | 256 | 1.0 | 1.7 | 99.7% | 0.75 | 0.006 | 94.0% |
| confidence | 512 | 1.0 | 1.4 | 99.7% | 0.72 | 0.004 | 95.9% |
| ancestral | 512 | 0.9 | 45.8 | 86.4% | 4.18 | 0.629 | 0.0% |
| real Telugu text | 16.9 | 88.7% | 4.32 | 0.73 |
- The model, not the sampler, is the limit. Four times as many steps (256 → 1,024) bring judge perplexity only from 71.5 to 64.9, the level of word-shuffled real text.
- Confidence order collapses. It fills the canvas with a few repeated tokens; the judge scores that as nearly perfect, which token entropy and repeated 4-grams expose.
- Temperature 0.9 gives the most fluent samples (45.8, 86.4% real words), at some cost in diversity.
An unedited excerpt from a sample at step 76,000 (training sampler: ancestral, 256 steps, temperature 1.0):
వంతెన, వేలం సీజన్ లో వారి నిర్ణయాలు తీసుకుంటుంటారు. వారు పర్యావరణ నిర్మాణం లేదని చేయాలి, ఇది ప్రైవేట్ గదులు, చాలా గది వాతావరణం ఒక సౌకర్యం ఆస్వాదించడం
Growth compatibility
The model is shaped so that larger models can be initialised from it function-preservingly: the grown model starts out computing exactly what this one does. Every stage keeps the head size (64), the RoPE base, the tokenizer and the 512-token canvas, and keeps the SwiGLU width at 8/3 × d_model; each width is a whole multiple of the previous one. Width grows by cloning (HyperCloning; Samragh et al., 2024), depth by inserting blocks whose output projections start at zero.
| Preset | d_model × layers | Heads × 64 | SwiGLU width | Parameters |
|---|---|---|---|---|
small-768 (this model) |
768 × 12 | 12 | 2,048 | 120,048,384 |
grow-1536x16 |
1,536 × 16 | 24 | 4,096 | 523,224,576 |
grow-1536x32 |
1,536 × 32 | 24 | 4,096 | 976,258,560 |
grow-2304x16 |
2,304 × 16 | 36 | 6,144 | 1,124,575,488 |
Files in this repository
| File | What it is |
|---|---|
model.safetensors |
The weights: EMA at step 76,000, float32, 74 tensors |
config.json |
Architecture, tokenizer, diffusion and training settings |
telugu_alldomain.tokenizer.json |
The tokenizer vocabulary |
chandohasam_mdlm/model.py |
The denoiser (Denoiser, ModelConfig), as trained |
chandohasam_mdlm/sampling.py |
The ancestral and confidence samplers, as evaluated |
chandohasam_mdlm/tokenizer/ |
The akshara tokenizer, with aksharanusarika v0.0.7a (MIT) for segmentation |
chandohasam_mdlm/__init__.py |
load, load_model, load_tokenizer, generate, invalid_bias |
generate.py |
Command-line sampling |
figures/ |
Figures of this card |
LICENSE |
MIT |
License
The weights and code in this repository are released under the MIT License. aksharanusarika.py is
© 2025 Aksharanusarika Contributors, also MIT (chandohasam_mdlm/tokenizer/LICENSE.aksharanusarika).
The training data has its own terms: Sangraha (CC BY 4.0), Telugu Wikipedia (CC BY-SA 3.0 and GFDL) and
IndicCorp v2 (no license stated on its dataset card).
Citation
@misc{chandohasam_pilot120m_2026,
title = {Chandohasam pilot120m: a 120M-parameter masked diffusion language model for Telugu},
author = {Rallabandi, Samvaran Kashyap and Emani, Mahesh and Salopanthula, Radhe Shyam},
year = {2026},
howpublished = {\url{https://huggingface.co/samvaran/chandohasam}}
}
References
- Sahoo, S. S., Arriola, M., Schiff, Y., Gokaslan, A., Marroquin, E., Chiu, J. T., Rush, A. and Kuleshov, V. (2024). Simple and Effective Masked Diffusion Language Models. Advances in Neural Information Processing Systems 37.
- Ou, J., Nie, S., Xue, K., Zhu, F., Sun, J., Li, Z. and Li, C. (2025). Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data. International Conference on Learning Representations.
- Zheng, K., Chen, Y., Mao, H., Liu, M.-Y., Zhu, J. and Zhang, Q. (2025). Masked Diffusion Models are Secretly Time-Agnostic Masked Models and Exploit Inaccurate Categorical Sampling. International Conference on Learning Representations.
- Nie, S., Zhu, F., You, Z., Zhang, X., Ou, J., Hu, J., Zhou, J., Lin, Y., Wen, J.-R. and Li, C. (2025). Large Language Diffusion Models. Advances in Neural Information Processing Systems 38.
- Samragh, M., Mirzadeh, I., Alizadeh Vahid, K., Faghri, F., Cho, M., Nabi, M., Naik, D. and Farajtabar, M. (2024). Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization. Proceedings of the 4th NeurIPS Efficient Natural Language and Speech Processing Workshop, PMLR 262.
- Khan, M. S. U. R., Mehta, P., Sankar, A., Kumaravelan, U., Doddapaneni, S., B, S., G, V., Jain, S., Kunchukuttan, A., Kumar, P., Dabre, R. and Khapra, M. M. (2024). IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages. Proceedings of ACL 2024.
- Doddapaneni, S., Aralikatte, R., Ramesh, G., Goyal, S., Khapra, M. M., Kunchukuttan, A. and Kumar, P. (2023). Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages. Proceedings of ACL 2023.
- Gemma Team (2025). Gemma 3 Technical Report. arXiv:2503.19786.
- Downloads last month
- 284





