🪔 MANAS

Model for Awadhi Natural Autoregressive Sequences

A character-level causal Transformer trained on the literary corpus of Goswami Tulsidas

License: MIT PyTorch Language Parameters Status


Overview

MANAS is a small, experimental character-level causal Transformer language model trained on a Devanagari literary corpus focused on the works of Goswami Tulsidas — the 16th-century poet-saint and author of the Ramcharitmanas.

Unlike modern large language models that operate on subword tokens (BPE, SentencePiece), MANAS processes text one Unicode character at a time. The model was built entirely from scratch in PyTorch, without any pretrained weights or transfer learning, as an educational experiment into whether a small Transformer can learn the statistical patterns of classical Awadhi poetry from raw characters.

⚠️ This is Experimental Checkpoint 1. The dataset still contains OCR artifacts and some non-literary material. Do not use this as a finished, production Awadhi language model. A cleaned, page-verified corpus rebuild is planned for v2.


What's in This Repository

File Description
tulsidas_model.pth Trained PyTorch model state dictionary (~52MB)
README.md This file

The training dataset and inference code are not included in this public release.


Model Architecture

MANAS is a GPT-style decoder-only causal Transformer — the same fundamental architecture family as GPT-2. The key difference is that it works directly on Devanagari Unicode code points rather than BPE tokens.

Input Characters
      ↓
Character Embedding  (384-dim)
      +
Positional Embedding (256 positions)
      ↓
6 × Transformer Blocks
      ├── LayerNorm
      ├── Multi-Head Causal Self-Attention (6 heads × 64 dim)
      ├── Residual Connection
      ├── LayerNorm
      ├── Feed-Forward Network (384 → 1536 → 384)
      └── Residual Connection
      ↓
Final LayerNorm
      ↓
LM Head (Linear: 384 → 101)
      ↓
Logits over 101 Devanagari characters

Hyperparameters

Parameter Value
Embedding Dimension (n_embd) 384
Transformer Layers (n_layer) 6
Attention Heads (n_head) 6
Head Dimension 64
Context Window (block_size) 256 characters
Vocabulary Size 101 characters
Dropout 0.2
Total Parameters 10,816,613 (~10.9M)

Training

Dataset

The model was trained on a corpus of approximately 2.83 million Devanagari characters derived from OCR-processed editions of Tulsidas's works:

  • Ramcharitmanas (रामचरितमानस)
  • Vinay Patrika (विनय पत्रिका)
  • Kavitavali (कवितावली)
  • Parvatimangal (पार्वतीमंगल)
  • Other works from the Tulsi Granthavali compilation

The corpus contains zero Latin alphabet characters. The vocabulary consists of 101 unique Devanagari characters, punctuation marks, and verse/number markers common in classical Hindi poetry.

⚠️ Known Limitation: The current corpus is OCR-derived and may contain publisher front matter, Hindi commentary (टीका), table-of-contents material, and extraction errors alongside the literary verses. The corpus has not been manually page-verified.

Training Configuration

Config Value
Optimizer AdamW
Learning Rate 3e-4
Batch Size 64
Training Iterations 15,000
Train / Val Split 90% / 10%
Train Tokens 2,550,420
Validation Tokens 283,381
Hardware Single consumer GPU (NVIDIA CUDA)

Loss Curve

Step Train Loss Val Loss
0 4.85 4.85
500 2.29 2.61
2,000 1.62 2.15
5,000 1.35 2.05
10,000 1.16 2.07
15,000 1.03 2.11

The gap between training loss (1.03) and validation loss (2.11) indicates the model has overfit to the relatively small corpus — expected behavior for a 10.9M parameter model on ~2.8M characters.


How to Use

To run inference, you need to reconstruct the model architecture in PyTorch and also reconstruct the character vocabulary from the original training corpus (since the stoi/itos maps are derived from it).

import torch
import torch.nn as nn
from torch.nn import functional as F

# --- Hyperparameters (must match training) ---
block_size = 256
n_embd = 384
n_head = 6
n_layer = 6
dropout = 0.0  # Set to 0 for inference
vocab_size = 101  # Must match your character mapping
device = 'cuda' if torch.cuda.is_available() else 'cpu'

# --- Model Architecture ---
class Head(nn.Module):
    def __init__(self, head_size):
        super().__init__()
        self.key   = nn.Linear(n_embd, head_size, bias=False)
        self.query = nn.Linear(n_embd, head_size, bias=False)
        self.value = nn.Linear(n_embd, head_size, bias=False)
        self.register_buffer('tril', torch.tril(torch.ones(block_size, block_size)))
        self.dropout = nn.Dropout(dropout)
    def forward(self, x):
        B, T, C = x.shape
        k, q = self.key(x), self.query(x)
        wei = q @ k.transpose(-2, -1) * (k.shape[-1] ** -0.5)
        wei = wei.masked_fill(self.tril[:T, :T] == 0, float('-inf'))
        wei = F.softmax(wei, dim=-1)
        return self.dropout(wei) @ self.value(x)

class MultiHeadAttention(nn.Module):
    def __init__(self, num_heads, head_size):
        super().__init__()
        self.heads = nn.ModuleList([Head(head_size) for _ in range(num_heads)])
        self.proj  = nn.Linear(n_embd, n_embd)
        self.dropout = nn.Dropout(dropout)
    def forward(self, x):
        return self.dropout(self.proj(torch.cat([h(x) for h in self.heads], dim=-1)))

class FeedForward(nn.Module):
    def __init__(self, n_embd):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(n_embd, 4 * n_embd), nn.ReLU(),
            nn.Linear(4 * n_embd, n_embd), nn.Dropout(dropout),
        )
    def forward(self, x): return self.net(x)

class Block(nn.Module):
    def __init__(self, n_embd, n_head):
        super().__init__()
        head_size = n_embd // n_head
        self.sa, self.ffwd = MultiHeadAttention(n_head, head_size), FeedForward(n_embd)
        self.ln1, self.ln2 = nn.LayerNorm(n_embd), nn.LayerNorm(n_embd)
    def forward(self, x):
        return x + self.ffwd(self.ln2(x + self.sa(self.ln1(x))))

class LanguageModel(nn.Module):
    def __init__(self):
        super().__init__()
        self.token_embedding_table    = nn.Embedding(vocab_size, n_embd)
        self.position_embedding_table = nn.Embedding(block_size, n_embd)
        self.blocks  = nn.Sequential(*[Block(n_embd, n_head=n_head) for _ in range(n_layer)])
        self.ln_f    = nn.LayerNorm(n_embd)
        self.lm_head = nn.Linear(n_embd, vocab_size)
    def forward(self, idx, targets=None):
        B, T = idx.shape
        x = self.token_embedding_table(idx) + self.position_embedding_table(torch.arange(T, device=device))
        logits = self.lm_head(self.ln_f(self.blocks(x)))
        return logits, None
    def generate(self, idx, max_new_tokens):
        for _ in range(max_new_tokens):
            logits, _ = self(idx[:, -block_size:])
            idx_next = torch.multinomial(F.softmax(logits[:, -1, :], dim=-1), num_samples=1)
            idx = torch.cat((idx, idx_next), dim=1)
        return idx

# --- Load weights ---
model = LanguageModel()
model.load_state_dict(torch.load('tulsidas_model.pth', map_location=device))
model.to(device)
model.eval()

# --- Build vocab from your corpus (required for encoding/decoding) ---
# with open('your_training_corpus.txt', 'r', encoding='utf-8') as f:
#     text = f.read()
# chars = sorted(list(set(text)))
# stoi = {ch: i for i, ch in enumerate(chars)}
# itos = {i: ch for i, ch in enumerate(chars)}
# encode = lambda s: [stoi[c] for c in s]
# decode = lambda l: ''.join([itos[i] for i in l])

# --- Inference ---
# prompt = "श्री राम"
# context = torch.tensor([encode(prompt)], dtype=torch.long, device=device)
# print(decode(model.generate(context, max_new_tokens=300)[0].tolist()))

Sample Output

Given the prompt श्री, the model produced (at step 14999):

श्रीरामचन्द्रजीकी परिश्रामचन्द्रजीने ही
सब माता आदि मिट ढीं । विभीषणजीने उसको हृदयमें उठा लिया
निदान दीन बचन गहि सोभा बढ़ावा। बालि और बिपुल बोलावड़े गावा ॥
मोरे आधीस मैं भी जान । ता कुन्ठ सद्य सुरुचि रसखावा ॥

Limitations

Limitation Detail
Small model 10.9M parameters is far below modern LLM scale. Outputs are statistically plausible Devanagari, not semantically coherent poetry.
Corpus noise OCR errors, Hindi commentary (टीका), and publisher front matter remain in the training data.
No factual knowledge Cannot answer questions. Only predicts next characters.
Short context 256-character window limits long-range coherence.
Overfitting Train loss (1.03) vs. val loss (2.11) gap indicates memorisation of the small corpus.
Devanagari-only Vocabulary has zero Latin characters. English input must be transliterated before encoding.

Roadmap

  • v2 Dataset: Manually page-verified OCR rebuild separating verse from prose commentary
  • v2 Model: Retrain on the cleaned corpus with a larger context window
  • Evaluation: Implement character-level perplexity benchmarks on a held-out verse set
  • Tokenizer: Experiment with syllable-level tokenization for better Hindi morpheme coverage

Citation

If you reference this model in academic work:

@misc{manas2026,
  author       = {JayF14},
  title        = {MANAS: Model for Awadhi Natural Autoregressive Sequences},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/JayF14/MANAS}},
  note         = {Experimental Checkpoint 1. Character-level causal Transformer trained on Tulsidas literary corpus.}
}

License

MIT License. See LICENSE for details.


जेहि पर कृपा राम कै होई। ता पर कृपा करहिं सब कोई॥
— Ramcharitmanas, Goswami Tulsidas
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support