ViuMini-Dense-360M (Sovereign 32-Layer Dense Transformer)
ViuMini-Dense-360M is a sovereign Indian Small Language Model (SLM) featuring a 32-Layer Deep Dense Transformer architecture optimized for hierarchical multi-step reasoning, Indic multilingual fluency (Hindi, Hinglish, English), and high-speed on-device inference (mobile phones, laptops, and consumer GPUs).
Unlike sparse MoE models where routing instability and inactive parameters limit depth and reasoning density, ViuMini-Dense-360M dedicates 100% of its 366.6 Million parameters to every single token, delivering 2.2Γ higher active cognitive capacity per forward pass with rock-solid gradient stability.
ποΈ Architecture Specifications
| Specification | Value | Technical Rationale |
|---|---|---|
| Total Parameters | 366,617,536 (~366.6M) | Exact SmolLM-360M blueprint scaled for Indic |
| Active Parameters / Token | 366.6M (100% Active) | No routing overhead, no expert collapse |
Layers (n_layers) |
32 Deep Layers | Ultra-deep hierarchical reasoning abstraction |
Hidden Dimension (dim) |
960 | Optimal representation width for SLM |
Attention Heads (n_heads) |
15 Query Heads | Head dimension $d_{\text{head}} = 64$ ($15 \times 64 = 960$) |
KV Heads (n_kv_heads) |
5 KV Heads (GQA) | 3:1 Grouped-Query Attention saving 66% KV cache |
Feed-Forward Dimension (hidden_dim) |
2,624 | SwiGLU MLP ($h = 64 \times \lceil 2592/64 \rceil$, mult=2.7, multiple_of=64) |
| Vocabulary Size | 48,000 Tokens | Custom BPE optimized for Devanagari, Hinglish & Code |
| Embeddings | Weight-Tied | Input embeddings tied to LM output head (saves 46M params) |
| Attention Stability | Gemma 4 Pure QK-Norm | RMSNorm on Q and K; zero gradient saturation |
| Positional Encoding | Rotary RoPE ($\theta = 500,000$) | High-frequency positional precision, long-context ready |
Context Length (max_seq_len) |
2,048 Tokens | 3-Stage Curriculum: $512 \rightarrow 1,024 \rightarrow 2,048$ tokens |
| Optimizer | PagedAdamW8bit | Stable, decoupled weight decay, zero OOM spikes |
| Learning Rate | $2.5 \times 10^{-4}$ (Cosine) | 1,000 warmup steps, grad clip 1.0 |
π Why 32-Layer Dense beats 242M MoE
- 2.2Γ Higher Cognitive Capacity per Token:
- Old
ViuMini-MoE-242Monly activated ~166M parameters per token (with 76M parameters sitting idle). - ViuMini-Dense-360M activates all 366.6M parameters on every token.
- Old
- 32 Layers of Hierarchical Reasoning:
- 32 sequential layers provide deep multi-step deduction, grammar parsing, and factual recall compared to shallower configurations.
- Rock-Solid Training Stability:
- MoE router collapse (where routers get stuck picking the same 1β2 experts) and auxiliary loss fighting are completely eliminated.
- PagedAdamW8bit with $2.5 \times 10^{-4}$ learning rate and Gemma 4 QK-Norm guarantees clean loss convergence from the initial theoretical loss ($\ln(48,000) \approx 10.78$).
- Zero-Friction Deployment (1-Click GGUF / Ollama):
- Dense models run natively in
llama.cpp, Ollama, vLLM, MLC-LLM, and LM Studio without custom MoE routing kernels. - 4-bit quantized footprint is under 220 MB RAM, running at 100+ tokens/second directly on mobile devices and laptops.
- Dense models run natively in
π Pretraining Dataset
- Dataset Repo:
ViuAI/viu-mini-raw-pretrain - Total Volume: 126.5B+ authentic tokens across 2,109 verified Parquet files
- Domain Coverage:
- Hindi & Devanagari: Valmiki Ramayana, Mahabharata, 16 Mahapuranas, 4 Vedas, Hindi literature & news.
- Hinglish & Colloquial: Bollywood cinema screenplays, conversational dialogues, social chat.
- Reasoning & Code: DeepSeek-R1 math, Orca math, GSM8K, Python code & algorithms.
- Indian Governance, Law & Finance: Constitution of India, BNS/BNSS/BSA 2023 statutes, Supreme Court judgments, RBI circulars & financial news.
- Parallel Translation: AI4Bharat Samanantar (26 Million English $\leftrightarrow$ Hindi sentence pairs).
β‘ Kaggle Dual Tesla T4 Pretraining
Pretraining is fully automated and session-proof on Kaggle Dual Tesla T4 (2Γ16GB):
# Clone the repository
git clone https://huggingface.co/ViuAI/ViuMini-Dense-360M /kaggle/working/ViuMini-Dense-360M
cd /kaggle/working/ViuMini-Dense-360M
# Run pretraining using the dedicated Kaggle config
torchrun --nproc_per_node=2 model/scripts/train.py --config model/configs/train_kaggle_t4.yaml
- Curriculum Batching: Constant 65,536 tokens/step across Dual T4:
- Stage 1 (0β50k steps): seq 512, micro-batch 8 per GPU, accum 8 $\rightarrow$ $512 \times 8 \times 16 = 65,536$ tok/step
- Stage 2 (50kβ200k steps): seq 1024, micro-batch 4 per GPU, accum 8 $\rightarrow$ $1024 \times 4 \times 16 = 65,536$ tok/step
- Stage 3 (200k+ steps): seq 2048, micro-batch 2 per GPU, accum 8 $\rightarrow$ $2048 \times 2 \times 16 = 65,536$ tok/step
- Throughput & Cadence: ~8,500 β 10,000 tokens/sec across Dual T4.
- Every 1,000 steps $\approx$ 65.5 Million tokens ($\approx$ 2.14 hours), pushed directly to Hugging Face Hub.
- Full 126.5B dataset pretraining: $\approx$ 1,930,000 steps.
- VRAM Utilization: ~3.4 GB / 16 GB per T4 (12.6 GB safety headroom, zero OOM risk).
π Repository Structure
ViuMini-Dense-360M/
βββ model/
β βββ configs/
β β βββ viu1_dense_config.yaml # 32L Dense model architecture (366.6M params)
β β βββ train_kaggle_t4.yaml # Kaggle Dual T4 DDP training config
β β βββ train_config.yaml # Universal pretraining config
β β βββ config.json # Hugging Face compatible architecture config
β βββ scripts/
β βββ viu1_dense.py # ViuMini Transformer engine (32L Dense + QK-Norm + GQA)
β βββ train.py # Distributed DDP pretrainer with HF streaming & auto-push
β βββ push_to_hf.py # Hub sync utility
βββ tokenizer/
β βββ configs/
β βββ outputs/ # Custom 48,000 Indic BPE tokenizer
βββ docs/
β βββ KAGGLE_RUN.md # Single-cell copy-paste Kaggle guide
β βββ PROGRESS.md # Development milestones & audit logs
βββ README.md
π License
Apache-2.0. Sovereign AI initiative by ViuAI.