ViuMini-Dense-360M (Sovereign 32-Layer Dense Transformer)

ViuMini-Dense-360M is a sovereign Indian Small Language Model (SLM) featuring a 32-Layer Deep Dense Transformer architecture optimized for hierarchical multi-step reasoning, Indic multilingual fluency (Hindi, Hinglish, English), and high-speed on-device inference (mobile phones, laptops, and consumer GPUs).

Unlike sparse MoE models where routing instability and inactive parameters limit depth and reasoning density, ViuMini-Dense-360M dedicates 100% of its 366.6 Million parameters to every single token, delivering 2.2Γ— higher active cognitive capacity per forward pass with rock-solid gradient stability.


πŸ›οΈ Architecture Specifications

Specification Value Technical Rationale
Total Parameters 366,617,536 (~366.6M) Exact SmolLM-360M blueprint scaled for Indic
Active Parameters / Token 366.6M (100% Active) No routing overhead, no expert collapse
Layers (n_layers) 32 Deep Layers Ultra-deep hierarchical reasoning abstraction
Hidden Dimension (dim) 960 Optimal representation width for SLM
Attention Heads (n_heads) 15 Query Heads Head dimension $d_{\text{head}} = 64$ ($15 \times 64 = 960$)
KV Heads (n_kv_heads) 5 KV Heads (GQA) 3:1 Grouped-Query Attention saving 66% KV cache
Feed-Forward Dimension (hidden_dim) 2,624 SwiGLU MLP ($h = 64 \times \lceil 2592/64 \rceil$, mult=2.7, multiple_of=64)
Vocabulary Size 48,000 Tokens Custom BPE optimized for Devanagari, Hinglish & Code
Embeddings Weight-Tied Input embeddings tied to LM output head (saves 46M params)
Attention Stability Gemma 4 Pure QK-Norm RMSNorm on Q and K; zero gradient saturation
Positional Encoding Rotary RoPE ($\theta = 500,000$) High-frequency positional precision, long-context ready
Context Length (max_seq_len) 2,048 Tokens 3-Stage Curriculum: $512 \rightarrow 1,024 \rightarrow 2,048$ tokens
Optimizer PagedAdamW8bit Stable, decoupled weight decay, zero OOM spikes
Learning Rate $2.5 \times 10^{-4}$ (Cosine) 1,000 warmup steps, grad clip 1.0

πŸš€ Why 32-Layer Dense beats 242M MoE

  1. 2.2Γ— Higher Cognitive Capacity per Token:
    • Old ViuMini-MoE-242M only activated ~166M parameters per token (with 76M parameters sitting idle).
    • ViuMini-Dense-360M activates all 366.6M parameters on every token.
  2. 32 Layers of Hierarchical Reasoning:
    • 32 sequential layers provide deep multi-step deduction, grammar parsing, and factual recall compared to shallower configurations.
  3. Rock-Solid Training Stability:
    • MoE router collapse (where routers get stuck picking the same 1–2 experts) and auxiliary loss fighting are completely eliminated.
    • PagedAdamW8bit with $2.5 \times 10^{-4}$ learning rate and Gemma 4 QK-Norm guarantees clean loss convergence from the initial theoretical loss ($\ln(48,000) \approx 10.78$).
  4. Zero-Friction Deployment (1-Click GGUF / Ollama):
    • Dense models run natively in llama.cpp, Ollama, vLLM, MLC-LLM, and LM Studio without custom MoE routing kernels.
    • 4-bit quantized footprint is under 220 MB RAM, running at 100+ tokens/second directly on mobile devices and laptops.

πŸ“š Pretraining Dataset

  • Dataset Repo: ViuAI/viu-mini-raw-pretrain
  • Total Volume: 126.5B+ authentic tokens across 2,109 verified Parquet files
  • Domain Coverage:
    • Hindi & Devanagari: Valmiki Ramayana, Mahabharata, 16 Mahapuranas, 4 Vedas, Hindi literature & news.
    • Hinglish & Colloquial: Bollywood cinema screenplays, conversational dialogues, social chat.
    • Reasoning & Code: DeepSeek-R1 math, Orca math, GSM8K, Python code & algorithms.
    • Indian Governance, Law & Finance: Constitution of India, BNS/BNSS/BSA 2023 statutes, Supreme Court judgments, RBI circulars & financial news.
    • Parallel Translation: AI4Bharat Samanantar (26 Million English $\leftrightarrow$ Hindi sentence pairs).

⚑ Kaggle Dual Tesla T4 Pretraining

Pretraining is fully automated and session-proof on Kaggle Dual Tesla T4 (2Γ—16GB):

# Clone the repository
git clone https://huggingface.co/ViuAI/ViuMini-Dense-360M /kaggle/working/ViuMini-Dense-360M
cd /kaggle/working/ViuMini-Dense-360M

# Run pretraining using the dedicated Kaggle config
torchrun --nproc_per_node=2 model/scripts/train.py --config model/configs/train_kaggle_t4.yaml
  • Curriculum Batching: Constant 65,536 tokens/step across Dual T4:
    • Stage 1 (0–50k steps): seq 512, micro-batch 8 per GPU, accum 8 $\rightarrow$ $512 \times 8 \times 16 = 65,536$ tok/step
    • Stage 2 (50k–200k steps): seq 1024, micro-batch 4 per GPU, accum 8 $\rightarrow$ $1024 \times 4 \times 16 = 65,536$ tok/step
    • Stage 3 (200k+ steps): seq 2048, micro-batch 2 per GPU, accum 8 $\rightarrow$ $2048 \times 2 \times 16 = 65,536$ tok/step
  • Throughput & Cadence: ~8,500 – 10,000 tokens/sec across Dual T4.
    • Every 1,000 steps $\approx$ 65.5 Million tokens ($\approx$ 2.14 hours), pushed directly to Hugging Face Hub.
    • Full 126.5B dataset pretraining: $\approx$ 1,930,000 steps.
  • VRAM Utilization: ~3.4 GB / 16 GB per T4 (12.6 GB safety headroom, zero OOM risk).

πŸ“ Repository Structure

ViuMini-Dense-360M/
β”œβ”€β”€ model/
β”‚   β”œβ”€β”€ configs/
β”‚   β”‚   β”œβ”€β”€ viu1_dense_config.yaml    # 32L Dense model architecture (366.6M params)
β”‚   β”‚   β”œβ”€β”€ train_kaggle_t4.yaml      # Kaggle Dual T4 DDP training config
β”‚   β”‚   β”œβ”€β”€ train_config.yaml         # Universal pretraining config
β”‚   β”‚   └── config.json               # Hugging Face compatible architecture config
β”‚   └── scripts/
β”‚       β”œβ”€β”€ viu1_dense.py             # ViuMini Transformer engine (32L Dense + QK-Norm + GQA)
β”‚       β”œβ”€β”€ train.py                  # Distributed DDP pretrainer with HF streaming & auto-push
β”‚       └── push_to_hf.py             # Hub sync utility
β”œβ”€β”€ tokenizer/
β”‚   β”œβ”€β”€ configs/
β”‚   └── outputs/                      # Custom 48,000 Indic BPE tokenizer
β”œβ”€β”€ docs/
β”‚   β”œβ”€β”€ KAGGLE_RUN.md                 # Single-cell copy-paste Kaggle guide
β”‚   └── PROGRESS.md                   # Development milestones & audit logs
└── README.md

πŸ“œ License

Apache-2.0. Sovereign AI initiative by ViuAI.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support