π MCR-Attention V2.0 β 29.5M
MCR-Attention V2.0 is a lightweight, custom-built language model featuring a Multi-Scale Context Router (MCR) attention mechanism. Trained on the TinyStories dataset, this architecture achieves constant-time $O(1)$ per-step generation speed regardless of sequence length, making it ultra-fast and memory-efficient for real-time streaming.
π Key Highlights & Metrics
| Metric | Value |
|---|---|
| Total Parameters | 29,479,825 (~29.5M) |
| Evaluation Loss | 0.9234 |
| Evaluation Perplexity | 2.52 |
| Final Training Loss | 0.9737 |
| Inference Speed (Peak) | ~403 tokens/sec (2.48 ms/token) |
| Inference Complexity | $O(1)$ per step constant time |
| Training Time | 127.5 minutes (~2.1 hrs) on 2x NVIDIA T4 |
ποΈ Architecture & Parameter Breakdown
The MCR-Attention mechanism uses an Attention Router that dynamically allocates compute across $K=4$ explicit memory scales: [32, 128, 512, 2048].
Model Configuration
d_model: 192d_state: 96num_scales: 4mlp_hidden: 384
Parameter Distribution
| Component | Parameters | Percentage |
|---|---|---|
| Embedding Layer | 9,649,344 | 32.7% |
| Memory Scales | 222,720 | 0.8% |
| Attention Router | 184,704 | 0.6% |
| Language Model Head | 19,423,057 | 65.9% |
| TOTAL | 29,479,825 | 100% |
β‘ Generation Speed Benchmark
Due to the fixed state-compression in the MCR memory blocks, generation latency remains constant as the sequence grows:
| Generated Tokens | Total Time (s) | Latency (ms/token) | Throughput (tok/s) | Complexity |
|---|---|---|---|---|
| 50 | 0.146 | 2.93 | 342 | $O(1)$ |
| 100 | 0.273 | 2.73 | 367 | $O(1)$ |
| 200 | 0.520 | 2.60 | 385 | $O(1)$ |
| 500 | 1.245 | 2.49 | 402 | $O(1)$ |
| 1000 | 2.484 | 2.48 | 403 | $O(1)$ |
| 2000 | 4.959 | 2.48 | 403 | $O(1)$ Constant |
Key Takeaway: As shown above, time per token asymptotically approaches 2.48 ms/token, confirming true $O(1)$ step-wise inference.
π Training Progress
- Dataset:
roneneldan/TinyStories - Hardware: 2x NVIDIA T4 GPUs (Kaggle)
- Precision:
fp16mixed precision - Optimizer Schedule: Warmup 500 steps $\rightarrow$ Cosine Decay down to
3e-05 - Total Training Time: 7,651 seconds (127.5 minutes)
π Sample Generation Output
Prompt: "One day, a brave"
One day, a brave little girl named Lily went to the park with her mommy. They saw a big tree with a shiny rock. She picked it up and started to eat it. "Can I have some of my doll?" asked her mom. "I'm looking at it, sweetie," said the fairy. "It's a beautiful flower. It's so pretty," replied the bird. Lily was very happy to see the butterfly and said, "Look at me! I'm so excited!" The man laughed and said, "That's great, Timmy! You are a good friend." The little boy was so happy. He learned that sometimes it's important to ask for help when he needed help.
Copyright (c) 2026 Gerson Fabian Buenahora Ormaza (BUEORM)
Licensed under the BUEORM MCR-Attention Community & Research License v1.0.
- Downloads last month
- 14
