- ARTN - Adaptive Ternary Recurrent Network
ARTN - Adaptive Ternary Recurrent Network
1.58-bit Ternary Quantization for Language Models
Experimental implementation of ternary (1.58-bit) quantization for training small language models on extremely limited hardware.
Status: EXPERIMENTAL - Not Production Ready
Project Overview
This project explores whether ternary quantization (weights constrained to -1, 0, +1) can be used to train small language models on hardware with:
| Constraint | Value |
|---|---|
| RAM | 4GB |
| CPU | 2-core (no GPU) |
| Storage | ~20GB |
Model Architecture
| Parameter | Value |
|---|---|
| Model type | BitLinearGPT |
| Parameters | 6,856,608 (~6.8M) |
| Hidden size (d_model) | 256 |
| Layers | 6 |
| Attention heads | 8 |
| Feed-forward size | 1024 |
| Max sequence length | 256 |
| Vocabulary size | 4,000 |
| Quantization | Ternary (1.58-bit) |
Checkpoints Included
1. checkpoint-step5000-experiment.pt
| Attribute | Value |
|---|---|
| Steps | 5,000 |
| Final Loss | 8.45 |
| Entropy | Not monitored |
| Status | SUCCESS (pipeline verified) |
| Description | First test training, proved architecture works |
| Duration | ~4 hours |
Purpose: Verified that training pipeline works correctly.
2. checkpoint-step50000-failed-collapsed.pt
| Attribute | Value |
|---|---|
| Steps | 50,000 |
| Final Loss | 3.16 |
| Final Entropy | 8.2% (COLLAPSED) |
| Top Token | "time" at 94% |
| Status | FAILED (collapsed) |
| Description | Beta1=None caused model collapse |
| Duration | 79 hours |
| Root Cause | Missing momentum (beta1=None) |
Warning: This checkpoint is COLLAPSED. Model outputs repetitive text like "time time time...". Use only as a learning example of what NOT to do.
3. checkpoint-step50000-experiment-gibberish.pt
| Attribute | Value |
|---|---|
| Steps | 50,000 |
| Final Loss | 7.56 |
| Final Entropy | 96.2% (HEALTHY) |
| Top Token | "the" at 0.56% |
| Status | HEALTHY but NOT USEFUL |
| Description | Beta1=0.9 prevented collapse but learning too slow |
| Duration | 83 hours |
| Issue | Beta1=0.9 too conservative |
Note: This checkpoint is HEALTHY (entropy normal) but produces gibberish. Model learned slowly due to overly conservative optimizer settings.
Training Configuration
| Setting | Value |
|---|---|
| Dataset | TinyStories (roneneldan/TinyStories) |
| Samples | 50,000 |
| Batch size | 4 |
| Gradient accumulation | 8 |
| Effective batch size | 32 |
| Optimizer | Adafactor |
| Max Learning Rate | 0.0003 |
| Min Learning Rate | 0.00003 |
| Warmup steps | 1,000 |
| Weight decay | 0.01 |
| Gradient clipping | 1.0 |
Key Findings
Beta1 Trade-off
| Beta1 Value | Learning Speed | Stability | Result |
|---|---|---|---|
| None (0.0) | FAST | UNSTABLE | COLLAPSED at 50k |
| 0.9 | SLOW | STABLE | HEALTHY but not useful |
| 0.5 - 0.7 | UNKNOWN | UNKNOWN | Need to test |
Critical Lesson: Beta1=None causes collapse. Beta1=0.9 too conservative. Need to find optimal value.
Training Metrics Comparison
| Training | Loss @ 50k | Entropy | Top Token | Status |
|---|---|---|---|---|
| Attempt 2 (beta1=None) | 3.16 | 8.2% | 94% "time" | COLLAPSED |
| Attempt 4 (beta1=0.9) | 7.56 | 96.2% | 0.56% "the" | HEALTHY (gibberish) |
Paradox: Lower loss does NOT mean better model. Attempt 2 had lower loss but was collapsed.
Total Time Invested
| Attempt | Duration | Outcome |
|---|---|---|
| 1 (5k test) | 4 hours | SUCCESS (pipeline verified) |
| 2 (50k collapsed) | 79 hours | FAILED (beta1=None) |
| 3 (stopped) | 1 hour | STOPPED (wrong accumulation) |
| 4 (50k gibberish) | 83 hours | COMPLETED (not useful) |
| Total | 167 hours |
Lessons Learned
- Entropy monitoring is CRITICAL - Loss alone doesn't indicate model health
- Beta1=None causes collapse - Model falls into repetitive patterns
- Beta1=0.9 is too conservative - Learning becomes extremely slow
- Lower loss ≠ better model - Collapsed model had lower loss (3.16 vs 7.56)
- Need optimal beta1 - Testing 0.5 and 0.7 as candidates
- Ternary quantization is challenging - Requires careful optimizer tuning
How to Use These Checkpoints
For Research/Learning
import torch
from src.model.bitlinear import BitLinearGPT
from src.tokenizer.bpe_tokenizer import BPETokenizer
# Load model architecture
model = BitLinearGPT(
vocab_size=4000,
d_model=256,
n_heads=8,
n_layers=6,
d_ff=1024,
max_seq_len=256
)
# Load checkpoint (choose one)
checkpoint = torch.load('checkpoint-step50000-experiment-gibberish.pt', map_location='cpu')
model.load_state_dict(checkpoint)
model.eval()
# Load tokenizer
tokenizer = BPETokenizer.load('tokenizer.json')
# Generate text
input_ids = torch.tensor([tokenizer.encode("Once upon a time")])
# ... generation code ...
Not For Production
These checkpoints are experimental. They will NOT produce coherent text. Use them to:
- Learn about ternary quantization challenges
- Understand model collapse patterns
- Compare optimizer settings
- Train your own models with different configs
Tokenizers
| File | Description |
|---|---|
| tokenizer.json | Main tokenizer (BPE, vocab size 4000) |
| tokenizer-backup-50k.json | Backup from collapsed training |
Source Code
Full source code available at GitHub:
https://github.com/kishxrx/ARTN
Includes:
- Complete training scripts
- Model architecture
- Training history documentation
- Monitoring tools
Training History
Complete documentation of all experiments:
https://github.com/kishxrx/ARTN/blob/main/TRAINING_HISTORY.md
Next Steps
| Phase | Plan |
|---|---|
| 1 | Test beta1=0.5 (10k steps) |
| 2 | Test beta1=0.7 (10k steps) |
| 3 | Compare entropy, loss, samples |
| 4 | Select optimal beta1 |
| 5 | Full 50k training with best config |
| 6 | Release v1.0 if successful |
Citation
If you use this work for learning or research:
@misc{artn-ternary-2026,
author = {kishxrx},
title = {ARTN: Adaptive Ternary Recurrent Network - 1.58-bit Quantization Experiments},
year = {2026},
publisher = {HuggingFace},
url = {https://huggingface.co/kishxrx/artn-ternary-experimental}
}
License
MIT License
Disclaimer
THIS IS EXPERIMENTAL SOFTWARE.
- Models do NOT produce coherent text
- Checkpoints are for learning purposes only
- Not suitable for any production use
- No guarantees of any kind
Contact
- GitHub: https://github.com/kishxrx/ARTN
- HuggingFace: https://huggingface.co/kishxrx
Last Updated: August 2026
Version: v0.1-experimental