πŸ¦β€β¬› RAVEN-1

A 29.5M parameter language model, built entirely from scratch.

No frameworks. No shortcuts. Just PyTorch, math, and stubbornness.


Python PyTorch Parameters Vocab Context License


What is RAVEN-1?

RAVEN-1 is a GPT-style decoder-only transformer language model with 29.51 million parameters. Every single component β€” tokenizer, model architecture, training loop, data pipeline, inference engine, and web UI β€” is written from scratch in Python and PyTorch.

No HuggingFace transformers. No pre-trained weights borrowed. No abstractions hiding the work.

Why Build From Scratch?

  • Full understanding β€” Every matrix multiply, every attention mask, every gradient step is explicit and visible.
  • Colab-friendly β€” Designed to train end-to-end on Google Colab's free T4 GPU (15GB VRAM, 12GB RAM).
  • Two-phase training β€” Pre-trained on a 389MB text corpus for language modeling, then post-trained on 977MB of curated conversational data for personality.
  • Production-grade tooling β€” Resume-safe checkpoints, mixed precision, Google Drive sync, streaming data pipelines, and a Gradio web UI for deployment.

Personality

RAVEN-1 was post-trained on conversational data that gives it a distinct personality β€” dry, sarcastic, concise. It's not trying to be helpful. It's trying to be honest. Sometimes annoyingly so.

User: I'm bored
Raven: Good. Boredom is where ideas start.

User: good morning
Raven: Morning. Let's keep it simple.

User: thanks!
Raven: Thanks. I'll file that away.


πŸš€ Quick Start (Download Weights from Hugging Face)

Because model checkpoints (.pt files) are large, they are not tracked in this GitHub repository. Instead, the full model weights and trained BPE tokenizer are hosted on Hugging Face.

To set up and run RAVEN-1 locally, follow these simple steps to download the repository and fetch the weights directly from Hugging Face:

1. Clone & Set Up Environment

# Clone the repository
git clone https://github.com/itslokeshx/Raven-1.git
cd Raven-1

# Create the checkpoints directory
mkdir -p checkpoints

# Install required packages
pip install -r requirements.txt

2. Download Tokenizer & Weights from Hugging Face

Download the files directly into the checkpoints/ directory:

# 1. Download BPE Tokenizer Config
wget -P checkpoints/ https://huggingface.co/itslokeshx/Raven-1/resolve/main/checkpoints/raven1_tokenizer.json

# 2. Download Post-Trained Best Weights (Highly Recommended - Chat & Logic)
wget -P checkpoints/ https://huggingface.co/itslokeshx/Raven-1/resolve/main/checkpoints/raven1_posttrain_best.pt

# 3. Download Pre-Trained Base Weights (Optional)
wget -P checkpoints/ https://huggingface.co/itslokeshx/Raven-1/resolve/main/checkpoints/raven1_best.pt

3. Run the Gradio Web Chat Interface

Once the weights are in checkpoints/, start the premium local web chatbot UI:

python app.py

Open your browser and navigate to http://localhost:7860 to chat with RAVEN-1!


Architecture

Input Tokens (sequence of integers)
     β”‚
     β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚    Token Embedding (8,192 β†’ 512)     β”‚  Weight-tied with LM Head
β”‚  + Learned Position Embedding (256)  β”‚
β”‚  + Dropout (0.1)                     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                   β”‚
                   β–Ό
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β”‚                     β”‚
        β”‚   8Γ— Transformer    β”‚  ◄── Pre-LayerNorm design
        β”‚       Block         β”‚
        β”‚                     β”‚
        β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
        β”‚  β”‚ LayerNorm     β”‚  β”‚
        β”‚  β”‚ ↓             β”‚  β”‚
        β”‚  β”‚ Multi-Head    β”‚  β”‚  8 heads Γ— 64 dim = 512
        β”‚  β”‚ Causal Attn   β”‚  β”‚  FlashAttention (auto-dispatched)
        β”‚  β”‚ + Residual    β”‚  β”‚  Combined Q/K/V projection
        β”‚  β”‚               β”‚  β”‚
        β”‚  β”‚ LayerNorm     β”‚  β”‚
        β”‚  β”‚ ↓             β”‚  β”‚
        β”‚  β”‚ Feed-Forward  β”‚  β”‚  512 β†’ 2048 β†’ 512
        β”‚  β”‚ (GELU)        β”‚  β”‚  No bias terms
        β”‚  β”‚ + Residual    β”‚  β”‚
        β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
        β”‚                     β”‚
        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                   β”‚
                   β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚         Final LayerNorm              β”‚
β”‚         LM Head (512 β†’ 8,192)       β”‚  Weight-tied with Token Embedding
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                   β”‚
                   β–Ό
            Logits (8,192)

Model Specifications

Parameter Value Notes
Total Parameters 29.51M Counted with weight tying
Vocab Size 8,192 Byte-level BPE
Context Length 256 tokens Maximum sequence length
Embedding Dim 512 n_embd
Attention Heads 8 n_head
Head Dimension 64 n_embd / n_head
Transformer Layers 8 n_layer
FFN Inner Dim 2,048 4 Γ— n_embd
Activation GELU In feed-forward network
Normalization Pre-LayerNorm Before attention & FFN
Dropout 0.1 Embedding, attention, FFN residual
Weight Tying Yes LM head shares token embedding weights
Attention FlashAttention Auto-dispatched via F.scaled_dot_product_attention
Bias Terms None All Linear layers are bias-free

Special Tokens

Token ID Purpose
<|pad|> 0 Padding
<|user|> 1 User turn marker
<|raven|> 2 Model turn marker
<|eos|> 3 End of sequence

Project Structure

Raven-1/
β”‚
β”œβ”€β”€ config.py                  # Central configuration β€” ALL hyperparameters
β”œβ”€β”€ model.py                   # Transformer architecture + generation
β”œβ”€β”€ train.py                   # Pre-training loop
β”œβ”€β”€ posttrain.py               # Post-training / fine-tuning loop
β”œβ”€β”€ tokenizer_train.py         # Train byte-level BPE tokenizer
β”œβ”€β”€ data_prep.py               # Tokenize text β†’ binary memmap files
β”œβ”€β”€ posttrain_data_prep.py     # Streaming post-train data pipeline
β”œβ”€β”€ inference.py               # Terminal chat with slash commands
β”œβ”€β”€ app.py                     # Gradio web UI (HF Spaces compatible)
β”‚
β”œβ”€β”€ eval/                      # Evaluation & benchmarks
β”‚   β”œβ”€β”€ run_300_test.py        # 300-prompt benchmark suite
β”‚   β”œβ”€β”€ stress_test.py         # Stress test with long inputs
β”‚   β”œβ”€β”€ pretrain_results.txt   # Pre-train benchmark results
β”‚   └── posttrain_results.txt  # Post-train benchmark results
β”‚
β”œβ”€β”€ notebooks/                 # Colab training notebooks
β”‚   β”œβ”€β”€ colab_train.ipynb      # One-click pre-training notebook
β”‚   └── colab_posttrain.ipynb  # One-click post-training notebook
β”‚
β”œβ”€β”€ data/                      # Training corpora (not in git)
β”‚   β”œβ”€β”€ corpus.txt             # Pre-training corpus (389 MB)
β”‚   └── posttrain.txt          # Post-training data (977 MB, ~13M lines)
β”‚
β”œβ”€β”€ checkpoints/               # Generated at runtime (not in git)
β”‚   β”œβ”€β”€ raven1_tokenizer.json  # Trained BPE tokenizer (557 KB)
β”‚   β”œβ”€β”€ raven1_best.pt         # Best pre-trained weights (338 MB)
β”‚   β”œβ”€β”€ raven1_posttrain_best.pt  # Best post-trained weights (338 MB)
β”‚   β”œβ”€β”€ train.bin / val.bin    # Pre-training binary data
β”‚   └── posttrain_train.bin / posttrain_val.bin
β”‚
β”œβ”€β”€ requirements.txt
β”œβ”€β”€ README.md
└── .gitignore

Total hand-written code: ~2,480 lines across 11 Python files.


Quick Start

1. Install Dependencies

pip install -r requirements.txt

2. Run Inference (Terminal Chat)

python inference.py

The model automatically loads the best available weights:

  1. Post-train weights (raven1_posttrain_best.pt) β€” tried first
  2. Pre-train weights (raven1_best.pt) β€” fallback
  3. Random weights β€” if no checkpoints found

3. Run Web UI

python app.py
# β†’ Opens at http://localhost:7860

Training Pipeline

The full pipeline has 5 stages, each with a dedicated script. Every stage is independently runnable, resume-safe, and Colab-compatible.

Stage 1 β€” Train Tokenizer

python tokenizer_train.py

Trains a byte-level BPE tokenizer on data/corpus.txt using HuggingFace's tokenizers library (Rust backend for speed).

Detail Value
Algorithm Byte-level BPE
Vocab size 8,192
Min frequency 2
Special tokens <|pad|>, <|user|>, <|raven|>, <|eos|>
Output checkpoints/raven1_tokenizer.json
Corpus data/corpus.txt (389 MB)

Stage 2 β€” Prepare Data

# Pre-training data
python data_prep.py

# Post-training data (streaming, handles 1GB+ without RAM issues)
python posttrain_data_prep.py

Tokenizes text corpora into binary uint16 memmap files for zero-copy random access during training.

Key features:

  • Chunked processing β€” reads 10MB (data_prep) or 8MB (posttrain_data_prep) at a time to stay within Colab's 12GB RAM
  • Streaming parser β€” reconstructs <|eos|>-delimited samples across chunk boundaries
  • Batch tokenization β€” uses encode_batch() with Rust multi-threading (4,096 samples per batch)
  • Metadata cache β€” skips rebuild if source file and tokenizer haven't changed
  • Malformed sample detection β€” warns if >1% of samples lack required special tokens
Data Split Pre-training Post-training
Split ratio 95% / 5% 90% / 10%
Source data/corpus.txt (389 MB) data/posttrain.txt (977 MB)
Output format uint16 memmap uint16 memmap

Stage 3 β€” Pre-train

python train.py

Full pre-training loop with production-grade features:

Feature Implementation
⚑ Mixed precision AMP + GradScaler (automatic loss scaling)
πŸ“ˆ LR schedule Cosine decay with linear warmup
πŸ”„ Gradient accumulation Effective batch = micro_batch Γ— accum_steps
βœ‚οΈ Gradient clipping Max norm 1.0
πŸ’Ύ Auto-resume Loads latest checkpoint automatically
⭐ Best model tracking Saves on val_loss improvement
☁️ Drive sync Auto-backup to Google Drive on Colab
πŸ›‘ Clean interrupt Ctrl+C saves checkpoint before exit
πŸ“Š Parameter groups Weight decay on 2D+ params only, none on biases/norms
Hyperparameter Value
Peak learning rate 3e-4
Min learning rate 3e-5 (cosine floor)
Warmup steps 300
Total steps 10,000
Micro batch size 32
Gradient accumulation 2 steps
Effective batch size 64
Tokens per step 64 Γ— 256 = 16,384
Weight decay 0.1
Optimizer AdamW (β₁=0.9, Ξ²β‚‚=0.95)
Eval interval Every 500 steps
Eval batches 100

Pre-train results:

  • Best checkpoint: step 9,000
  • Best val loss: 0.5906

Stage 4 β€” Post-train

python posttrain.py              # Full post-training
python posttrain.py --resume     # Resume from checkpoint
python posttrain.py --dry-run    # Test 20 steps only

Fine-tunes the pre-trained model on curated conversational data. Uses the same training infrastructure with lower learning rates and generates before/after sample comparisons.

Hyperparameter Value
Peak learning rate 5e-5
Min learning rate 5e-6
Warmup steps 100
Total steps 2,000
Eval interval Every 250 steps
Log interval Every 25 steps

Post-train results:

  • Best checkpoint: step 2,000
  • Best val loss: 0.5950

Stage 5 β€” Test & Evaluate

python eval/run_300_test.py     # 300-prompt benchmark across 6 categories
python eval/stress_test.py      # Stress test with very long inputs
python inference.py             # Interactive terminal chat

Inference

Terminal Chat (inference.py)

python inference.py

Interactive chat with conversation history (last 3 turns), multiple sampling modes, and slash commands.

══════════════════════════════
  RAVEN-1  |  29.5M parameters
  Loaded: posttrain
  Device: cpu
══════════════════════════════

You: hey there
Raven: You found me. Now what.

You: I'm tired of everything
Raven: Everything is temporary. Including this conversation, hopefully.

You: /mode chaos
  Mode β†’ chaos

You: /exit
Bye.

Slash commands:

Command Action
/mode [name] Switch sampling mode (default, sharp, chaos, cold)
/modes List all modes with current parameters
/clear Clear terminal screen
/reset Clear conversation history
/exit Exit the chat

Web UI (app.py)

python app.py
# β†’ http://localhost:7860

Premium dark monochrome Gradio interface. Same generation logic as the terminal chat. Features mode selection dropdown and conversation history. Deployable directly to Hugging Face Spaces.

UI tech stack: Gradio with custom dark theme, Inter font (Google Fonts), #0a0a0a background, 12px border-radius, 700px max-width container.

Generation Pipeline

Both inference methods use the same pipeline:

  1. Build prompt with conversation context (last 3 turns)
  2. Encode with BPE tokenizer
  3. Crop to block_size (256) if needed (keeps rightmost tokens)
  4. Generate with autoregressive sampling:
    • Temperature scaling
    • Top-K filtering
    • Top-P (nucleus) filtering
    • Repetition penalty on previously generated tokens
  5. Decode only new tokens
  6. Clean leaked special tokens
  7. Fallback to witty canned response if output is empty

Training Results

Pre-training

Metric Value
Corpus 389 MB raw text
Steps completed 10,000
Best step 9,000
Best val loss 0.5906
Checkpoint size 338 MB

Post-training

Metric Value
Corpus 977 MB conversational data (~13M lines)
Steps completed 2,000
Best step 2,000
Best val loss 0.5950
Checkpoint size 338 MB

Loss Trajectory

The model converges smoothly with cosine decay:

Pre-train:   random init β†’ ~7.0 loss β†’ 0.5906 val loss (9,000 steps)
Post-train:  0.5906 pretrain β†’ 0.5950 posttrain val loss (2,000 steps)

300-Prompt Benchmark

Both the pre-trained and post-trained models were evaluated on 300 diverse prompts across 6 categories using identical sampling parameters (default mode).

Benchmark Comparison

Metric Pre-train Post-train Ξ”
Total time 358.02s 181.56s 1.97Γ— faster
Avg latency 1.19s 0.61s 49% reduction
Min latency 0.08s β€” β€”
Max latency 3.28s β€” β€”

Category Distribution

Category Count Description
casual 53 Greetings, small talk, emotions
coding 52 Code snippets, technical prompts
creative 63 Stories, poems, hypotheticals
edge_cases 43 Empty inputs, repeated chars, symbols
math 36 Arithmetic, algebra, number theory
questions 53 Factual knowledge queries

Sample Post-train Outputs

Casual / Conversational (strong improvement):

Input Output
hello Hello. Make it quick.
I'm bored Good. Boredom is where ideas start.
what's up? Morning. Let's keep it simple.
good morning The universe trembles. Or not.
good night Night. The quiet suits you.
bye Still here. The quiet suits you.
I am happy Noted. Now use it.
class MyClass: The body is being patient. Do not test that.

Personality-driven responses:

Input Output
make me laugh 1. You survived basic arithmetic.
who are you? I was here before you typed that. Go ahead.
how does electricity work? Electricity's just a phase. Probably.
what is gravity? Mass pulling on mass. Keeps you grounded.
what is AI? The answer existed before you asked.

Note: The post-train corpus included some competitive programming data, which occasionally causes reasoning-style leakage (e.g., "To solve this problem...") on certain prompts. This is a data quality issue, not a model architecture issue.


Sampling Modes

Mode Temperature Top-K Top-P Repetition Penalty Character
default 0.65 25 0.85 1.10 Balanced, reliable
sharp 0.55 20 0.80 1.10 More focused, deterministic
chaos 0.85 45 0.92 1.05 Creative, unpredictable
cold 0.45 12 0.75 1.08 Most conservative, factual

All modes generate up to 60 new tokens per response and stop on <|eos|>.


Configuration Reference

All hyperparameters live in config.py β€” zero hardcoded values anywhere else in the codebase. Every script imports from this single source of truth.

# ── Model Architecture ──
vocab_size   = 8192        # BPE vocabulary size
block_size   = 256         # Maximum context length (tokens)
n_embd       = 512         # Embedding dimension
n_head       = 8           # Number of attention heads
n_layer      = 8           # Number of transformer blocks
dropout      = 0.1         # Dropout rate during training

# ── Pre-training ──
batch_size   = 32          # Micro-batch size per step
gradient_accumulation_steps = 2   # Effective batch = 64
max_iters    = 10000       # Total pre-training steps
learning_rate = 3e-4       # Peak learning rate
min_lr       = 3e-5        # Cosine floor
warmup_iters = 300         # Linear warmup steps
weight_decay = 0.1         # AdamW weight decay
grad_clip    = 1.0         # Gradient norm clipping
eval_interval = 500        # Evaluate every N steps
eval_iters   = 100         # Batches per evaluation
use_amp      = True        # Mixed precision training

# ── Post-training ──
posttrain_max_iters       = 2000
posttrain_learning_rate   = 5e-5
posttrain_min_lr          = 5e-6
posttrain_warmup_iters    = 100

# ── Special Tokens ──
user_token   = "<|user|>"
raven_token  = "<|raven|>"
eos_token    = "<|eos|>"
pad_token    = "<|pad|>"

# ── Device ──
device = "cuda" if torch.cuda.is_available() else "cpu"

Design Decisions

Architecture Choices

Decision Rationale
Pre-LayerNorm LayerNorm before attention/FFN (not after) stabilizes training for small models. Standard in GPT-2+ and modern transformers.
Weight tying LM head shares weights with token embedding. Reduces parameter count by ~4M and improves generalization.
FlashAttention F.scaled_dot_product_attention auto-dispatches to FlashAttention on supported hardware (A100, H100, T4). Falls back to standard attention elsewhere.
No bias All nn.Linear layers use bias=False. Modern practice from LLaMA/PaLM β€” reduces parameters and doesn't hurt quality.
GELU activation Smoother than ReLU, standard in transformer FFNs since GPT-2.
Combined QKV projection Single nn.Linear(n_embd, 3 * n_embd) then chunk, instead of three separate projections. More efficient, same result.

Training Choices

Decision Rationale
Cosine schedule + warmup Prevents instability at start (warmup) and enables graceful convergence (cosine decay). Industry standard.
Two parameter groups Weight decay on 2D+ parameters (weight matrices) only. Biases and LayerNorm parameters get zero weight decay. Prevents regularization interference with normalization.
AdamW (β₁=0.9, Ξ²β‚‚=0.95) Ξ²β‚‚=0.95 instead of default 0.999 β€” better for transformers, reduces sensitivity to gradient spikes.
Gradient accumulation Achieves effective batch size of 64 with only 32 samples per micro-step. Essential for Colab's limited VRAM.
Two-phase training Phase 1 (pre-train) teaches general language modeling. Phase 2 (post-train) teaches conversational structure and personality. Separate LR schedules for each phase.

Data Pipeline Choices

Decision Rationale
Byte-level BPE Handles all Unicode without unknown tokens. 8,192 vocab is compact enough for a 29.5M model.
Chunked tokenization Processes 8–10 MB chunks to stay within Colab's 12GB RAM. Never loads full corpus.
Memmap data Training data stored as memory-mapped uint16 arrays. Zero-copy random access, no RAM overhead.
Streaming parser posttrain_data_prep.py reads samples across chunk boundaries without holding the full file in memory. Handles 1GB+ datasets.
Metadata cache Saves source file hash/size/mtime. Skips rebuild if nothing changed.

Inference Choices

Decision Rationale
60 token max Keeps responses concise. RAVEN-1 is designed for short, punchy responses, not essays.
Repetition penalty Penalizes tokens that already appeared in context. Prevents degenerate repetition loops.
Top-K + Top-P Combined filtering: Top-K removes long tail, Top-P (nucleus) adapts to probability distribution shape. Together they produce diverse but coherent text.
Funny fallbacks If the model generates empty output, a random witty fallback is returned instead of an error. Keeps the UX clean.
3-turn history Conversation context is limited to the last 3 user/raven exchanges. Keeps the prompt within context window and focuses on recent conversation.

Google Colab Guide

RAVEN-1 is designed to train end-to-end on Google Colab's free tier (T4 GPU, 15GB VRAM, 12GB RAM). Two notebooks are provided:

Pre-training (notebooks/colab_train.ipynb)

  1. Mount Drive & check GPU β†’ verify T4
  2. Install dependencies β†’ tokenizers, numpy
  3. Create directories β†’ upload project files
  4. Train tokenizer β†’ python tokenizer_train.py
  5. Prepare data β†’ python data_prep.py (chunked, RAM-safe)
  6. Pre-train β†’ python train.py (auto-resumes, syncs to Drive)
  7. Post-train β†’ python posttrain.py
  8. Test β†’ python inference.py
  9. Download weights β†’ .pt + tokenizer .json

Post-training (notebooks/colab_posttrain.ipynb)

  1. Mount Drive & check GPU
  2. Install dependencies
  3. Upload code files (6 files)
  4. Copy posttrain.txt from Drive
  5. Copy pre-trained weights from Drive
  6. Pre-flight check (validates all files)
  7. Tokenize data β†’ python posttrain_data_prep.py (streaming, ~1 min for 1GB)
  8. Post-train β†’ python posttrain.py --resume (resume-safe)
  9. Test the model
  10. Download weights

πŸ’‘ Tip: If the Colab session disconnects, just re-run from the training cell β€” it auto-resumes from the latest checkpoint synced to Google Drive. Checkpoints include optimizer state, scaler state, step count, and best val loss for exact continuation.

Colab Resource Usage

Resource Pre-training Post-training
GPU T4 (15GB VRAM) T4 (15GB VRAM)
VRAM used ~3-4 GB ~3-4 GB
RAM used < 6 GB < 6 GB
Disk ~2 GB ~2 GB
Time ~2-3 hours ~30-45 minutes

Requirements

torch>=2.1.0
tokenizers>=0.15.0
numpy>=1.24.0
gradio>=4.0.0
pip install -r requirements.txt

Compatibility

  • Python: 3.10+
  • PyTorch: 2.1+ (for F.scaled_dot_product_attention / FlashAttention)
  • CUDA: Optional. Runs on CPU (slower) or any CUDA-capable GPU
  • OS: Linux, macOS, Windows
  • Hardware: Colab T4 (minimum), any modern GPU, or CPU

File Reference

File Lines Purpose
config.py 65 Central configuration β€” every hyperparameter
model.py 299 Transformer: attention, FFN, blocks, generation
tokenizer_train.py 87 Train byte-level BPE from corpus
data_prep.py 124 Tokenize corpus β†’ binary memmap (pre-train)
posttrain_data_prep.py 447 Streaming tokenizer β†’ binary memmap (post-train)
train.py 329 Pre-training loop
posttrain.py 424 Post-training / fine-tuning loop
inference.py 240 Terminal chat with slash commands
app.py 269 Gradio web UI
eval/run_300_test.py 175 300-prompt benchmark suite
eval/stress_test.py 51 Stress test with long inputs

Checkpoint Format

Each .pt checkpoint contains:

{
    "step": int,                 # Training step number
    "model_state": OrderedDict,  # Model weights (69 tensors)
    "optimizer_state": dict,     # AdamW state for exact resume
    "scaler_state": dict,        # AMP GradScaler state
    "best_val_loss": float,      # Best validation loss seen
    "config": {                  # Architecture config snapshot
        "vocab_size": 8192,
        "block_size": 256,
        "n_embd": 512,
        "n_head": 8,
        "n_layer": 8,
        "dropout": 0.1,
    }
}

Checkpoint size: ~338 MB (includes optimizer momentum buffers).


License

MIT


Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support