YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
SuperSmallJokeClaude
A 1.64M-parameter tiny transformer trained from scratch on Fable 5 Claude Code assistant traces.
Specs
- Params: 1,644,032 (1.64M)
- Architecture: 4-layer transformer, d_model=128, 4 heads, d_ff=256, tied embeddings
- Vocab: BPE-8192 (from Claude Code tokenizer)
- Context: 512 tokens
- Training data: 24,249 tokens (Fable 5 Claude Code assistant traces)
- Training: 2000 steps, batch 8, seq 256, AdamW lr=3e-4 cosine schedule
- Final loss: 2.90
- Training time: 7 seconds on RTX 5090
What it does
It's a joke โ a tiny model trained on tiny data. It will produce degenerate output (repeated tokens, gibberish). It exists to demonstrate that you can train a working transformer in 7 seconds on a single GPU.
Files
final.ptโ full checkpoint (model state dict + config + training metadata)tokenizer.jsonโ BPE-8192 tokenizer (Claude Code)
Load it
import torch
from tokenizers import Tokenizer
ckpt = torch.load('final.pt', map_location='cpu')
config = ckpt['config']
# Build model (see train_jokeclaude_v5.py for the class definition)
model = Transformer(**config)
model.load_state_dict(ckpt['model_state_dict'])
model.eval()
tokenizer = Tokenizer.from_file('tokenizer.json')
Disclaimer
This is an experimental/fun artifact. The model is not useful for any real task. It exists because someone asked for a "super small joke claude" and we said why not.
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support