YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

SuperSmallJokeClaude

A 1.64M-parameter tiny transformer trained from scratch on Fable 5 Claude Code assistant traces.

Specs

  • Params: 1,644,032 (1.64M)
  • Architecture: 4-layer transformer, d_model=128, 4 heads, d_ff=256, tied embeddings
  • Vocab: BPE-8192 (from Claude Code tokenizer)
  • Context: 512 tokens
  • Training data: 24,249 tokens (Fable 5 Claude Code assistant traces)
  • Training: 2000 steps, batch 8, seq 256, AdamW lr=3e-4 cosine schedule
  • Final loss: 2.90
  • Training time: 7 seconds on RTX 5090

What it does

It's a joke โ€” a tiny model trained on tiny data. It will produce degenerate output (repeated tokens, gibberish). It exists to demonstrate that you can train a working transformer in 7 seconds on a single GPU.

Files

  • final.pt โ€” full checkpoint (model state dict + config + training metadata)
  • tokenizer.json โ€” BPE-8192 tokenizer (Claude Code)

Load it

import torch
from tokenizers import Tokenizer

ckpt = torch.load('final.pt', map_location='cpu')
config = ckpt['config']

# Build model (see train_jokeclaude_v5.py for the class definition)
model = Transformer(**config)
model.load_state_dict(ckpt['model_state_dict'])
model.eval()

tokenizer = Tokenizer.from_file('tokenizer.json')

Disclaimer

This is an experimental/fun artifact. The model is not useful for any real task. It exists because someone asked for a "super small joke claude" and we said why not.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support