wskan-1m-tinystories
985,956 parameters · WSKAN-11 · byte-level LM trained on TinyStories (local-coherence tier)
Introduction
wskan-1m-tinystories is a checkpoint of WSKAN-11 (Wavelet-like State-Space KAN, version 11), a
Kolmogorov-Arnold Network in which every edge function is the impulse response of a stable
state-space model — a wavelet-like damped oscillator. One parameterization supports two modes: a
closed-form wavelet edge (static, WavKAN-style) and a recurrent SSM edge (Mamba-style selective
scan); this checkpoint runs the recurrent mode through a fused Triton associative-scan kernel.
This model is one cell of the 3-epoch release matrix (5 sizes x 3 datasets, seed 42) trained for the WSKAN architecture study. It is released to make the study's qualitative claims directly inspectable: the generation results below are real, unedited outputs of these exact weights.
- GitHub repository (code, reports, training scripts): https://github.com/llaa33219/wskan
- Architecture document: models/V11_README.md
- Definitive interpretation report: experiments/W11_FINAL_INTERPRETATION_REPORT.md
Features
- Kolmogorov-Arnold Network: learnable univariate edge functions replace scalar weights.
- Wavelet-like SSM edges: each edge function is the impulse response of a stable state-space model — a damped oscillator (decay sigma, frequency omega) that acts as a wavelet-like kernel.
- Fully selective recurrence: input-dependent step size dt and input/output projections (B, C), so the wavelet transform is content-warped per token (Mamba-style selectivity).
- Fused Triton scan: the associative scan is a single fused GPU kernel (log-depth), verified against the reference scan (fwd err 7e-7, grads <= 4e-6).
- Byte-level: vocabulary is the 256 byte values — no tokenizer, no BPE, no special tokens.
- Interpretable by construction: every edge exposes its learned (sigma, omega) spectrum; see the interpretation reports in the GitHub repo.
Model collection
| Model | Params | Dataset | Final eval loss | Byte PPL |
|---|---|---|---|---|
| wskan-1k-tinystories | 3,610 | TinyStories | 2.1416 | 8.51 |
| wskan-10k-tinystories | 12,546 | TinyStories | 1.4401 | 4.22 |
| wskan-100k-tinystories | 113,612 | TinyStories | 0.8224 | 2.28 |
| wskan-1m-tinystories (this model) | 985,956 | TinyStories | 0.6135 | 1.85 |
| wskan-10m-tinystories | 10,153,996 | TinyStories | 0.5453 | 1.73 |
| wskan-1k-ultrachat | 3,610 | UltraChat 200k | 2.4298 | 11.36 |
| wskan-10k-ultrachat | 12,546 | UltraChat 200k | 1.8724 | 6.5 |
| wskan-100k-ultrachat | 113,612 | UltraChat 200k | 1.2316 | 3.43 |
| wskan-1m-ultrachat | 985,956 | UltraChat 200k | 0.9854 | 2.68 |
| wskan-10m-ultrachat | 10,153,996 | UltraChat 200k | 0.8472 | 2.33 |
| wskan-1k-wikitext | 3,610 | WikiText-103 | 2.3627 | 10.62 |
| wskan-10k-wikitext | 12,546 | WikiText-103 | 1.9053 | 6.72 |
| wskan-100k-wikitext | 113,612 | WikiText-103 | 1.3508 | 3.86 |
| wskan-1m-wikitext | 985,956 | WikiText-103 | 1.1541 | 3.17 |
| wskan-10m-wikitext | 10,153,996 | WikiText-103 | 1.0266 | 2.79 |
This model is trained on children's stories (roneneldan/TinyStories); the other rows cover the remaining two domains and sizes. All rows use the identical WSKAN-11 architecture, differing only in depth/width.
Generation results (real outputs)
The following are actual outputs of this checkpoint, generated with
experiments/RELEASE_infer.py from the released weights (multinomial sampling, temperature 0.8,
seed 1234, 400 new bytes). Nothing is cherry-picked or edited — these are the first and only samples
drawn for this card.
Prompt:
Once upon a time
Output (seed 1234, temperature 0.8, 400 new bytes):
Once upon a time there was a blanket. It was red and had many leaves on it. One day, it was green and it was big and rain. The water was bright and had wings and blocks. The sun was shining and the wind was blue.
"Wow, look at that car!" Tom said. "It sounds like no one else there."
They ran to the car and sat down on the rock. Lily saw the box and the cat. She wanted to try it. She followed him and saw a big s
Prompt:
One day, a little girl named Lily found a
Output (seed 1234, temperature 0.8, 400 new bytes):
One day, a little girl named Lily found a scarf in her bag. She loved the red car seeds, so she whispered to the dog. She hugged her big smile and said, "We need to learn a lesson."
The little girl said, "You're welcome, Lucy. I was frustrated because I could have lovely and me. Do you want to come with me?"
Lucy smiled. "I don't know, sweetie. I will call it weigh a dragon!" her mom said.
Lucy was so excited. She ran to the shop and
Prompt:
Tom was very sad because he lost his
Output (seed 1234, temperature 0.8, 400 new bytes):
Tom was very sad because he lost his toy car. He was brave.
Spot hugged his mom and said, "See, you can have the carrots and sharp teeth. Can we be friends with Sam?" Missy smiled and hugged her mom.
"Yes, Max, you can't have it!" she learnt.
Max smiled and laughed and smiled. He knew that sharing is a good sister.
Once upon a time, there was a little girl named Lily. She loved to play with her toys, especially her teddy bear to
Evaluation of the outputs
Sentences are grammatical and locally consistent, with concrete imagery ("It was red and had many leaves on it") and working dialogue attribution. Multi-sentence coherence now holds for roughly 3-5 sentences before topic drift sets in. The remaining errors are semantic rather than grammatical — a blanket that is simultaneously "green and it was big and rain" — and entity tracking across a paragraph is unreliable (Tom speaks, then "she" follows the box). Story beats (problem, reaction, resolution) appear spontaneously. A competent stylistic model of children's narrative.
Objective metrics for this checkpoint: final eval loss 0.6135 (byte-level perplexity 1.85) on a held-out slice of TinyStories, at training step 100,000.
Configuration
The released config.yaml (verbatim):
model_name: wskan-1m-tinystories
architecture_family: WSKAN-11 (Wavelet-like State-Space KAN, fused Triton scan)
architecture:
vocab_size: 256
d_model: 80
n_layers: 6
n_states: 6
use_feature_bc: true
bc_rank: 32
wz_diag: false
g_rank: null
oscillatory: true
bf16_scan: true
tokenizer:
type: byte-level
description: raw UTF-8 bytes; vocab ids 0-255; no BPE, no special tokens
training:
dataset: roneneldan/TinyStories
steps: 100000
checkpoint_step: 100000
batch_size: 64
block_size: 256
lr: 0.001
lr_schedule: cosine
optimizer: AdamW (weight_decay=0.0, grad clip 1.0)
seed: 42
epochs: ~3
precision: fp32 weights, bf16 scan
evaluation:
final_eval_loss: 0.6135
final_eval_byte_ppl: 1.85
generation_defaults:
max_new: 400
temperature: 0.8
seed: 1234
Architecture quick reference: d_model=80, n_layers=6, n_states=6, bc_rank=32, vocab=256 (raw bytes). Training: 100,000 steps, batch 64, block 256, lr 0.001 (cosine), AdamW, seed 42, ~3 epochs, bf16 scan.
Local run
Inference requires an NVIDIA GPU with Triton (the scan kernel is fused Triton; there is no CPU path) and Python 3.12+.
git clone https://github.com/llaa33219/wskan.git
cd wskan
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -r requirements.txt
uv pip install --python .venv/bin/python safetensors pyyaml huggingface_hub
# generate from this model (downloads config.yaml + model.safetensors from the Hub)
.venv/bin/python experiments/RELEASE_infer.py \
--model llaa33219/wskan-1m-tinystories \
--prompt "Once upon a time" \
--max-new 400 --temperature 0.8 --seed 1234
--model also accepts a local directory containing config.yaml and model.safetensors.
Weights are fp32; set bf16_scan: false in config.yaml for an fp32 scan (outputs will differ
slightly from the samples above, which used the training-time bf16 scan).
To retrain from scratch under the identical protocol:
uv pip install --python .venv/bin/python datasets transformers
.venv/bin/python experiments/V1_train_tinystories_lm.py \
--model wskan11 --scale 1m --dataset tinystories --seed 42 \
--batch 64 --block 256 \
--steps 100000 --lr 0.001 --lr-schedule cosine \
--ckpt-every 25000 --eval-every 1000 --out-tag 3ep --compile --bf16
Limitations
- Research artifact. These are 3.6k-10M parameter proof-of-concept models from an architecture study, not production LMs.
- No factual reliability. All models confabulate names, dates, and facts; the larger ones merely do so more fluently. Never use outputs as a source of truth.
- Not instruction-tuned. Even the UltraChat models only imitate conversational format; they do not follow instructions.
- Short-range memory. The interpretation reports show the models' effective memory is local (dozens of bytes); long-range consistency is not to be expected.
- English-only, single-domain. Each model knows only its training slice.
- CUDA GPU required for inference. The recurrent scan is a fused Triton kernel; there is no CPU
path. (
config.yaml: architecture.bf16_scan: falseswitches the scan to fp32 on GPU.) - Sampling only. Generation is multinomial sampling; outputs vary with seed and temperature.
Provenance
- Checkpoint:
checkpoints/wskan11_tinystories_1m_3ep_s42/latest.pt(step 100,000) of the WSKAN repository, converted to safetensors (fp32) without any weight modification. - Training data: TinyStories, byte-encoded UTF-8.
- Seed 42 is the canonical seed of the 5-seed campaign; the released sample outputs above used sampling seed 1234.
- Downloads last month
- 18