WSKAN logo

wskan-1m-tinystories

985,956 parameters · WSKAN-11 · byte-level LM trained on TinyStories (local-coherence tier)

Introduction

wskan-1m-tinystories is a checkpoint of WSKAN-11 (Wavelet-like State-Space KAN, version 11), a Kolmogorov-Arnold Network in which every edge function is the impulse response of a stable state-space model — a wavelet-like damped oscillator. One parameterization supports two modes: a closed-form wavelet edge (static, WavKAN-style) and a recurrent SSM edge (Mamba-style selective scan); this checkpoint runs the recurrent mode through a fused Triton associative-scan kernel.

This model is one cell of the 3-epoch release matrix (5 sizes x 3 datasets, seed 42) trained for the WSKAN architecture study. It is released to make the study's qualitative claims directly inspectable: the generation results below are real, unedited outputs of these exact weights.

Features

  • Kolmogorov-Arnold Network: learnable univariate edge functions replace scalar weights.
  • Wavelet-like SSM edges: each edge function is the impulse response of a stable state-space model — a damped oscillator (decay sigma, frequency omega) that acts as a wavelet-like kernel.
  • Fully selective recurrence: input-dependent step size dt and input/output projections (B, C), so the wavelet transform is content-warped per token (Mamba-style selectivity).
  • Fused Triton scan: the associative scan is a single fused GPU kernel (log-depth), verified against the reference scan (fwd err 7e-7, grads <= 4e-6).
  • Byte-level: vocabulary is the 256 byte values — no tokenizer, no BPE, no special tokens.
  • Interpretable by construction: every edge exposes its learned (sigma, omega) spectrum; see the interpretation reports in the GitHub repo.

Model collection

Model Params Dataset Final eval loss Byte PPL
wskan-1k-tinystories 3,610 TinyStories 2.1416 8.51
wskan-10k-tinystories 12,546 TinyStories 1.4401 4.22
wskan-100k-tinystories 113,612 TinyStories 0.8224 2.28
wskan-1m-tinystories (this model) 985,956 TinyStories 0.6135 1.85
wskan-10m-tinystories 10,153,996 TinyStories 0.5453 1.73
wskan-1k-ultrachat 3,610 UltraChat 200k 2.4298 11.36
wskan-10k-ultrachat 12,546 UltraChat 200k 1.8724 6.5
wskan-100k-ultrachat 113,612 UltraChat 200k 1.2316 3.43
wskan-1m-ultrachat 985,956 UltraChat 200k 0.9854 2.68
wskan-10m-ultrachat 10,153,996 UltraChat 200k 0.8472 2.33
wskan-1k-wikitext 3,610 WikiText-103 2.3627 10.62
wskan-10k-wikitext 12,546 WikiText-103 1.9053 6.72
wskan-100k-wikitext 113,612 WikiText-103 1.3508 3.86
wskan-1m-wikitext 985,956 WikiText-103 1.1541 3.17
wskan-10m-wikitext 10,153,996 WikiText-103 1.0266 2.79

This model is trained on children's stories (roneneldan/TinyStories); the other rows cover the remaining two domains and sizes. All rows use the identical WSKAN-11 architecture, differing only in depth/width.

Generation results (real outputs)

The following are actual outputs of this checkpoint, generated with experiments/RELEASE_infer.py from the released weights (multinomial sampling, temperature 0.8, seed 1234, 400 new bytes). Nothing is cherry-picked or edited — these are the first and only samples drawn for this card.

Prompt:

Once upon a time

Output (seed 1234, temperature 0.8, 400 new bytes):

Once upon a time there was a blanket. It was red and had many leaves on it. One day, it was green and it was big and rain. The water was bright and had wings and blocks. The sun was shining and the wind was blue.

"Wow, look at that car!" Tom said. "It sounds like no one else there."

They ran to the car and sat down on the rock. Lily saw the box and the cat. She wanted to try it. She followed him and saw a big s

Prompt:

One day, a little girl named Lily found a

Output (seed 1234, temperature 0.8, 400 new bytes):

One day, a little girl named Lily found a scarf in her bag. She loved the red car seeds, so she whispered to the dog. She hugged her big smile and said, "We need to learn a lesson."

The little girl said, "You're welcome, Lucy. I was frustrated because I could have lovely and me. Do you want to come with me?"

Lucy smiled. "I don't know, sweetie. I will call it weigh a dragon!" her mom said.

Lucy was so excited. She ran to the shop and 

Prompt:

Tom was very sad because he lost his

Output (seed 1234, temperature 0.8, 400 new bytes):

Tom was very sad because he lost his toy car. He was brave.

Spot hugged his mom and said, "See, you can have the carrots and sharp teeth. Can we be friends with Sam?" Missy smiled and hugged her mom.

"Yes, Max, you can't have it!" she learnt.

Max smiled and laughed and smiled. He knew that sharing is a good sister.
Once upon a time, there was a little girl named Lily. She loved to play with her toys, especially her teddy bear to 

Evaluation of the outputs

Sentences are grammatical and locally consistent, with concrete imagery ("It was red and had many leaves on it") and working dialogue attribution. Multi-sentence coherence now holds for roughly 3-5 sentences before topic drift sets in. The remaining errors are semantic rather than grammatical — a blanket that is simultaneously "green and it was big and rain" — and entity tracking across a paragraph is unreliable (Tom speaks, then "she" follows the box). Story beats (problem, reaction, resolution) appear spontaneously. A competent stylistic model of children's narrative.

Objective metrics for this checkpoint: final eval loss 0.6135 (byte-level perplexity 1.85) on a held-out slice of TinyStories, at training step 100,000.

Configuration

The released config.yaml (verbatim):

model_name: wskan-1m-tinystories
architecture_family: WSKAN-11 (Wavelet-like State-Space KAN, fused Triton scan)
architecture:
  vocab_size: 256
  d_model: 80
  n_layers: 6
  n_states: 6
  use_feature_bc: true
  bc_rank: 32
  wz_diag: false
  g_rank: null
  oscillatory: true
  bf16_scan: true
tokenizer:
  type: byte-level
  description: raw UTF-8 bytes; vocab ids 0-255; no BPE, no special tokens
training:
  dataset: roneneldan/TinyStories
  steps: 100000
  checkpoint_step: 100000
  batch_size: 64
  block_size: 256
  lr: 0.001
  lr_schedule: cosine
  optimizer: AdamW (weight_decay=0.0, grad clip 1.0)
  seed: 42
  epochs: ~3
  precision: fp32 weights, bf16 scan
evaluation:
  final_eval_loss: 0.6135
  final_eval_byte_ppl: 1.85
generation_defaults:
  max_new: 400
  temperature: 0.8
  seed: 1234

Architecture quick reference: d_model=80, n_layers=6, n_states=6, bc_rank=32, vocab=256 (raw bytes). Training: 100,000 steps, batch 64, block 256, lr 0.001 (cosine), AdamW, seed 42, ~3 epochs, bf16 scan.

Local run

Inference requires an NVIDIA GPU with Triton (the scan kernel is fused Triton; there is no CPU path) and Python 3.12+.

git clone https://github.com/llaa33219/wskan.git
cd wskan
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -r requirements.txt
uv pip install --python .venv/bin/python safetensors pyyaml huggingface_hub

# generate from this model (downloads config.yaml + model.safetensors from the Hub)
.venv/bin/python experiments/RELEASE_infer.py \
    --model llaa33219/wskan-1m-tinystories \
    --prompt "Once upon a time" \
    --max-new 400 --temperature 0.8 --seed 1234

--model also accepts a local directory containing config.yaml and model.safetensors. Weights are fp32; set bf16_scan: false in config.yaml for an fp32 scan (outputs will differ slightly from the samples above, which used the training-time bf16 scan).

To retrain from scratch under the identical protocol:

uv pip install --python .venv/bin/python datasets transformers
.venv/bin/python experiments/V1_train_tinystories_lm.py \
    --model wskan11 --scale 1m --dataset tinystories --seed 42 \
    --batch 64 --block 256 \
    --steps 100000 --lr 0.001 --lr-schedule cosine \
    --ckpt-every 25000 --eval-every 1000 --out-tag 3ep --compile --bf16

Limitations

  • Research artifact. These are 3.6k-10M parameter proof-of-concept models from an architecture study, not production LMs.
  • No factual reliability. All models confabulate names, dates, and facts; the larger ones merely do so more fluently. Never use outputs as a source of truth.
  • Not instruction-tuned. Even the UltraChat models only imitate conversational format; they do not follow instructions.
  • Short-range memory. The interpretation reports show the models' effective memory is local (dozens of bytes); long-range consistency is not to be expected.
  • English-only, single-domain. Each model knows only its training slice.
  • CUDA GPU required for inference. The recurrent scan is a fused Triton kernel; there is no CPU path. (config.yaml: architecture.bf16_scan: false switches the scan to fp32 on GPU.)
  • Sampling only. Generation is multinomial sampling; outputs vary with seed and temperature.

Provenance

  • Checkpoint: checkpoints/wskan11_tinystories_1m_3ep_s42/latest.pt (step 100,000) of the WSKAN repository, converted to safetensors (fp32) without any weight modification.
  • Training data: TinyStories, byte-encoded UTF-8.
  • Seed 42 is the canonical seed of the 5-seed campaign; the released sample outputs above used sampling seed 1234.
Downloads last month
18
Safetensors
Model size
1.01M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including llaa33219/wskan-1m-tinystories