Standard attention ablation model (1.50M params, 3072 context)

Part of a controlled ablation: BananaMind/dsa-model-ablation (DSA) vs BananaMind/normal-model-ablation (standard attention). Same seed, data order, tokenizer, architecture and hyper-parameters. Only the attention differs.

  • Attention: Standard dense causal multi-head attention.
  • Architecture: 5 layers, hidden 128, 4 heads, SwiGLU 336, RoPE, RMSNorm, tied embeddings, vocab 4096 (custom BPE)
  • Context length: 3072
  • Params: 1,498,496
  • Data: HuggingFaceFW/fineweb-edu (sample-10BT), 2000M tokens, 1 epoch, seed 1337
  • Optimiser: AdamW lr 0.002 (cosine), batch 16x3072 tokens, bf16 autocast

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "BananaMind/normal-model-ablation"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True)

(No KV cache is implemented; generate recomputes the full sequence each step.)

Benchmarks (0-shot, lm-evaluation-harness; acc / acc_norm where available)

Model Params Val loss Val ppl piqa hellaswag arc_easy arc_challenge
DSA 1,530,496 3.4150 30.42 54.90 / 53.05 26.84 / 26.72 30.18 / 31.02 16.98 / 20.73
Standard 1,498,496 3.4248 30.72 53.81 / 53.75 26.86 / 26.75 30.68 / 31.44 17.83 / 20.99

Per-model raw results: benchmarks/results.json. Comparison with the other model: benchmarks/comparison.md (other repo: BananaMind/dsa-model-ablation). Training curve: train_log.json.

Note: at this size models are close to chance on these benchmarks, and the benchmark prompts are shorter than the top-k (512), where DSA selects every token and behaves identically to dense attention at inference, so differences mostly reflect how training differed.

Downloads last month
328
Safetensors
Model size
1.5M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support