GDN2-MLP 1.3B โ€” 152B-token control

This is the matched GDN2-MLP control for Emender E97 1.3B. It is a 1.287B-parameter recurrent base language model trained for 152,280,498,176 tokens. It is not a 152B-parameter model.

The model uses NVIDIA Gated DeltaNet 2 token mixing plus a post-mixer SwiGLU MLP, reshaped by CMA-ES to match the E97 parameter budget. Both controls used the same Pile/p50k_base stream, 2,048-token chunks, BF16 Schedule-Free AdamW, and eight-island DiLoCo harness.

Artifact

  • Final step: 2,323,616
  • Training-log last-100 loss: 2.4034 nats/token
  • Source checkpoint: 7,720,577,595 bytes
  • Source checkpoint SHA-256: ffae61cfeeeff820f469f7d66dd31c60e48c455872ff9102e57b15ad75bad59b
  • Export: BF16 Schedule-Free train/y weights
  • Tokenizer: p50k_base
  • Raw pickle checkpoint and optimizer state: not included

At the last common regular log point (150,793,420,800 tokens), the 80-point moving averages were E97 2.437045 and GDN2 2.426705; the difference was -0.010340 nats/token (GDN2 minus E97). These are training-log summaries, not a replacement for fixed held-out evaluation.

Loading

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("spinozans/gdn2-mlp-1.3b")
model = AutoModelForCausalLM.from_pretrained(
    "spinozans/gdn2-mlp-1.3b", trust_remote_code=True, dtype="auto"
)

The bundled remote code is a standalone portable PyTorch implementation of the published GDN2 recurrence. It does not redistribute NVIDIA source. It favors portability and exact weight loading over fused-kernel throughput.

The original fused GDN2 implementation used for training is available from NVIDIA GatedDeltaNet-2 at commit 95709fc250357c2dd109361c353192f2aa5913f9 under NVIDIA's Source Code License-NC. The Emender training wrapper is recorded at commit a3a862f30f5c3c9584e18490c986bc9a065d6653.

Validation

  • all 267 exported tensors match recovered source-checkpoint train/y weights exactly;
  • the Transformers loader reports zero missing, unexpected, or mismatched keys;
  • portable source-checkpoint and exported CPU logits are bit-identical;
  • CUDA fused and portable paths agree on greedy argmax for The theorem states (measured mean/max absolute logit delta 0.01244 / 0.125 from BF16 kernel evaluation order);
  • live-style greedy generation produces a finite continuation.

Machine-readable evidence is included in validation.json.

Limitations and license

This is a raw base LM, not an instruction or chat model. It may emit incorrect, biased, or unsafe text. Benchmark coverage is limited. No standalone model license has been selected; license: other is intentional. The separately available NVIDIA fused implementation has its own non-commercial source license, which is not replaced by this model card.

Downloads last month
285
Safetensors
Model size
1B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support