modelv8: small research language models (v8 architecture family)
Base language models from the speed-check-ai-traning research project, one folder each. They are research models, not assistants: they continue text, with no instruction or chat tuning.
| Folder | Model | Parameters | Trained on | Held-out result |
|---|---|---|---|---|
fineweb-v83-71m |
v8.3, width 768, 6 blocks | 71.3M (46.1M outside the embedding and output layers) | FineWeb-Edu, 1.42B tokens, one pass | loss 3.204, BLiMP grammar 77.8% |
tinystories-v813-w512 |
v8.13 (v8.3 + NextLat training loss), width 512, 6 blocks | 26.3M (24.7M used when generating) | all of TinyStories V2 (GPT-4), about 500M tokens, one pass | loss 1.1075 |
Each folder holds model.safetensors, config.json, tokenizer.json and its own README.md (architecture,
training, results, limitations, an example and how to load it).
The architecture
Decoder-only, 6 blocks of forgetting attention (FoX, no position encoding) with a 64-token window in blocks 1, 2, 4, 5 and whole-text attention in blocks 3 and 6, Canon layers (4-token causal convolution), QK-norm, ReLU^2 MLP, value residual, U-net skips and soft-capped logits. Trained with Muon and a learning rate that is held and then decayed to zero over the last half of the steps.
Loading
The v8 architecture is custom, so loading needs the model code from the project (src/hrm_text). In each
folder:
import json, torch
from safetensors.torch import load_file
from hrm_text.models.hrm import ModelConfig, build_model
config = ModelConfig.from_saved(json.load(open("config.json"))["model_config"])
model = build_model(config)
model.load_state_dict(load_file("model.safetensors"))
model.eval()