modelv8: small research language models (v8 architecture family)

Base language models from the speed-check-ai-traning research project, one folder each. They are research models, not assistants: they continue text, with no instruction or chat tuning.

Folder Model Parameters Trained on Held-out result
fineweb-v83-71m v8.3, width 768, 6 blocks 71.3M (46.1M outside the embedding and output layers) FineWeb-Edu, 1.42B tokens, one pass loss 3.204, BLiMP grammar 77.8%
tinystories-v813-w512 v8.13 (v8.3 + NextLat training loss), width 512, 6 blocks 26.3M (24.7M used when generating) all of TinyStories V2 (GPT-4), about 500M tokens, one pass loss 1.1075

Each folder holds model.safetensors, config.json, tokenizer.json and its own README.md (architecture, training, results, limitations, an example and how to load it).

The architecture

Decoder-only, 6 blocks of forgetting attention (FoX, no position encoding) with a 64-token window in blocks 1, 2, 4, 5 and whole-text attention in blocks 3 and 6, Canon layers (4-token causal convolution), QK-norm, ReLU^2 MLP, value residual, U-net skips and soft-capped logits. Trained with Muon and a learning rate that is held and then decayed to zero over the last half of the steps.

Loading

The v8 architecture is custom, so loading needs the model code from the project (src/hrm_text). In each folder:

import json, torch
from safetensors.torch import load_file
from hrm_text.models.hrm import ModelConfig, build_model

config = ModelConfig.from_saved(json.load(open("config.json"))["model_config"])
model = build_model(config)
model.load_state_dict(load_file("model.safetensors"))
model.eval()
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train ModhiSathvik/modelv8