Dense Middle · 336B tokens
Checkpoint of the dense_middle recipe at the 336B-token endpoint (optimizer step 80,000),
from Towards Looped Models Done Right. Part I: Topology, Input Injection, Recurrent-State Organization.
Loading this checkpoint requires the xLLM code. The weights are stored in xLLM's native format (BF16 Safetensors with an xLLM
config.jsonandartifact_manifest.json). This is not a Hugging Facetransformerscheckpoint:AutoModel.from_pretrainedcannot load it. Get the code at https://github.com/ifm-ai/xllm-loop and load the directory withxllm.paper_part1.native_inference.load_native_model.
Model details
| Field | Value |
|---|---|
| Recipe | dense_middle |
| Family | Dense middle loop |
Architecture (model.arch) |
huginn |
| Endpoint | 336b: optimizer step 80,000, 335,544,320,000 training tokens |
| Parameters (resident) | 1,496,667,648 |
| Parameters (activated per token) | 1,496,667,648 |
| Training seed | 1 |
| Optimizer | AdamW; weight decay 0.1 |
| LR schedule | cosine; 200-step warmup; 119,210-step horizon (the 500B endpoint) |
| Global batch | 512 sequences × 8,192 tokens = 4,194,304 tokens/step |
| Weights | BF16 Safetensors, 1 shard of at most 5 GiB |
| Tokenizer | tokenizer/, derived from IFM/K2-Horizon-3.7B (see NOTICE) |
All Part I endpoints share one 119,210-step learning-rate schedule, so this 336B-token checkpoint is taken mid-schedule, before the cosine decay completes at the 500B-token endpoint.
Download
hf download IFM/LoopedLM-P1-dense-middle-336b --local-dir models/dense_middle/336b
The xLLM loader checks the directory against artifact_manifest.json: it rejects symbolic links
and files the manifest does not list, apart from the .gitattributes file and the
.cache/huggingface/ folder that hf download --local-dir adds. Download into a directory as
above, not into the Hub cache (~/.cache/huggingface/hub), whose files are symbolic links.
Use
Native inference requires a CUDA GPU and FlashAttention 3; the xLLM repository documents installation.
export ENABLE_FLASH_ATTENTION_3=true
from xllm.paper_part1.native_inference import generate_native, load_native_model
model, tokenizer, config = load_native_model("models/dense_middle/336b")
tokens = generate_native(model, tokenizer, ["The capital of France is"],
max_gen_len=32, use_sampling=False)
print(tokenizer.decode(tokens[0]))
The loader verifies the directory against artifact_manifest.json, then strictly loads tensor
names, shapes and dtypes. Evaluate with the paper protocol:
ENABLE_FLASH_ATTENTION_3=true python eval_paper_part1.py --artifact models/dense_middle/336b \
--tasks-root /path/to/composite-eval-root --output /path/to/evaluation-output \
--protocol paper-part1
Retrain from scratch with the same recipe (the base config holds the data, tokenizer, logging and checkpoint-retention settings):
ENABLE_FLASH_ATTENTION_3=true python train_paper_part1.py --recipe dense_middle --target 336b \
--base-config /path/to/base.json --dump-dir /path/to/new-run
Files
config.json: the recipe'smodelfields with the training checkpoint's values, and thetokenizersettings withtokenizer.pathset totokenizer/.model.safetensors.index.jsonandmodel-*-of-*.safetensors: BF16 weights; no tensor is split across shards.tokenizer/:tokenizer.json,tokenizer_config.json,special_tokens_map.json.artifact_manifest.json: size and SHA-256 of every file.LICENSE,NOTICE.
Optimizer, dataloader and training state are not included; resuming a run mid-training needs the original training checkpoint.
Paper and citation
Towards Looped Models Done Right. Part I: Topology, Input Injection, Recurrent-State Organization: https://huskydoge.github.io/husky-blog/posts/recursive_models/towards-looped-models-done-right/
@misc{huang2026loopedmodels,
title = {Towards Looped Models Done Right. Part I: Topology, Input Injection, Recurrent-State Organization},
author = {Benhao Huang and Chufan Shi and Junlin Chen and Shicheng Wen and Zhengzhong Liu and Eric Xing and Xuezhe Ma},
year = {2026},
url = {https://huskydoge.github.io/husky-blog/posts/recursive_models/towards-looped-models-done-right/}
}
License
The weights and tokenizer are released under the Apache License 2.0 (LICENSE). NOTICE
records the tokenizer's attribution to
IFM/K2-Horizon-3.7B
and its modifications. The xLLM code is distributed under its own license.
- Downloads last month
- 122