d4-m
Table 4, M, D4 of Towards Looped Models Done Right. Part II: Rethinking at Fixed Points.
Loading this checkpoint requires the xLLM code. The weights are stored in xLLM's native format (BF16 Safetensors with an xLLM
config.jsonandartifact_manifest.json). This is not a Hugging Facetransformerscheckpoint:AutoModel.from_pretrainedcannot load it. Get the code at https://github.com/ifm-ai/xllm-loop and load the directory withxllm.paper_part2.artifacts.load_artifact.
Model details
| Model | dense Transformer of 4 blocks, untied output layer; width 3,072, 48 heads (12 KV heads), SwiGLU FFN 8,704 |
| Parameters | 810,052,608, including the embedding and output layers |
| Evaluation | full KV cache |
| Training | recipe m_d4: 20,480 updates of 512 x 8,192 tokens (85.90B), AdamW, warmup-stable-decay |
| Weights | BF16 Safetensors: the training checkpoint's FP32 weights rounded to BF16 |
| Tokenizer | tokenizer/, the Jais64k tokenizer of training |
Download
hf download IFM/LoopedLM-P2-d4-m --local-dir d4-m
The xLLM loader checks the directory against artifact_manifest.json: it rejects symbolic links
and files the manifest does not list, apart from the .gitattributes file and the
.cache/huggingface/ folder that hf download --local-dir adds. Download into a directory as
above, not into the Hub cache (~/.cache/huggingface/hub), whose files are symbolic links.
Use
Evaluate with eval_paper_part2.py from the xLLM repository:
ENABLE_FLASH_ATTENTION_3=true python eval_paper_part2.py --artifact d4-m \
--data /path/to/eval-data/data.json --out out ppl
data.json and the evaluation inputs come from release/paper-part2/prepare-eval-data.py --output /path/to/eval-data.
xllm.paper_part2.artifacts.load_artifact("d4-m") verifies the directory and returns the model
config fields and the BF16 state dict on CPU.
artifact_manifest.json records the size and SHA-256 of every file in this repository.
Paper and citation
Towards Looped Models Done Right. Part II: Rethinking at Fixed Points: https://github.com/ifm-ai/xllm-loop/blob/main/papers/part2.pdf
@misc{huang2026fixedpoints,
title = {Towards Looped Models Done Right. Part II: Rethinking at Fixed Points},
author = {Benhao Huang and Chufan Shi and Junlin Chen and Shicheng Wen and Zhengzhong Liu and Eric Xing and Xuezhe Ma},
year = {2026},
url = {https://github.com/ifm-ai/xllm-loop/blob/main/papers/part2.pdf}
}
License
The weights are released under the Apache License 2.0 (LICENSE); NOTICE records the
tokenizer's attribution and how the artifact was prepared. The xLLM code is distributed under its
own license.
- Downloads last month
- 145