Dense Middle · 336B tokens

Checkpoint of the dense_middle recipe at the 336B-token endpoint (optimizer step 80,000), from Towards Looped Models Done Right. Part I: Topology, Input Injection, Recurrent-State Organization.

Loading this checkpoint requires the xLLM code. The weights are stored in xLLM's native format (BF16 Safetensors with an xLLM config.json and artifact_manifest.json). This is not a Hugging Face transformers checkpoint: AutoModel.from_pretrained cannot load it. Get the code at https://github.com/ifm-ai/xllm-loop and load the directory with xllm.paper_part1.native_inference.load_native_model.

Model details

Field Value
Recipe dense_middle
Family Dense middle loop
Architecture (model.arch) huginn
Endpoint 336b: optimizer step 80,000, 335,544,320,000 training tokens
Parameters (resident) 1,496,667,648
Parameters (activated per token) 1,496,667,648
Training seed 1
Optimizer AdamW; weight decay 0.1
LR schedule cosine; 200-step warmup; 119,210-step horizon (the 500B endpoint)
Global batch 512 sequences × 8,192 tokens = 4,194,304 tokens/step
Weights BF16 Safetensors, 1 shard of at most 5 GiB
Tokenizer tokenizer/, derived from IFM/K2-Horizon-3.7B (see NOTICE)

All Part I endpoints share one 119,210-step learning-rate schedule, so this 336B-token checkpoint is taken mid-schedule, before the cosine decay completes at the 500B-token endpoint.

Download

hf download IFM/LoopedLM-P1-dense-middle-336b --local-dir models/dense_middle/336b

The xLLM loader checks the directory against artifact_manifest.json: it rejects symbolic links and files the manifest does not list, apart from the .gitattributes file and the .cache/huggingface/ folder that hf download --local-dir adds. Download into a directory as above, not into the Hub cache (~/.cache/huggingface/hub), whose files are symbolic links.

Use

Native inference requires a CUDA GPU and FlashAttention 3; the xLLM repository documents installation.

export ENABLE_FLASH_ATTENTION_3=true
from xllm.paper_part1.native_inference import generate_native, load_native_model

model, tokenizer, config = load_native_model("models/dense_middle/336b")
tokens = generate_native(model, tokenizer, ["The capital of France is"],
                         max_gen_len=32, use_sampling=False)
print(tokenizer.decode(tokens[0]))

The loader verifies the directory against artifact_manifest.json, then strictly loads tensor names, shapes and dtypes. Evaluate with the paper protocol:

ENABLE_FLASH_ATTENTION_3=true python eval_paper_part1.py --artifact models/dense_middle/336b \
    --tasks-root /path/to/composite-eval-root --output /path/to/evaluation-output \
    --protocol paper-part1

Retrain from scratch with the same recipe (the base config holds the data, tokenizer, logging and checkpoint-retention settings):

ENABLE_FLASH_ATTENTION_3=true python train_paper_part1.py --recipe dense_middle --target 336b \
    --base-config /path/to/base.json --dump-dir /path/to/new-run

Files

  • config.json: the recipe's model fields with the training checkpoint's values, and the tokenizer settings with tokenizer.path set to tokenizer/.
  • model.safetensors.index.json and model-*-of-*.safetensors: BF16 weights; no tensor is split across shards.
  • tokenizer/: tokenizer.json, tokenizer_config.json, special_tokens_map.json.
  • artifact_manifest.json: size and SHA-256 of every file.
  • LICENSE, NOTICE.

Optimizer, dataloader and training state are not included; resuming a run mid-training needs the original training checkpoint.

Paper and citation

Towards Looped Models Done Right. Part I: Topology, Input Injection, Recurrent-State Organization: https://huskydoge.github.io/husky-blog/posts/recursive_models/towards-looped-models-done-right/

@misc{huang2026loopedmodels,
  title  = {Towards Looped Models Done Right. Part I: Topology, Input Injection, Recurrent-State Organization},
  author = {Benhao Huang and Chufan Shi and Junlin Chen and Shicheng Wen and Zhengzhong Liu and Eric Xing and Xuezhe Ma},
  year   = {2026},
  url    = {https://huskydoge.github.io/husky-blog/posts/recursive_models/towards-looped-models-done-right/}
}

License

The weights and tokenizer are released under the Apache License 2.0 (LICENSE). NOTICE records the tokenizer's attribution to IFM/K2-Horizon-3.7B and its modifications. The xLLM code is distributed under its own license.

Downloads last month
122
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including IFM/LoopedLM-P1-dense-middle-336b