LoopWAM final checkpoints
Final trained policy weights from the LoopWAM experiments. This repository contains 17 checkpoints, separated by training data scope. Both original and repeated 4/4 runs are retained. All production checkpoints are from completed ten-epoch runs with global batch 128 and training seed 42.
October 9 additions: all-loop KV
| Directory | KV mode | Seed42/43/44 | Pooled SR |
|---|---|---|---|
| libero-long/v0-video4-action1-kv-concat | concat | 83/86/86 | 255/300 = 85.00% |
| libero-long/v0-video4-action1-kv-mix | mix | 90/89/87 | 266/300 = 88.67% |
These two additions require the KV_concat code branch, revision3a89bfac5cfccad66bfebbe00014f7f739a78adc or compatible later.
Their checkpoints record action_kv_mode; the original LoopWAM_NT code predates
these modes. Use eager BF16 inference, as native Inductor failed numerical
equivalence. Each directory includes normalization, training provenance,
three evaluation summaries and a tensor-verified optimizer-free native policy.
The Long aligned4/1 and both original/repeat4/4 weights were hash-checked and
are unchanged. Their additional evaluation seeds do not create new weights.
The full-suite v0 4/1 training-only checkpoint is now included at libero-all-suites/v0-video4-action1. It has not been evaluated after training.
Expanded seed42: concat404/500 versus aligned413/500 (paired p=0.4814). The original three-seed decision remains inconclusive. This grid overlaps the original states and must not be pooled as independent evidence. See the updated results report.
October 10 addition: full-suite aligned 1/4
libero-all-suites/v0-video1-action4 contains the final ten-epoch full-suite v0 aligned 1/4 policy (21,700 updates, global batch 128, training seed 42). Training took 10h 58m 46s on two H100s. Not evaluated: this run finished under the requested training-only workflow. Normalization, training metadata, fairness checks and export hashes are included.
October 10 addition: full-suite 4/1
| Directory | Model | Video/action loops | Parameters | Evaluation |
|---|---|---|---|---|
| libero-all-suites/v0-video4-action1 | v0 | 4/1 | 584,536,135 | Training only; evaluation not run |
The policy completed ten epochs on the four-suite full dataset with global batch 128. Evaluation results are not available for this checkpoint.
October 10 addition: full-suite 2/2
| Directory | Model | Video/action loops | Parameters | Evaluation |
|---|---|---|---|---|
| libero-all-suites/v0-video2-action2 | v0 | 2/2 | 584,536,135 | Training only; evaluation not run |
The checkpoint completed ten epochs on the four-suite full dataset with global batch 128. Evaluation has not been run.
October 11 addition: full-suite mix 4/1
| Directory | Model | Video/action loops | Parameters | Evaluation |
|---|---|---|---|---|
| libero-all-suites/v0-video4-action1-kv-mix | v0 | 4/1 mix | 584,536,423 | Training only; evaluation not run |
The checkpoint completed ten epochs on the four-suite full dataset with global batch 128. Evaluation has not been run.
Checkpoint index
| Directory | Model | Video/action loops | Parameters | Final evaluation |
|---|---|---|---|---|
| libero-all-suites/v0-video1-action4 | v0 aligned | 1/4 | 584536135 | Not evaluated (training only) |
| libero-long/v0-video4-action4-original | v0 | 4/4 | 584,536,135 | 96/100 (96%) |
| libero-long/v1-video4-action4 | v1 | 4/4 | 584,536,135 | 82/100 (82%) |
| libero-long/dense-s12 | dense_s12 | 1/1 | 584,536,135 | 81/100 (81%) |
| libero-long/v2-video4-action4 | v2 | 4/4 | 584,536,135 | 81/100 (81%) |
| libero-long/dense-s30 | dense_s30 | 1/1 | 1,416,114,247 | 86/100 (86%) |
| libero-all-suites/v0-video4-action4 | v0 | 4/4 | 584,536,135 | 388/400 (97%) |
| libero-long/v0-video4-action1 | v0 | 4/1 | 584,536,135 | 81/100 (81%) |
| libero-long/v0-video4-action2 | v0 | 4/2 | 584,536,135 | 86/100 (86%) |
| libero-long/v0-video2-action2 | v0 | 2/2 | 584,536,135 | 81/100 (81%) |
| libero-long/v0-video1-action4 | v0 | 1/4 | 584,536,135 | 86/100 (86%) |
| libero-long/v0-video4-action4-repeat | v0 | 4/4 | 584,536,135 | 91/100 (91%) |
libero-long/: matched 344-training / 44-validation demonstration split, 7,250 updates. Dense-S12/S30 execute their independent blocks once; their 1/1 metadata does not mean the dense networks have equal depth. S12 has 12 independent block pairs; S30 has 30. Loop models use three prelude, six shared core and three coda pairs, with 3 + 6K + 3 effective block applications per expert.
libero-all-suites/: one v0 trained on all 1,712 locally available Spatial/Object/Goal/Long demonstrations, 21,700 updates. Its 388/400 result aggregates all four suites, so it is a different training/evaluation scope from the Long-only rows. The same four-suite checkpoint also scored 720/2,000 (36%) on LIBERO-Pro.
The original checkpoint-index rows below use 100 final-checkpoint rollouts at evaluation seed 42; the new KV table above reports three seeds. Standard-suite evaluations use 700 policy steps, 30 settling steps, 32-action predictions, replanning every ten actions, ten denoising steps and CFG 1. The original 4/4 checkpoint additionally scored 96%, 88%, 91% across evaluation seeds 42/43/44. These are one training seed and cannot establish across-training-seed significance.
Files and integrity
Each directory contains policy.pt, dataset_stats.json, data_manifest.json, training_manifest.json, training_timing.json, evaluation_summary.json, and export.json. policy.pt is the native loopwam-s-v1 checkpoint, containing FP32 model tensors, architecture/depth metadata and the training contract. Adam optimizer states are omitted; use it for inference or fresh-optimizer fine-tuning, not exact optimizer-state resume. All exported model tensors were checked for exact equality with the evaluated final checkpoint. export.json records both the original checkpoint SHA-256 and the exported file SHA-256; these hashes differ because optimizer removal changes serialization.
The frozen VAE and text encoder/embeddings are external inference dependencies. The native loader does not require the donor initialization artifact when checkpoint_path is provided. Normalization files are specific to each checkpoint and must remain matched. Videos, dataset videos and original optimizer checkpoints remain on the training server.
Loading
Use the compatible LoopWAM code, at revision 8a29ffdce1537409e2b7e975e838d990ebbe068a or a compatible later revision. This is a custom policy, not a Transformers AutoModel checkpoint. Install that repository’s inference dependencies first. The repository is publicly readable.
import torch
from huggingface_hub import hf_hub_download
from fastwam.models.wan22.loopwam import create_loopwam
repo = "anhdao69/LoopWAM_NT"
variant = "libero-long/v0-video4-action4-repeat"
checkpoint = hf_hub_download(repo, f"{variant}/policy.pt")
stats = hf_hub_download(repo, f"{variant}/dataset_stats.json")
vae = hf_hub_download("Wan-AI/Wan2.1-T2V-1.3B", "Wan2.1_VAE.pth")
model = create_loopwam(checkpoint_path=checkpoint, vae_path=vae,
model_dtype=torch.float32, device="cuda").eval()
Depth/version are reconstructed from checkpoint metadata. Use the repository’s observation adapter, matched normalization, text embeddings with their padding masks, and BF16 autocast during prediction. See scripts/evaluate_loopwam_libero.py for the complete simulator pipeline. Evaluation summaries retain the original checkpoint hashes; use the tensor equality and source/export mapping in export.json when auditing exports. The evaluator computes the exported file’s own hash for a new evaluation.
Provenance and rights
These research checkpoints derive from pretrained Wan and the FastWAM/LoopWAM implementation. Consult the upstream model and code licenses before redistribution or use; this model card does not grant additional upstream rights. The repository is currently public.
For detailed comparison caveats, runtime accounting and per-task results, see the consolidated report. That report’s earlier snapshot labels the repeat pending; the repeat is now complete at 91/100, as recorded here.