LoopWAM final checkpoints

Final trained policy weights from the LoopWAM experiments. This repository contains 17 checkpoints, separated by training data scope. Both original and repeated 4/4 runs are retained. All production checkpoints are from completed ten-epoch runs with global batch 128 and training seed 42.

October 9 additions: all-loop KV

Directory KV mode Seed42/43/44 Pooled SR
libero-long/v0-video4-action1-kv-concat concat 83/86/86 255/300 = 85.00%
libero-long/v0-video4-action1-kv-mix mix 90/89/87 266/300 = 88.67%

These two additions require the KV_concat code branch, revision3a89bfac5cfccad66bfebbe00014f7f739a78adc or compatible later. Their checkpoints record action_kv_mode; the original LoopWAM_NT code predates these modes. Use eager BF16 inference, as native Inductor failed numerical equivalence. Each directory includes normalization, training provenance, three evaluation summaries and a tensor-verified optimizer-free native policy. The Long aligned4/1 and both original/repeat4/4 weights were hash-checked and are unchanged. Their additional evaluation seeds do not create new weights. The full-suite v0 4/1 training-only checkpoint is now included at libero-all-suites/v0-video4-action1. It has not been evaluated after training.

Expanded seed42: concat404/500 versus aligned413/500 (paired p=0.4814). The original three-seed decision remains inconclusive. This grid overlaps the original states and must not be pooled as independent evidence. See the updated results report.

October 10 addition: full-suite aligned 1/4

libero-all-suites/v0-video1-action4 contains the final ten-epoch full-suite v0 aligned 1/4 policy (21,700 updates, global batch 128, training seed 42). Training took 10h 58m 46s on two H100s. Not evaluated: this run finished under the requested training-only workflow. Normalization, training metadata, fairness checks and export hashes are included.

October 10 addition: full-suite 4/1

Directory Model Video/action loops Parameters Evaluation
libero-all-suites/v0-video4-action1 v0 4/1 584,536,135 Training only; evaluation not run

The policy completed ten epochs on the four-suite full dataset with global batch 128. Evaluation results are not available for this checkpoint.

October 10 addition: full-suite 2/2

Directory Model Video/action loops Parameters Evaluation
libero-all-suites/v0-video2-action2 v0 2/2 584,536,135 Training only; evaluation not run

The checkpoint completed ten epochs on the four-suite full dataset with global batch 128. Evaluation has not been run.

October 11 addition: full-suite mix 4/1

Directory Model Video/action loops Parameters Evaluation
libero-all-suites/v0-video4-action1-kv-mix v0 4/1 mix 584,536,423 Training only; evaluation not run

The checkpoint completed ten epochs on the four-suite full dataset with global batch 128. Evaluation has not been run.

Checkpoint index

Directory Model Video/action loops Parameters Final evaluation
libero-all-suites/v0-video1-action4 v0 aligned 1/4 584536135 Not evaluated (training only)
libero-long/v0-video4-action4-original v0 4/4 584,536,135 96/100 (96%)
libero-long/v1-video4-action4 v1 4/4 584,536,135 82/100 (82%)
libero-long/dense-s12 dense_s12 1/1 584,536,135 81/100 (81%)
libero-long/v2-video4-action4 v2 4/4 584,536,135 81/100 (81%)
libero-long/dense-s30 dense_s30 1/1 1,416,114,247 86/100 (86%)
libero-all-suites/v0-video4-action4 v0 4/4 584,536,135 388/400 (97%)
libero-long/v0-video4-action1 v0 4/1 584,536,135 81/100 (81%)
libero-long/v0-video4-action2 v0 4/2 584,536,135 86/100 (86%)
libero-long/v0-video2-action2 v0 2/2 584,536,135 81/100 (81%)
libero-long/v0-video1-action4 v0 1/4 584,536,135 86/100 (86%)
libero-long/v0-video4-action4-repeat v0 4/4 584,536,135 91/100 (91%)

libero-long/: matched 344-training / 44-validation demonstration split, 7,250 updates. Dense-S12/S30 execute their independent blocks once; their 1/1 metadata does not mean the dense networks have equal depth. S12 has 12 independent block pairs; S30 has 30. Loop models use three prelude, six shared core and three coda pairs, with 3 + 6K + 3 effective block applications per expert.

libero-all-suites/: one v0 trained on all 1,712 locally available Spatial/Object/Goal/Long demonstrations, 21,700 updates. Its 388/400 result aggregates all four suites, so it is a different training/evaluation scope from the Long-only rows. The same four-suite checkpoint also scored 720/2,000 (36%) on LIBERO-Pro.

The original checkpoint-index rows below use 100 final-checkpoint rollouts at evaluation seed 42; the new KV table above reports three seeds. Standard-suite evaluations use 700 policy steps, 30 settling steps, 32-action predictions, replanning every ten actions, ten denoising steps and CFG 1. The original 4/4 checkpoint additionally scored 96%, 88%, 91% across evaluation seeds 42/43/44. These are one training seed and cannot establish across-training-seed significance.

Files and integrity

Each directory contains policy.pt, dataset_stats.json, data_manifest.json, training_manifest.json, training_timing.json, evaluation_summary.json, and export.json. policy.pt is the native loopwam-s-v1 checkpoint, containing FP32 model tensors, architecture/depth metadata and the training contract. Adam optimizer states are omitted; use it for inference or fresh-optimizer fine-tuning, not exact optimizer-state resume. All exported model tensors were checked for exact equality with the evaluated final checkpoint. export.json records both the original checkpoint SHA-256 and the exported file SHA-256; these hashes differ because optimizer removal changes serialization.

The frozen VAE and text encoder/embeddings are external inference dependencies. The native loader does not require the donor initialization artifact when checkpoint_path is provided. Normalization files are specific to each checkpoint and must remain matched. Videos, dataset videos and original optimizer checkpoints remain on the training server.

Loading

Use the compatible LoopWAM code, at revision 8a29ffdce1537409e2b7e975e838d990ebbe068a or a compatible later revision. This is a custom policy, not a Transformers AutoModel checkpoint. Install that repository’s inference dependencies first. The repository is publicly readable.

import torch
from huggingface_hub import hf_hub_download
from fastwam.models.wan22.loopwam import create_loopwam

repo = "anhdao69/LoopWAM_NT"
variant = "libero-long/v0-video4-action4-repeat"
checkpoint = hf_hub_download(repo, f"{variant}/policy.pt")
stats = hf_hub_download(repo, f"{variant}/dataset_stats.json")
vae = hf_hub_download("Wan-AI/Wan2.1-T2V-1.3B", "Wan2.1_VAE.pth")
model = create_loopwam(checkpoint_path=checkpoint, vae_path=vae,
                       model_dtype=torch.float32, device="cuda").eval()

Depth/version are reconstructed from checkpoint metadata. Use the repository’s observation adapter, matched normalization, text embeddings with their padding masks, and BF16 autocast during prediction. See scripts/evaluate_loopwam_libero.py for the complete simulator pipeline. Evaluation summaries retain the original checkpoint hashes; use the tensor equality and source/export mapping in export.json when auditing exports. The evaluator computes the exported file’s own hash for a new evaluation.

Provenance and rights

These research checkpoints derive from pretrained Wan and the FastWAM/LoopWAM implementation. Consult the upstream model and code licenses before redistribution or use; this model card does not grant additional upstream rights. The repository is currently public.

For detailed comparison caveats, runtime accounting and per-task results, see the consolidated report. That report’s earlier snapshot labels the repeat pending; the repeat is now complete at 91/100, as recorded here.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading