THETA Bench DreamZero5B: Simulation training, 3,003 segments
The validated final policy checkpoint is available in this repository.
This is a THETA training-result repository. It does not substitute an upstream pretrained policy for a THETA-trained checkpoint.
| Setting | Value |
|---|---|
| Training stage | Simulation training, 3,003 segments |
| Target optimizer updates | 40000 |
| Per-GPU batch / GPUs / global batch | 16 / 8 / 128 |
| Gradient accumulation | 1 |
| Conditions per global batch | 18 |
| Dataset revision | 8b2cd31e107b64cb13f812ea217a63a20845c78a |
The simulation pool contains 1,200 successful L1/L2 demonstrations and 1,803 extracted L0 prefixes, spanning 18 conditions. The 3,003 segments are not 3,003 independent demonstrations.
Use the model's native THETA adapter and model-specific dependencies. This repository does not claim compatibility with arbitrary Transformers or simulation loaders. No evaluation score is claimed by checkpoint publication.
Complete model-only trained DreamZero safetensors including frozen text/image/VAE tensors; all tensor keys, shapes, dtypes and bytes preserved, including learned THETA projection heads. Use load_policy.py:load_policy with the pinned RLinf, DreamZero and THETA source environments. It uses strict native loading and the source-checked THETA startup patch to skip overwritten base component reads. No DROID initialization exporter, projection reset, optimizer, or RNG is included. CPU export parity passed; GPU reload/evaluation is separate.
Training uses independent model optimizers and shared GPU execution through MPS. Publication is performed by a CPU uploader after final checkpoint validation.
- Downloads last month
- -