YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
World-Model Post-Training Audit Artifacts
This repository contains selected fine-tuned checkpoints and cleaned training artifacts for Does Learning to Predict the World Help Agents Act? Auditing World-Model Post-Training.
Contents
weights/: selected final checkpoints for ALFWorld, ScienceWorld, and VisualWebArena.data/alfworld/: 20,000-sample GT and mismatched-target training parquet files.data/sciworld/: 20,000-sample GT and mismatched-target training parquet files.data/vwa/: VisualWebArena COIN training samples and the 201-task evaluation ID list.
Evaluation results are intentionally omitted from this repository.
Cleaning applied
- Removed the absolute ALFWorld
gamefilepath fromextra_info; prompts and training targets are unchanged. - Preserved the generated VisualWebArena
extra_infoprovenance fields (site, URLs, action, run and episode identifiers, step index, and snapshot index); these fields do not identify the authors. Prompts and images are unchanged. tasks_eval.jsonlkeeps onlysiteandtask_id; it omits natural-language intents, difficulty labels, and evaluation-pool metadata while preserving the full 201-task evaluation set.- Retained
teacher_prompt_idsfor OPSD-compatible training. - Retained
true_next_statein mismatched-target files for auditability.
The fine-tuned weights inherit the terms of their base models. Benchmark data and derived artifacts remain subject to the upstream benchmark licenses and terms; please consult the cited upstream projects before redistribution.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support