YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

World-Model Post-Training Audit Artifacts

This repository contains selected fine-tuned checkpoints and cleaned training artifacts for Does Learning to Predict the World Help Agents Act? Auditing World-Model Post-Training.

Contents

  • weights/: selected final checkpoints for ALFWorld, ScienceWorld, and VisualWebArena.
  • data/alfworld/: 20,000-sample GT and mismatched-target training parquet files.
  • data/sciworld/: 20,000-sample GT and mismatched-target training parquet files.
  • data/vwa/: VisualWebArena COIN training samples and the 201-task evaluation ID list.

Evaluation results are intentionally omitted from this repository.

Cleaning applied

  • Removed the absolute ALFWorld gamefile path from extra_info; prompts and training targets are unchanged.
  • Preserved the generated VisualWebArena extra_info provenance fields (site, URLs, action, run and episode identifiers, step index, and snapshot index); these fields do not identify the authors. Prompts and images are unchanged.
  • tasks_eval.jsonl keeps only site and task_id; it omits natural-language intents, difficulty labels, and evaluation-pool metadata while preserving the full 201-task evaluation set.
  • Retained teacher_prompt_ids for OPSD-compatible training.
  • Retained true_next_state in mismatched-target files for auditability.

The fine-tuned weights inherit the terms of their base models. Benchmark data and derived artifacts remain subject to the upstream benchmark licenses and terms; please consult the cited upstream projects before redistribution.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support