HD-PPO Agent playing CartPole-v1

This is a trained HD-PPO (Hyperdimensional Proximal Policy Optimization) agent playing CartPole-v1 using gradient-adaptive Fractional Power Encoding (FPE) with a prune-and-fine-tune pipeline.

Published by LTU-AI.

Pipeline

  1. Train a teacher at D=128 with gradient-adaptive single-beta FPE.
  2. Prune by actor-weight importance through D=128 → 32 → 16.
  3. Fine-tune each pruned checkpoint with PPO.

Published checkpoint: seed 123, compact D=16 model (held-out eval mean reward 500.00 ± 0.00).

Usage

Install dependencies:

pip install -r requirements.txt

Evaluate the local checkpoint:

python enjoy.py --weights hdppo-CartPole-v1/weights.npz --episodes 10

Render episodes:

python enjoy.py --weights hdppo-CartPole-v1/weights.npz --render --episodes 3

Record a replay video:

python record_video.py --weights hdppo-CartPole-v1/weights.npz --output replay.mp4

Load from Hugging Face Hub:

python enjoy.py --weights LTU-AI/hdppo-CartPole-v1 --episodes 10

Training pipeline

Reproduce the teacher → prune → fine-tune workflow:

python run_prune_finetune_128_32_16.py

Hyperparameters

{
    "env": "CartPole-v1",
    "algo": "HD-PPO (gradient-adaptive FPE, discrete)",
    "teacher_D": 128,
    "pruned_D": 16,
    "beta_base": 1.0,
    "timesteps_per_stage": 1000000,
    "rollout_steps": 1024,
    "actor_lr": 0.001,
    "critic_lr": 0.005,
    "seed": 123
}

Environment Arguments

{
    "render_mode": "rgb_array"
}

Model files

File Description
hdppo-CartPole-v1/weights.npz Published actor (+ critic if HD) and FPE encoder (D=16)
hdppo-CartPole-v1/weights_D128_teacher.npz Teacher checkpoint (D=128)
replay.mp4 Sample rollout video from the published min-D checkpoint
results.json Evaluation summary for the published checkpoint
results_D128_teacher.json Evaluation summary for the teacher
config.yml Training hyperparameters
train_hdppo.py / training modules Self-contained training code

Citation

If you use this model, please cite the HD-PPO / Hybrid-HD-PPO work.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results