vlabki/rr-speed-v4

Self-contained recurrent player-policy checkpoint.

  • Source directory: best_reliable
  • PPO update: 18250
  • Environment steps: 112128000
  • Action support: bc

Evaluation values below come from deterministic fixed-seed argmax runs. Frame averages exclude DNF races.

Best reliable evaluation

Metric Value
Checkpoint update 18250
Cohort 20
Finish rate 95.0%
Finished mean frames 10706.58
Finished median frames 10697.00
Fastest finish frames 10645
Finished P90 frames 10726.60
First-place rate among finishes 100.0%
Mean wall events 0.65
Mean respawns 0.00

Best fastest evaluation

Metric Value
Checkpoint update 17000
Cohort 20
Finish rate 75.0%
Finished mean frames 10687.47
Finished median frames 10681.00
Fastest finish frames 10633
Finished P90 frames 10712.60
First-place rate among finishes 100.0%
Mean wall events 1.30
Mean respawns 0.20

Files

Runtime weights, model config, normalization statistics, route reference, training config, and portable checksummed fallbacks are included. Raw rollout traces, optimizer state, and full training logs are excluded.

Downloads last month
28
Safetensors
Model size
575k params
Tensor type
F32
·
Video Preview
loading