SuperMarioBros-Nes-v0 β€” Level1-2 β€” PPO

Stable-Baselines3 PPO policy for SuperMarioBros-Nes-v0 Level1-2, trained and evaluated with rlab.

At a Glance

Item Value
Task Complete SuperMarioBros-Nes-v0 Level1-2
Provider supermariobrosnes-turbo
Algorithm ppo
Checkpoint Step 7500000
Evaluation stochastic full evaluation, 100 episodes
Success minimum 97.0%, mean 97.0%
Mean return 3077.086
Release v1
Preview Root replay.mp4
YouTube Watch on YouTube

Quick Start

git clone https://github.com/tsilva/rlab
cd rlab
git checkout 40b8613743e41d36489701888cafcee24306604e
uv sync --frozen

Import the ROM, then play or evaluate the immutable checkpoint:

uv run rlab import-roms ~/roms --game SuperMarioBros-Nes-v0
uv run rlab play https://huggingface.co/tsilva/Level1-2_stable-baselines3-ppo_91df110f/resolve/v1/model.zip
uv run rlab eval https://huggingface.co/tsilva/Level1-2_stable-baselines3-ppo_91df110f/resolve/v1/model.zip

Evaluation

Action selection was stochastic under the published evaluation environment contract.

Start Episodes Successes Success rate Mean return
Level1-2 100 97 97.0% 3077.086

Environment and Policy Contract

Item Value
Environment supermariobrosnes-turbo:SuperMarioBros-Nes-v0
Environment hash sha256:3ed482a01669b9e31884d67157ab5724d1bd45f97a1db19ae2515331e4387f73
Preprocessing {"frame_skip":4,"frame_stack":4,"max_pool_frames":false,"obs_copy":"safe_view","obs_crop":[32,0,0,0],"obs_crop_fill":0,"obs_crop_mode":"remove","obs_grayscale":true,"obs_resize":[84,84],"obs_resize_algorithm":"area","pipeline":"supermariobrosnes_turbo_native_vec_env","policy_observation_layout":"channel_first","sticky_action_prob":0.0}
Action contract {"set":"simple"}

Provenance

Item Value
Source rlab
Run Level1-2_base_s1_20260704T072752Z
Recipe base
Seed 1
Source commit 40b8613743e41d36489701888cafcee24306604e
Evaluated artifact tsilva/SuperMarioBros-Nes-v0/Level1-2_base_s1_20260704T072752Z-checkpoint:step-7500000

Files

File Purpose
model.zip Stable-Baselines3 policy checkpoint
model.json Versioned checkpoint identity, policy type, provenance, and recipe binding
recipe.json Versioned execution and evaluation contract
release_manifest.json Release identity, evaluation evidence, and artifact hashes
replay.mp4 Browser-safe representative episode
LICENSE License for rlab-authored policy weights and publication material

Limitations

Evaluation establishes performance only for the published environment hash, start distribution, policy preprocessing, and action-selection protocol. It does not establish generalization to other levels, environments, ROM revisions, or contracts.

Licensing

The rlab-authored policy weights and publication material are licensed under the MIT License in LICENSE. Emulator/runtime software and game assets remain governed by their own licenses and terms. This repository does not redistribute a game ROM.

Policy Lineage

This is a legacy rlab policy trained by Stable-Baselines3 PPO. It is not a GradLab-trained checkpoint and no GradLab compatibility is claimed.

  • Trainer: Stable-Baselines3
  • Algorithm: PPO
  • Model class: stable_baselines3.ppo.ppo.PPO
  • Full lineage digest: 91df110f1c98e4200d38c20a204f67a8fbaf4e3b6ec1a0fc2fc382a65c7814ea
  • Immutable release: hf://tsilva/Level1-2_stable-baselines3-ppo_91df110f@v1
  • Exact checkpoint tag: checkpoint-7500000

The mutable main branch contains this corrected card; the original v1 release files, commit, and tag remain unchanged.

Downloads last month
37
Video Preview
loading

Collection including tsilva/Level1-2_stable-baselines3-ppo_91df110f