ViZDoom single-player Deathmatch score attack β GradLab PPO @ 464M env steps
An immutable PPO research release for VizdoomDeathmatch-v1 / VizdoomDeathmatch-v1, backed by accepted exact-contract evaluation evidence and a separately labeled representative replay.
Identity
| Item | Value |
|---|---|
| Goal | VizdoomDeathmatch-v1 β ViZDoom single-player Deathmatch score attack |
| Environment | VizdoomDeathmatch-v1 (vizdoom-turbo:VizdoomDeathmatch-v1) |
| Trainer | GradLab |
| Algorithm | PPO |
| Model class | gradlab.ppo.GradLabPPO |
| Compatibility | SB3-compatible |
| Checkpoint | Exact step 463970304 |
| Immutable release | v3 and checkpoint-463970304 |
| Lineage | b033024760cb83f03f906eb9d49f83a3ba68f8638e2ecaa4879dd17615fc2093 |
| Environment index | GradLab β VizdoomDeathmatch-v1 |
| Featured | True |
Evaluation Status
Status: Accepted under a stochastic full evaluation of 100 episodes.
| Metric | Meaning | Observed | Requirement | Outcome |
|---|---|---|---|---|
eval/full/episode/return/shaped/mean |
Full-eval return mean | 100.99 | >= 10.0 |
Pass |
Ranking metrics:
| Metric | Meaning | Observed | Direction |
|---|---|---|---|
eval/full/episode/return/shaped/mean |
Full-eval return mean | 100.99 | max |
eval/full/episode/return/shaped/max |
Full-eval return max | 208 | max |
leader/checkpoint/step |
Leader checkpoint step | 463970304 | min |
The normalized per-episode results, aggregates, rules, contracts, outcomes, and authoritative hashes are in evaluation_evidence.json. Evidence SHA-256: 8a0ae072c5fdcba89a82179878df325a3b975853b426a790d35bf4734deebb7b.
Trainer, Algorithm, and Compatibility
- Trainer: GradLab
- Algorithm: PPO
- Model class:
gradlab.ppo.GradLabPPO - Compatibility: SB3-compatible
- Hugging Face library:
gradlab
Comparable Changes
Comparable: v3 republishes the same checkpoint and exact evaluation contract as v2. Releases are compared only when lineage, run, seed, and evaluation-contract hash all match.
Representative Replay
Watch on YouTube. The root replay.mp4 is one validated training episode. It is representative media and is not evaluation or promotion evidence.
| Item | Value |
|---|---|
| Outcome | terminated |
| Goal success | False |
| Seed | 123 |
| Start | default |
| Policy steps | 1331 |
| Episode return | 96.0 |
| Player runtime | {"asset":null,"contract_mode":"training","device_type":"mps","environment_hash":"sha256:5ce5419576605f9e2f1f532b147067952c8a5876a7bc08f4ef96ccf6eaec412a","execution_target":"local_player","overrides":[],"provider_id":"vizdoom-turbo","provider_version":"1.3.0.post23","qualified_environment_id":"vizdoom-turbo:VizdoomDeathmatch-v1","runtime_image_digest":"docker:ghcr.io/tsilva/gradlab/gradlab-train@sha256:5c8d1e2e09e9781942cc3038344b5ec4cc902ad9e3cf95f1219dd3951104743e","runtime_versions":{"breakout_turbo_env":"0.5.2","stable_baselines3":"2.8.0","stable_retro_turbo":"1.0.1.post37","supermariobrosnes_turbo":"0.6.2","vizdoom_turbo":"1.3.0.post23"},"seed":123,"source":{"distribution":"gradlab","git_commit":"1713894536640a6a27e30bddb787eb6b6a902187","kind":"checkout","source_tree_sha256":"533b692e2afe698221133f1f92cf1086df5723ffb2d0670cb059ffb1635081c5","version":"0.1.1","working_tree_dirty":true}} |
Quick Start
uvx --from gradlab gradlab play hf://tsilva/VizdoomDeathmatch-v1_gradlab-ppo_b0330247@v3
The immutable model reference is hf://tsilva/VizdoomDeathmatch-v1_gradlab-ppo_b0330247@v3.
Full Evaluation
Evaluation used stochastic action selection for all 100 declared episodes. See evaluation_evidence.json for every normalized episode and all materialized acceptance and ranking outcomes.
Contracts
| Item | Value |
|---|---|
| Environment hash | sha256:95a7666af8d6c9bf03df393875fed8bf90d81981a45a647c8ab68c3e0a9eaff8 |
| Preprocessing and model inputs | {"model_inputs":{"context":{"armor":{"encoding":{"clip":true,"high":1.0,"kind":"continuous","low":0.0,"offset":0.0,"scale":0.005},"signal":"armor","update":"transition"},"health":{"encoding":{"clip":true,"high":2.0,"kind":"continuous","low":0.0,"offset":0.0,"scale":0.01},"signal":"health","update":"transition"},"selected_weapon":{"encoding":{"kind":"categorical","values":[1,2,3,4,5,6]},"signal":"selected_weapon","update":"transition"},"selected_weapon_ammo":{"encoding":{"clip":true,"high":1.0,"kind":"continuous","low":0.0,"offset":0.0,"scale":0.0033333333333333335},"signal":"selected_weapon_ammo","update":"transition"},"weapon_ammo":{"encoding":{"clip":true,"high":1.0,"kind":"continuous","low":0.0,"offset":0.0,"scale":[1.0,0.005,0.02,0.005,0.02,0.0033333333333333335]},"signal":"weapon_ammo","update":"transition"},"weapons_owned":{"encoding":{"clip":true,"high":1.0,"kind":"continuous","low":0.0,"offset":0.0,"scale":1.0},"signal":"weapons_owned","update":"transition"}},"schema_version":1},"preprocessing":{"frame_skip":2,"frame_stack":4,"max_pool_frames":false,"obs_copy":"safe_view","obs_crop":[0,32,0,0],"obs_crop_fill":0,"obs_crop_mode":"mask","obs_grayscale":true,"obs_resize":[84,84],"obs_resize_algorithm":"area","pipeline":"vizdoom_turbo_native_vec_env","policy_observation_layout":"dict_observation_context_v1","sticky_action_prob":0.0}} |
| Action semantics | {"contract_hash":"f1aa454eeec78c4a1f5ad4f056d1ceec7881d397649cfa0cb505b3b7cd790634","execution_hash":"abfee888d741c472cbab5c9925937b12c34f99f381e243c5264bcb00c689b437","policy":{"codec":{"type":"identity"},"semantics":{"encoding":"explicit","entries":[{"controls":[{"atoms":[],"inputs":[],"player":1}],"label":"noop","semantic_id":"noop","value":0},{"controls":[{"atoms":["attack"],"inputs":["a"],"player":1}],"label":"attack","semantic_id":"attack","value":1},{"controls":[{"atoms":["move_forward"],"inputs":["up"],"player":1}],"label":"move forward","semantic_id":"move_forward","value":2},{"controls":[{"atoms":["move_backward"],"inputs":["down"],"player":1}],"label":"move backward","semantic_id":"move_backward","value":3},{"controls":[{"atoms":["move_left"],"inputs":["left"],"player":1}],"label":"move left","semantic_id":"move_left","value":4},{"controls":[{"atoms":["move_right"],"inputs":["right"],"player":1}],"label":"move right","semantic_id":"move_right","value":5},{"controls":[{"atoms":["turn_left"],"inputs":["turn_left"],"player":1}],"label":"turn left","semantic_id":"turn_left","value":6},{"controls":[{"atoms":["turn_right"],"inputs":["turn_right"],"player":1}],"label":"turn right","semantic_id":"turn_right","value":7},{"controls":[{"atoms":["speed","move_forward"],"inputs":["x","up"],"player":1}],"label":"speed move forward","semantic_id":"speed_move_forward","value":8},{"controls":[{"atoms":["attack","move_forward"],"inputs":["a","up"],"player":1}],"label":"attack move forward","semantic_id":"attack_move_forward","value":9},{"controls":[{"atoms":["attack","move_backward"],"inputs":["a","down"],"player":1}],"label":"attack move backward","semantic_id":"attack_move_backward","value":10},{"controls":[{"atoms":["attack","move_left"],"inputs":["a","left"],"player":1}],"label":"attack move left","semantic_id":"attack_move_left","value":11},{"controls":[{"atoms":["attack","move_right"],"inputs":["a","right"],"player":1}],"label":"attack move right","semantic_id":"attack_move_right","value":12},{"controls":[{"atoms":["attack","turn_left"],"inputs":["a","turn_left"],"player":1}],"label":"attack turn left","semantic_id":"attack_turn_left","value":13},{"controls":[{"atoms":["attack","turn_right"],"inputs":["a","turn_right"],"player":1}],"label":"attack turn right","semantic_id":"attack_turn_right","value":14},{"controls":[{"atoms":["select_next_weapon"],"inputs":["select_next_weapon"],"player":1}],"label":"select next weapon","semantic_id":"select_next_weapon","value":15},{"controls":[{"atoms":["select_prev_weapon"],"inputs":["select_prev_weapon"],"player":1}],"label":"select prev weapon","semantic_id":"select_prev_weapon","value":16}],"status":"available"},"space":{"dtype":"int64","n":17,"start":0,"type":"discrete"}},"provider":{"mode":"custom_discrete","preset":null,"provider_id":"vizdoom-turbo","semantics":{"encoding":"explicit","entries":[{"controls":[{"atoms":[],"inputs":[],"player":1}],"label":"noop","semantic_id":"noop","value":0},{"controls":[{"atoms":["attack"],"inputs":["a"],"player":1}],"label":"attack","semantic_id":"attack","value":1},{"controls":[{"atoms":["move_forward"],"inputs":["up"],"player":1}],"label":"move forward","semantic_id":"move_forward","value":2},{"controls":[{"atoms":["move_backward"],"inputs":["down"],"player":1}],"label":"move backward","semantic_id":"move_backward","value":3},{"controls":[{"atoms":["move_left"],"inputs":["left"],"player":1}],"label":"move left","semantic_id":"move_left","value":4},{"controls":[{"atoms":["move_right"],"inputs":["right"],"player":1}],"label":"move right","semantic_id":"move_right","value":5},{"controls":[{"atoms":["turn_left"],"inputs":["turn_left"],"player":1}],"label":"turn left","semantic_id":"turn_left","value":6},{"controls":[{"atoms":["turn_right"],"inputs":["turn_right"],"player":1}],"label":"turn right","semantic_id":"turn_right","value":7},{"controls":[{"atoms":["speed","move_forward"],"inputs":["x","up"],"player":1}],"label":"speed move forward","semantic_id":"speed_move_forward","value":8},{"controls":[{"atoms":["attack","move_forward"],"inputs":["a","up"],"player":1}],"label":"attack move forward","semantic_id":"attack_move_forward","value":9},{"controls":[{"atoms":["attack","move_backward"],"inputs":["a","down"],"player":1}],"label":"attack move backward","semantic_id":"attack_move_backward","value":10},{"controls":[{"atoms":["attack","move_left"],"inputs":["a","left"],"player":1}],"label":"attack move left","semantic_id":"attack_move_left","value":11},{"controls":[{"atoms":["attack","move_right"],"inputs":["a","right"],"player":1}],"label":"attack move right","semantic_id":"attack_move_right","value":12},{"controls":[{"atoms":["attack","turn_left"],"inputs":["a","turn_left"],"player":1}],"label":"attack turn left","semantic_id":"attack_turn_left","value":13},{"controls":[{"atoms":["attack","turn_right"],"inputs":["a","turn_right"],"player":1}],"label":"attack turn right","semantic_id":"attack_turn_right","value":14},{"controls":[{"atoms":["select_next_weapon"],"inputs":["select_next_weapon"],"player":1}],"label":"select next weapon","semantic_id":"select_next_weapon","value":15},{"controls":[{"atoms":["select_prev_weapon"],"inputs":["select_prev_weapon"],"player":1}],"label":"select prev weapon","semantic_id":"select_prev_weapon","value":16}],"status":"available"},"space":{"dtype":"int64","n":17,"start":0,"type":"discrete"},"table_hash":"0bd9dd28d67a88ef6bc54734f53d55bc4af597e672665a7f20d4b204098036af"},"requested":{"meanings":["noop","attack","move_forward","move_backward","move_left","move_right","turn_left","turn_right","speed_move_forward","attack_move_forward","attack_move_backward","attack_move_left","attack_move_right","attack_turn_left","attack_turn_right","select_next_weapon","select_prev_weapon"],"mode":"custom_discrete","preset":null,"table":[[],["ATTACK"],["MOVE_FORWARD"],["MOVE_BACKWARD"],["MOVE_LEFT"],["MOVE_RIGHT"],["TURN_LEFT"],["TURN_RIGHT"],["SPEED","MOVE_FORWARD"],["ATTACK","MOVE_FORWARD"],["ATTACK","MOVE_BACKWARD"],["ATTACK","MOVE_LEFT"],["ATTACK","MOVE_RIGHT"],["ATTACK","TURN_LEFT"],["ATTACK","TURN_RIGHT"],["SELECT_NEXT_WEAPON"],["SELECT_PREV_WEAPON"]],"table_hash":"0bd9dd28d67a88ef6bc54734f53d55bc4af597e672665a7f20d4b204098036af"},"schema_version":1,"semantic_hash":"7a5a73ecc7f770d673626e5aaec26aa1ad3c0db7c87a1b715c79019b066009d2"} |
Provenance
| Item | Value |
|---|---|
| Source | gradlab |
| Run | gradlab-294ea8ebb1c7c049caedfb29221eca2f |
| Recipe | ppo |
| Seed | 123 |
| Source commit | bfad768f66d68d3ddc01f69691ef1e5dbab9a931 |
| Evaluated artifact | https://pub-fc35c0b186ce4aad8eea5e93d38c99db.r2.dev/runs/gradlab-294ea8ebb1c7c049caedfb29221eca2f/checkpoints/463970304-e5596d939cef0b6d34e75874b4d917660e7b6efab672ad8d7bea85445a7bb100/model.zip |
Limitations
Evaluation establishes performance only for the published environment, policy, action-selection, start, boundary, and evaluation contracts. It does not establish generalization to other tasks or contracts. The representative replay cannot establish or change acceptance, ranking, promotion, or featured status.
Files
| File | Purpose |
|---|---|
model.zip |
Immutable policy checkpoint |
model.json |
Trainer, algorithm, model class, checkpoint identity, and provenance |
recipe.json |
Materialized training, playback, and evaluation contracts |
evaluation_evidence.json |
Normalized authoritative evaluation evidence and rule outcomes |
release_manifest.json |
Release identity, evaluation evidence, and artifact hashes |
replay.mp4 |
Browser-safe representative completed episode |
LICENSE |
License for GradLab-authored weights and publication material |
Licensing
GradLab-authored policy weights and publication material are licensed under the MIT License in LICENSE. Emulator/runtime software and game assets remain governed by their own licenses and terms. This repository does not redistribute a game ROM.