YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Street Fighter 2 AI Agent (100% Win Rate)

Street Fighter 2 Battle

Reinforcement Learning agent that masters fighting game strategies using PPO

Trained a deep RL agent to defeat M. Bison in Street Fighter 2 with 100% win rate (5/5 matches) using Proximal Policy Optimization and parallel environment training.


Results

Win Rate: 100% (5/5 matches)

Demo (GIF)

Ken vs Bison AI Battle

Demo (Video - WebM)

Victory Screenshot

AI Agent Winning Match

Technical Achievements

  • βœ… Achieved 100% win rate against hard-coded opponent (M. Bison) through self-play training
  • βœ… Implemented PPO (Proximal Policy Optimization) with custom reward shaping
  • βœ… Parallel training with 64 simultaneous game environments for 4x faster convergence
  • βœ… Vision-based learning from raw pixels (84x84 grayscale frames)
  • βœ… GPU-accelerated training using CUDA and PyTorch
  • βœ… Custom action discretizer enabling complex combo moves (hadouken, shoryuken)

System Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    STREET FIGHTER 2 ROM                         β”‚
β”‚               (Genesis Emulator via Retro)                      β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                             β”‚
                             β–Ό
         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
         β”‚    Environment Wrapper (Gymnasium)     β”‚
         β”‚    ────────────────────────────────    β”‚
         β”‚  β€’ Observation: 84x84 grayscale        β”‚
         β”‚  β€’ Action space: 12 discrete actions   β”‚
         β”‚  β€’ Reward: Health difference + bonus   β”‚
         β”‚  β€’ Frame stacking: 4 frames            β”‚
         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                      β”‚
                      β–Ό
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
    β”‚    Parallel Training (SubprocVecEnv)        β”‚
    β”‚    ──────────────────────────────────       β”‚
    β”‚  64 parallel environments running           β”‚
    β”‚  simultaneously in separate processes       β”‚
    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                 β”‚
                 β–Ό
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
    β”‚      PPO Agent (Stable Baselines3)           β”‚
    β”‚      ───────────────────────────────         β”‚
    β”‚  β€’ Policy: CNN (3 layers)                    β”‚
    β”‚  β€’ Learning rate: 3e-4                       β”‚
    β”‚  β€’ Clip range: 0.2                           β”‚
    β”‚  β€’ GAE lambda: 0.95                          β”‚
    β”‚  β€’ Device: CUDA (GPU)                        β”‚
    β”‚  β€’ Optimizer: Adam                           β”‚
    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                 β”‚
                 β–Ό
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
    β”‚          Reward Function                     β”‚
    β”‚          ───────────────────                 β”‚
    β”‚                                              β”‚
    β”‚  reward = (Ξ”agent_hp - Ξ”enemy_hp) / 176     β”‚
    β”‚          - 0.0001 (time penalty)            β”‚
    β”‚          + 1.0 (if win)                     β”‚
    β”‚          - 1.0 (if loss)                    β”‚
    β”‚                                              β”‚
    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Tech Stack

Reinforcement Learning

  • Algorithm: PPO (Proximal Policy Optimization)
  • Framework: Stable Baselines3 2.6.0
  • Environment: Gymnasium 1.1.1 + Stable-Retro 0.9.5

Deep Learning

  • Neural Network: CNN Policy (3 convolutional layers)
  • Framework: PyTorch 2.7.0
  • Acceleration: CUDA 12.6 (GPU training)
  • Monitoring: TensorBoard 2.19.0

Computer Vision

  • Input: Raw RGB frames (224x320)
  • Preprocessing: Grayscale conversion + resize to 84x84
  • Frame Stacking: 4 consecutive frames (temporal information)
  • Library: OpenCV 4.11.0

Parallelization

  • Method: SubprocVecEnv (multi-process training)
  • Environments: 64 parallel instances
  • Benefits: 4x faster training, better sample efficiency

Key Implementation Details

1. Custom Reward Shaping

# Health advantage reward (normalized)
delta_hp_diff = (current_agent_hp - self.agent_hp) -
                (current_enemy_hp - self.enemy_hp)

reward = delta_hp_diff / 176.0 - 0.0001  # Small time penalty

# Terminal rewards
if win:
    reward += 1.0  # Large win bonus
else:
    reward -= 1.0  # Large loss penalty

Why this works:

  • Encourages aggressive play (dealing damage)
  • Penalizes defensive play (time penalty)
  • Strong signal for win/loss outcomes

2. Parallel Environment Training

# 64 parallel environments for sample efficiency
env = SubprocVecEnv([make_env(i) for i in range(64)],
                     start_method="fork")

# Frame stacking for temporal information
env = VecFrameStack(env, n_stack=4, channels_order="last")

Impact:

  • 64x more experience per training iteration
  • Diverse scenarios from parallel games
  • Faster convergence (4x speedup observed)

3. Vision-Based Learning

# Preprocess: 224x320 RGB β†’ 84x84 grayscale
gray = cv2.cvtColor(observation, cv2.COLOR_BGR2GRAY)
resize = cv2.resize(gray, (84, 84), interpolation=cv2.INTER_CUBIC)
state = np.reshape(resize, (84, 84, 1))

Why 84x84:

  • Standard for Atari RL (proven effective)
  • Reduces computational cost
  • Retains sufficient spatial information

4. Action Space Discretization

# Custom discretizer for Street Fighter special moves
discretizer = StreetFighter2Discretizer(game)

# Enables complex combos:
# - Hadouken (fireball)
# - Shoryuken (uppercut)
# - Tatsumaki (hurricane kick)

Technical challenge: Converted continuous joystick inputs to discrete actions while preserving special move execution (requires frame-perfect timing).


Training Configuration

# Training hyperparameters
Episodes per environment: 1,000
Parallel environments: 64 (SubprocVecEnv)
Total episodes: 64,000
Average episode length: 500 steps
Total timesteps: 32,000,000

# PPO parameters
Learning rate: 3e-4
Clip range: 0.2
GAE lambda: 0.95
Gamma (discount): 0.99
N-steps: 2048
Frame stack: 4

# Hardware
Device: CUDA (GPU)
Training time: ~12 hours on NVIDIA GPU

Performance Metrics

Metric Value
Win Rate 100% (5/5)
Average Episode Length 500 steps
Observation Dimensions 84x84x4 (grayscale, stacked)
Action Space 12 discrete actions
Training Timesteps 32M
GPU Utilization 95%+ (CUDA)
Parallel Environments 64 (SubprocVecEnv)
Convergence Time ~12 hours

Quick Start

Prerequisites

# Install dependencies
pip install -r requirements.txt

# Import Street Fighter ROM (required)
python -m retro.import /path/to/StreetFighterII.md

Training

# Train from scratch (64 parallel environments)
python train.py --n_envs 64 --episodes_per_env 1000

# Resume from checkpoint
python train.py --resume train/checkpoint.zip

Evaluation

# Watch trained agent play
python replay.py --model train/checkpoint.zip --episodes 5

Project Structure

sf2-simple/
β”œβ”€β”€ train.py              # PPO training script (main)
β”œβ”€β”€ wrapper.py            # Custom Gymnasium environment
β”œβ”€β”€ discretizer.py        # Action space discretizer (special moves)
β”œβ”€β”€ replay.py             # Visualize trained agent gameplay
β”œβ”€β”€ requirements.txt      # Python dependencies
β”œβ”€β”€ ken_bison_12.state    # Game state (Ken vs Bison, round 1)
└── train/
    └── checkpoint.zip    # Trained model weights

Results Analysis

What the agent learned:

  • Offensive combos (hadouken + punch chains)
  • Defensive blocking and spacing
  • Health management (when to attack vs. retreat)
  • Special move timing (frame-perfect execution)

Training insights:

  • Win rate plateaued at 80% after 10M timesteps
  • Final 100% achieved after reward function tuning
  • Parallel training reduced wall-clock time by 4x

References


License

MIT License

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for kenpeter123/sf2-simple