Instructions to use hwihwalab/pusher-v5-ppo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- stable-baselines3
How to use hwihwalab/pusher-v5-ppo with stable-baselines3:
from huggingface_sb3 import load_from_hub checkpoint = load_from_hub( repo_id="hwihwalab/pusher-v5-ppo", filename="{MODEL FILENAME}.zip", ) - Notebooks
- Google Colab
- Kaggle
- π¦Ύ Pusher-v5 PPO // AI Hub & Live Control Cockpit
- π Model Specifications & Benchmark Performance
- ποΈ System Architecture
- πΉοΈ Interactive Cockpit Features
- π Quickstart & Usage
- π Quick Python Evaluation Snippet
- β¨οΈ Keyboard Shortcuts Reference
- π‘οΈ AI Governance & Documentation Architecture
- π Repository Contents
- π Open Source Hubs & Project Links
- π License
- π Model Specifications & Benchmark Performance
π¦Ύ Pusher-v5 PPO // AI Hub & Live Control Cockpit
MuJoCo 7-DOF Robotic Continuous Control Telemetry & PPO Deep Reinforcement Learning Platform
π English Documentation | π°π· νκ΅μ΄ λ§€λ΄μΌ
This repository contains an advanced continuous deep reinforcement learning system (PPO) and a real-time engineering telemetry cockpit for 7-DOF robotic arm manipulation in Gymnasium MuJoCo Pusher-v5.
π Model Specifications & Benchmark Performance
| Parameter | Specification |
|---|---|
| Environment | Gymnasium MuJoCo Pusher-v5 (7-DOF Robotic Arm) |
| Observation Space | 23-dimensional continuous vector (Joints, Velocities, Tip 3D, Object 3D, Goal 3D) |
| Action Space | 7-dimensional continuous motor torques (Box[-2.0, 2.0], float32) |
| Algorithm | Proximal Policy Optimization (PPO) with MlpPolicy |
| Deep Learning Framework | Stable-Baselines3 / PyTorch backend |
| Observation Normalization | Raw MuJoCo coordinates & velocities |
| Baseline Return (Step 0) | -57.51 pts (Random exploration, arm-to-object dist ~0.215m) |
| Converged Return (Step 300k+) | -32.42 Β± 4.30 pts *(Peak: -26.15 pts)* |
| Arm-to-Object Proximity | 0.028 m (Precise contact & cylinder grasp alignment) |
| Goal Proximity Accuracy | 0.054 m (Target zone reached & pushed) |
ποΈ System Architecture
flowchart TD
subgraph Web_Cockpit ["1-Screen Zero-Scroll Robotics Telemetry Cockpit"]
W1["HTML5 / CSS3 / Vanilla JS Client"] <-->|"WebSocket /ws/simulation @ 30 FPS"| S1["FastAPI High-Performance Engine"]
S1 -->|"Base64 JPEG Physics Stream"| W1
S1 -->|"7-DOF Bipolar Torques (-2 to +2 Nm)"| W1
S1 -->|"3D Vector Coordinates (Tip, Obj, Goal)"| W1
W1 -->|"Control Commands (Start, Pause, Step, Reset, Policy)"| S1
end
subgraph Analytics_Deck ["4-Tab Analytics & Replay Deck"]
T1["Tab 1: Live Telemetry Dynamics (Raw & 20-Ep Moving Average)"]
T2["Tab 2: Milestone Replay Deck (16:9 Widescreen Video Gallery)"]
T3["Tab 3: Live PPO Logs (Algorithmic Console Stream)"]
T4["Tab 4: Environment & Reward Math Specifications"]
end
subgraph Deep_RL_Pipeline ["Stable-Baselines3 PPO Training Loop"]
TR1["train.py / Background Thread"] --> TR2["MuJoCo Pusher-v5 Physics"]
TR2 --> TR3["VisualProgressCallback"]
TR3 --> TR4["Step 0 to 300k MP4 & GIF Videos"]
TR3 --> TR5["Training Plots & Metrics JSON"]
TR4 & TR5 --> TR6["Single-Click ZIP Archive: ppo_pusher_bundle.zip"]
end
πΉοΈ Interactive Cockpit Features
- High-Fidelity 30 FPS Physics Stream:
- Ultra low-latency canvas streaming via WebSocket.
- 7-DOF Action Space Motor Torque Bipolar Gauge (
[-2.0, +2.0] Nm) with positive (Cyan) and negative (Rose) deflection. - 3D Cartesian coordinates tracker for Fingertip, Object, and Goal in real meters.
- Deep RL Training Budget Presets:
500 Ep (50k Steps β’ ~12s) - Quick Test2,000 Ep (200k Steps β’ ~45s) - Basic Pushing5,000 Ep (500k Steps β’ ~1.8m) β Recommended Mature10,000 Ep (1M Steps β’ ~3.5m) - High-Precision
- Widescreen Checkpoint Replay Gallery:
- Side-by-side comparative video cards displaying the robotic arm's learning trajectory from random exploration (Step 0) to mature convergence (Step 30.7k).
- Instant 1-click export for MP4 videos and animated GIFs.
π Quickstart & Usage
1. Installation
git clone https://github.com/Hwihwa-Lab/pusher-v5-ppo.git
cd pusher-v5-ppo
pip install -r requirements.txt
2. Launch Local Web Control Cockpit
python app.py
Open your browser at http://localhost:8000.
3. One-Click Deploy to Hugging Face
python deploy_to_hf.py
4. Standalone CLI Training & Evaluation
# Train PPO agent
python train.py --timesteps 300000 --eval_freq 30000
# Evaluate trained model
python evaluate.py --model_path ./results/ppo_pusher.zip --episodes 5
π Quick Python Evaluation Snippet
You can load and evaluate this pre-trained agent in 5 lines of Python using Stable-Baselines3:
import gymnasium as gym
from stable_baselines3 import PPO
# 1. Initialize Pusher-v5 environment & load model
env = gym.make("Pusher-v5", render_mode="human")
model = PPO.load("results/ppo_pusher.zip")
# 2. Run deterministic pushing evaluation
obs, _ = env.reset()
done = False
while not done:
action, _ = model.predict(obs, deterministic=True)
obs, reward, terminated, truncated, _ = env.step(action)
done = terminated or truncated
env.close()
β¨οΈ Keyboard Shortcuts Reference
| Key | Action | Description |
|---|---|---|
Space |
Start / Pause | Toggle 30 FPS MuJoCo physical simulation stream |
R |
Reset Environment | Reset robotic arm, cylinder object, and target goal to new random positions |
S |
Step Once | Advance physics engine forward by 1 discrete timestep (0.05s) |
H |
Toggle HUD | Show or hide on-canvas telemetry data overlay |
π‘οΈ AI Governance & Documentation Architecture
This repository is governed by rigorous engineering protocols to ensure simulation fidelity and prevent vibe-coding drift (hosted on GitHub):
.cursorrules: AI Vibe-Coding Defense Master ProtocolDOCS_AI_CODING_PROTOCOL.md: Coding Standards & Master Documentation MapDOCS_SYSTEM_ARCHITECTURE.md: Full-Stack System & WebSocket Architecture SpecDOCS_DATA_SCHEMA.md: WebSocket Telemetry Protocol & REST Data SchemaDOCS_MODEL_EVALUATION_AND_HF_DEPLOY.md: Benchmark Evaluation & Hugging Face Hub Pipeline
π Repository Contents
README.md: English Model Card and benchmark performance guide.README_KR.md: Full Korean comprehensive manual (νκ΅μ΄ λ§€λ΄μΌ).app.py: FastAPI high-performance backend & 30 FPS WebSocket simulation server.train.py: Stable-Baselines3 PPO 7-DOF training engine withVisualProgressCallback.evaluate.py: Standalone 5-episode deterministic policy evaluator and video recorder.visualizer.py: Standalone Matplotlib visualizer and benchmark plotter.web/: 1-Screen zero-scroll telemetry cockpit frontend (app.js,index.html,style.css).results/ppo_pusher.zip: Pre-trained PPO neural network weights (300,000 steps, -32.4 pts).ppo_pusher_bundle.zip: Complete production archive with weights, 12 checkpoint videos, and plots.deploy_to_hf.py: One-click automated Hugging Face Model Hub deployer.requirements.txt&packages.txt: Python and system dependency manifests.
π Open Source Hubs & Project Links
- π GitHub Repository: https://github.com/Hwihwa-Lab/pusher-v5-ppo
- π€ Hugging Face Model Hub: https://huggingface.co/hwihwalab/pusher-v5-ppo
π License
This project is licensed under the MIT License - see the LICENSE file for details.
Trained and deployed with Pusher AI Hub by hwihwalab.
- Downloads last month
- -
Evaluation results
- Mean Evaluation Reward (5-Ep Average) on Gymnasium MuJoCo Pusher-v5self-reported-32.420