Q-Learning Agent playing Taxi-v3
This is a trained model of a Q-Learning agent playing Taxi-v3.
Evaluation
- Environment: Taxi-v3
- Mean reward: 7.56
- Standard deviation: 2.71
- Leaderboard score (mean_reward - std_reward): 4.8533
- Required course score: 4.5
- Evaluation episodes: 100
Training configuration
- Training episodes: 25000
- Learning rate: 0.7
- Gamma: 0.95
- Maximum steps: 99
- Maximum epsilon: 1.0
- Minimum epsilon: 0.05
- Epsilon decay rate: 0.005
- Q-table shape: (500, 6)
Model files
- q-learning.pkl — trained Q-table and configuration
- results.json — evaluation results
- taxi-replay.mp4 — recorded agent replay
This model was created as part of the Hugging Face Deep Reinforcement Learning Course Unit 2 hands-on.
Evaluation results
- mean_reward on Taxi-v3self-reported7.56 +/- 2.71