Q-learning Taxi-v3

Trained from scratch for Kay Zheng's Hugging Face Deep RL coursework with AI coding assistance. 50,000 training episodes, NumPy seed 42, Gymnasium 0.29.1. Independent evaluation: 1,000 episodes, reset seeds 100000–100999. Mean reward: 7.965; population standard deviation: 2.565107210235081. Mean minus standard deviation: 5.399893.

Reproduce

Install gymnasium==0.29.1 and numpy, then run python train_taxi.py. qtable.npy is the trained policy; choose qtable[state].argmax(). evaluation.json includes every evaluation episode reward.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results