Q-Learning Agent playing Taxi-v3
This is a trained model of a Tabular Q-Learning agent playing Taxi-v3.
The agent was trained and evaluated directly using the Bellman optimality update equation.
Evaluation Results
- Mean Reward:
7.69 +/- 2.87(Passing threshold: $\ge 4.0$) - Number of Evaluation Episodes: 100
- Model weights: Saved in
q_table.pkl
Trained and submitted by Subhash3008.
Evaluation results
- mean_reward on Taxi-v3self-reported7.69 +/- 2.87