q-Taxi-v3
This model was trained as part of the Hugging Face Deep Reinforcement Learning Course.
Model Description
- Environment:
Taxi-v3 - Library:
q-learning - Algorithm:
q-learning - Mean Reward:
8.00 +/- 0.50
Evaluation Results
The model was evaluated on Taxi-v3 and achieved an average reward of 8.00 with a standard deviation of 0.50.
Evaluation results
- mean_reward on Taxi-v3self-reported8.00 +/- 0.50