›Reinforcement Learning›Balancing Exploration and Exploitation with Upper-Confidence-Bound Action Selection