Reinforcement LearningOptimistic Initial Values for Exploration

ARTICLE

How Optimistic Initial Values Change Bandit Performance

Loading lesson…