Reinforcement LearningOff-policy Prediction with Importance Sampling

ARTICLE

Why Off-policy Estimates Can Have Infinite Variance

Loading lesson…