Reinforcement LearningUnifying n-step Action-Value Backups with Q(σ)

ARTICLE

Estimating Q-values with Off-policy n-step Q(σ)

Loading lesson…