Concepts / Off-policy TD control

Off-policy TD control

Q-learning is an off-policy TD control algorithm in reinforcement learning.

  • Programming

Finding Q-learning’s Place

Q-learning is the name of a particular algorithm. The supplied definition places it within reinforcement learning and identifies it more precisely as an off-policy TD control algorithm. Reading that description accurately is the first skill: identify the broad field, the learning approach, the control role, and the policy relationship without adding mechanics that have not yet been introduced.

includesincludesincludesReinforcementlearningTD controlOff-policy controlQ-learning
Where does Q-learning fit in the hierarchy from reinforcement learning to off-policy control?

Q-learning is an off-policy TD control algorithm in reinforcement learning.

Reading the Definition Precisely

Classifying a named algorithm

Use the supplied definition to classify Q-learning without adding an unsupported defining expression or implementation detail.

Identify the field: The phrase reinforcement learning identifies the broad area in which Q-learning belongs.

Identify the learning approach: The abbreviation TD identifies temporal-difference learning in the supplied description.

Identify the role: The word control indicates that the method belongs to TD control rather than being described only as value prediction.

Identify the policy category: The word off-policy identifies the relationship between the policy being improved or evaluated and the actions used to produce experience.

Q-learning is a reinforcement learning algorithm in the off-policy TD control category.

Two Jobs in TD Control

TD control combines two activities. First, value prediction makes the value function better at predicting returns for the current policy. Second, local policy improvement uses that value information to improve the policy. These activities work together: experience supports value prediction, and the resulting value information supports a local improvement in action preference.

helps select actionssupportsinformsCurrent policybasis for action choiceExperienceobserved interactionValue predictionpredict returnsLocal policyimprovementuse value information
How does a TD control process use experience to estimate returns and improve a policy locally?

The word control matters because the process is not limited to predicting the returns of a fixed policy. The policy is also improved locally using the value function. Therefore, the on-policy or off-policy label does not say whether prediction or improvement occurs; TD control includes both. The label identifies the relationship between the policy being improved or evaluated and the actions used to produce experience.

Policy Relationships

The key diagnostic question is: does the current policy select the actions used to produce experience? In an on-policy method, the current policy is used to select actions. In an off-policy method, the current policy is not used to select those actions. The policy being improved need not be the policy that selected the observed actions.

producesseparated fromOn-policyCurrent policyselects actionsExperiencefrom current policyObserved actionsnot selected by currentpolicyCurrent policyimproved or evaluatedOff-policy
What is the difference between the policy generating the experience and the policy being improved or evaluated?
MethodPolicy categorySource classification
SarsaOn-policyThe current policy selects actions
Q-learningOff-policyThe current policy does not select the actions used to produce experience
Expected SarsaOff-policyThe current policy does not select the actions used to produce experience

Exploration and Learning

An agent must choose actions while it is still learning. The resulting experience helps the value function estimate returns, and that value information can then improve the policy locally. The central design issue is maintaining sufficient exploration while learning from experience. Action choices matter twice: they determine what experience is collected, and they are connected to the policy being evaluated or improved.

may includesupportsinformssupportsAction choiceduring learningExplorationsufficient experienceExperienceinformative observationsValue estimatespredict returnsPolicy improvementlocal change
How does choosing exploratory actions affect the experience collected, value estimates, and policy improvement?

Classification Practice

A policy-category check

A learner is given three names from the material: Sarsa, Q-learning, and Expected Sarsa. Classify each as on-policy or off-policy.

Classify Sarsa: The source identifies Sarsa as an on-policy TD control method because the current policy is used to select actions.

Classify Q-learning: The source identifies Q-learning as an off-policy TD control algorithm.

Classify Expected Sarsa: The source also classifies Expected Sarsa as off-policy.

Check the diagnostic question: The distinction concerns the relationship between the current policy and the actions used to produce experience, not whether TD control includes prediction or improvement.

Sarsa is on-policy. Q-learning and Expected Sarsa are off-policy.

EASY

Explain in one or two sentences why the statement Q-learning is a reinforcement learning algorithm is less precise than the statement Q-learning is an off-policy TD control algorithm in reinforcement learning.

Hints
  • Name the broad field first.
  • Then identify the learning approach, control role, and policy category.
  • Treating Q-learning as the name of reinforcement learning as a whole.

    The source identifies Q-learning as a particular algorithm within reinforcement learning.

    Fix: Describe reinforcement learning as the broad field and Q-learning as an algorithm within it.

  • Assuming that off-policy means TD control does not perform value prediction.

    TD control combines value prediction for the current policy with local policy improvement.

    Fix: Use off-policy to describe the policy relationship, while retaining both prediction and improvement as parts of TD control.

  • Inferring a complete algorithmic expression from the classification alone.

    The supplied material does not provide that defining expression.

    Fix: Keep the conclusion at the supported level: Q-learning belongs to the off-policy TD control category.

  • Classifying a method without asking which policy selects actions.

    That relationship is the key diagnostic question in the supplied explanation.

    Fix: Ask whether the current policy is used to select the actions that produce experience.

Key Takeaways

  1. Q-learning is an off-policy TD control algorithm in reinforcement learning.
  2. TD control combines value prediction for the current policy with local policy improvement.
  3. On-policy methods use the current policy to select actions; Sarsa is the supplied example.
  4. Off-policy methods do not use the current policy to select actions; Q-learning and Expected Sarsa are the supplied examples.
  5. Exploration is central because the agent must collect informative experience while its value function and policy are being improved.

Key Takeaways

  • Q-learning is a particular reinforcement learning algorithm, not the name of reinforcement learning as a whole.
  • Its precise supplied classification is off-policy TD control.
  • TD control joins value prediction with local policy improvement.
  • The on-policy or off-policy distinction concerns the relationship between the current policy and the actions producing experience.
  • Sarsa is on-policy, while Q-learning and Expected Sarsa are off-policy in the supplied material.