Q-learning: Off-policy TD Control
The cliff-walking task separates the shortest optimal route from the safer route under exploration.
The Dangerous Shortcut
Q-learning becomes easier to understand when the shortest route is not the safest route. The cliff-walking task creates exactly this situation: an agent must travel from a start state to a goal state, while a dangerous region lies beside one possible route. The task exposes a difference between two questions: should learned values describe the policy the agent is actually following, or should they describe the best policy the agent could follow?
A Walk Through the Grid
Imagine a gridworld with a start state and a goal state. At every step, the agent can move up, down, right, or left. Most transitions provide a reward of −1. A cliff occupies one region of the grid. If the agent enters the cliff, it receives a much worse reward of −100 and is immediately returned to the start state.
Comparing Two Complete Journeys
Compare an edge route that is shortest under perfect action choices with an upper route that stays farther from the cliff.
Choose the edge route: The agent travels along the edge toward the goal. This is the optimal route when the optimal choices are followed, but it places the agent next to the cliff.
Explore near the cliff: If action selection explores a random action near the cliff, the agent can enter the cliff, receive −100, and return to the start.
Choose the upper route: The agent takes a longer path through the upper part of the grid. This route is less exposed to the cliff, so exploratory choices do not create the same cliff risk along the route.
Compare the objectives: The edge route is shortest when the optimal actions are followed. The upper route is safer for a policy that still sometimes explores.
The task separates shortest-route optimality from safety during exploratory behavior.
Sarsa Follows Its Own Behavior
Sarsa is on-policy. Its learned values correspond to the policy the agent is actually following. In the cliff-walking task, that policy includes ε-greedy action selection, so the agent sometimes chooses a random action instead of its currently preferred action. Sarsa therefore accounts for the danger created by exploration near the cliff.
Because the edge route places the agent beside the cliff, an exploratory action there can produce the −100 penalty. Sarsa treats that possibility as part of the route's value. As a result, it prefers the longer path through the upper part of the grid. The route is not shorter, but it is safer for the exploratory policy the agent is actually using.
Sarsa does not merely ask which action would be best in an ideal, non-exploring run. It learns values for the policy that includes the agent's exploratory choices.
Q-learning Chooses an Ideal Target
Q-learning is off-policy. It learns the values of the optimal policy rather than the values of the exploratory policy currently being used. In the cliff-walking task, this means Q-learning can prefer the edge route because that route reaches the goal most effectively when the optimal choices are followed.
The behavior policy and the learned policy are therefore different ideas. The behavior policy determines what the agent actually does while it is exploring. The policy whose values Q-learning learns is the optimal policy. Q-learning can learn that the edge route is optimal even while an ε-greedy behavior policy occasionally takes an action that sends the agent into the cliff.
How ε-greedy Exploration Works
The example uses ε-greedy action selection with ε = 0.1. Under this rule, the greedy action is selected with probability 1 − ε, while a random action is selected with probability ε. Thus, the agent normally chooses its currently preferred action but still sometimes explores another action.
What do you think happens?
Suppose Q-learning has learned that the edge route is optimal, but ε-greedy exploration is still active. What can happen during an actual run?
Reveal answer
Answer: The agent can still choose a random action and fall into the cliff.
Q-learning learns values for the optimal policy, but the behavior policy still uses ε-greedy selection. With ε = 0.1, a random action is still selected with probability ε.
Performance While Learning
Q-learning can perform worse online than Sarsa while both agents are learning. Its ε-greedy behavior can occasionally take the agent from the edge route into the cliff, producing the −100 penalty and returning the agent to the start. Sarsa's values reflect this risk, so Sarsa favors the safer upper route.
This does not mean that Q-learning has learned the wrong policy. It means that online performance includes the actions actually taken during exploration, whereas the policy whose values are learned by Q-learning is the optimal policy. The shortest optimal route and the safest route under continued exploration can therefore be different.
Mistakes in the Comparison
Assuming that Q-learning must always take the edge route once it has learned that the route is optimal.
The learned policy and the behavior policy are different ideas in off-policy learning.
Fix:
Distinguish the policy whose values are learned from the actions the agent actually takes while exploring.Calling the upper route optimal simply because it is safer during exploration.
The edge route is the optimal route when the optimal choices are followed, even though it is dangerous under exploratory behavior.
Fix:
State which policy is being evaluated: the optimal policy or the ε-greedy policy being followed.Treating Sarsa and Q-learning as if they respond to the same risk.
Sarsa is on-policy and Q-learning is off-policy.
Fix:
Remember that Sarsa learns the policy being followed, whereas Q-learning learns the optimal policy.Using online penalties as proof that Q-learning learned a worse route.
Online performance includes exploratory actions, not only the policy represented by the learned values.
Fix:
Analyze performance and learned-policy values separately.
Check Your Understanding
In the cliff-walking task, explain why Sarsa can prefer the longer upper route even though Q-learning learns the edge route as optimal. Include the role of ε-greedy action selection and explain why Q-learning can still experience a cliff fall during learning.
Hints
- Start by stating what policy Sarsa learns values for.
- Then state what policy Q-learning learns values for.
- Connect ε = 0.1 to the possibility of a random action.
- Separate the route's learned value from the agent's online behavior.
On-policy means that the learned values correspond to the policy the agent is actually following. Off-policy means that the learned values can correspond to a different target policy from the behavior policy used to collect experience.
The Central Distinction
- The cliff-walking task separates the shortest optimal route from the safer route under exploration. Sarsa is on-policy, so it includes the risk of ε-greedy actions and prefers the longer upper route. Q-learning is off-policy, so it learns the optimal policy and prefers the short edge route. Q-learning can nevertheless perform worse during learning because its exploratory behavior can still enter the cliff and receive −100. Always distinguish the policy whose values are learned from the actions the agent actually takes.
Key Takeaways
- The cliff-walking task places a short route beside a dangerous cliff and a longer route farther from it.
- Sarsa learns the ε-greedy policy that the agent is actually following, so it accounts for exploration risk and prefers the safer route.
- Q-learning learns the optimal policy, which favors the shortest edge route when optimal choices are followed.
- With ε = 0.1, the agent still sometimes selects a random action, so Q-learning can fall into the cliff during exploration.
- Learned-policy quality and online performance are different: one describes what values are being learned, while the other describes what the exploring agent actually experiences.