Concepts / Q-learning: Off-policy TD Control

Q-learning: Off-policy TD Control

The cliff-walking task separates the shortest optimal route from the safer route under exploration.

  • Programming

The Dangerous Shortcut

Q-learning becomes easier to understand when the shortest route is not the safest route. The cliff-walking task creates exactly this situation: an agent must travel from a start state to a goal state, while a dangerous region lies beside one possible route. The task exposes a difference between two questions: should learned values describe the policy the agent is actually following, or should they describe the best policy the agent could follow?

short pathreach goalexploration near edgelonger pathreach goalStartEdge routeshortest; beside cliffCliff-100; return to startUpper routelonger; saferGoal
How do the two routes from start to goal differ in length, proximity to the cliff, and risk of receiving a large negative reward?

A Walk Through the Grid

Imagine a gridworld with a start state and a goal state. At every step, the agent can move up, down, right, or left. Most transitions provide a reward of −1. A cliff occupies one region of the grid. If the agent enters the cliff, it receives a much worse reward of −100 and is immediately returned to the start state.

Comparing Two Complete Journeys

Compare an edge route that is shortest under perfect action choices with an upper route that stays farther from the cliff.

Choose the edge route: The agent travels along the edge toward the goal. This is the optimal route when the optimal choices are followed, but it places the agent next to the cliff.

Explore near the cliff: If action selection explores a random action near the cliff, the agent can enter the cliff, receive −100, and return to the start.

Choose the upper route: The agent takes a longer path through the upper part of the grid. This route is less exposed to the cliff, so exploratory choices do not create the same cliff risk along the route.

Compare the objectives: The edge route is shortest when the optimal actions are followed. The upper route is safer for a policy that still sometimes explores.

The task separates shortest-route optimality from safety during exploratory behavior.

value estimates developvalue estimates developSarsa learninglonger, safer routeQ-learning duringlearningedge route plus cliff fallsSarsa after learningsafer routeQ-learning afterlearningshortest optimal route
How do the agents' chosen paths and experienced rewards differ during learning and after their value estimates have converged?

Sarsa Follows Its Own Behavior

Sarsa is on-policy. Its learned values correspond to the policy the agent is actually following. In the cliff-walking task, that policy includes ε-greedy action selection, so the agent sometimes chooses a random action instead of its currently preferred action. Sarsa therefore accounts for the danger created by exploration near the cliff.

Because the edge route places the agent beside the cliff, an exploratory action there can produce the −100 penalty. Sarsa treats that possibility as part of the route's value. As a result, it prefers the longer path through the upper part of the grid. The route is not shorter, but it is safer for the exploratory policy the agent is actually using.

selecttake actionselect next actionlearn from actual policyCurrent stateChosen actionε-greedy behaviorNext stateNext chosen actionalso ε-greedyUpdated route valuescliff risk included
How does Sarsa update values using the action the agent actually intends to take next, causing it to prefer the longer route away from the cliff?

Sarsa does not merely ask which action would be best in an ideal, non-exploring run. It learns values for the policy that includes the agent's exploratory choices.

Q-learning Chooses an Ideal Target

Q-learning is off-policy. It learns the values of the optimal policy rather than the values of the exploratory policy currently being used. In the cliff-walking task, this means Q-learning can prefer the edge route because that route reaches the goal most effectively when the optimal choices are followed.

The behavior policy and the learned policy are therefore different ideas. The behavior policy determines what the agent actually does while it is exploring. The policy whose values Q-learning learns is the optimal policy. Q-learning can learn that the edge route is optimal even while an ε-greedy behavior policy occasionally takes an action that sends the agent into the cliff.

acttransitionevaluate best optionlearnCurrent stateBehavior actionε-greedy choiceNext stateBest next actionoptimal targetOptimal-policy valuesshort edge route
How does Q-learning learn from the best possible next action even when the behavior policy may choose an exploratory action?

How ε-greedy Exploration Works

The example uses ε-greedy action selection with ε = 0.1. Under this rule, the greedy action is selected with probability 1 − ε, while a random action is selected with probability ε. Thus, the agent normally chooses its currently preferred action but still sometimes explores another action.

at each stepusuallysometimesactactStateChoose actionε = 0.1Greedy actionprobability 1 − εRandom actionprobability εEnvironmenttransition
At each state, how does ε determine whether the agent follows the highest-valued action or randomly explores another action?

What do you think happens?

Suppose Q-learning has learned that the edge route is optimal, but ε-greedy exploration is still active. What can happen during an actual run?

  • The agent must always stay on the edge route
  • The agent can still choose a random action and fall into the cliff
  • The agent can no longer reach the goal
Reveal answer

Answer: The agent can still choose a random action and fall into the cliff.

Q-learning learns values for the optimal policy, but the behavior policy still uses ε-greedy selection. With ε = 0.1, a random action is still selected with probability ε.

Performance While Learning

Q-learning can perform worse online than Sarsa while both agents are learning. Its ε-greedy behavior can occasionally take the agent from the edge route into the cliff, producing the −100 penalty and returning the agent to the start. Sarsa's values reflect this risk, so Sarsa favors the safer upper route.

This does not mean that Q-learning has learned the wrong policy. It means that online performance includes the actions actually taken during exploration, whereas the policy whose values are learned by Q-learning is the optimal policy. The shortest optimal route and the safest route under continued exploration can therefore be different.

guides preferred actionexploration can causeaccounts for riskQ-learning valuesoptimal edge routeSarsa valuesexploration risk includedQ-learning behaviorε-greedy actionsUpper routelonger, saferCliff fall−100; return to start
Why can Q-learning learn the shortest optimal route while still sometimes falling off the cliff during exploration?

Mistakes in the Comparison

  • Assuming that Q-learning must always take the edge route once it has learned that the route is optimal.

    The learned policy and the behavior policy are different ideas in off-policy learning.

    Fix: Distinguish the policy whose values are learned from the actions the agent actually takes while exploring.

  • Calling the upper route optimal simply because it is safer during exploration.

    The edge route is the optimal route when the optimal choices are followed, even though it is dangerous under exploratory behavior.

    Fix: State which policy is being evaluated: the optimal policy or the ε-greedy policy being followed.

  • Treating Sarsa and Q-learning as if they respond to the same risk.

    Sarsa is on-policy and Q-learning is off-policy.

    Fix: Remember that Sarsa learns the policy being followed, whereas Q-learning learns the optimal policy.

  • Using online penalties as proof that Q-learning learned a worse route.

    Online performance includes exploratory actions, not only the policy represented by the learned values.

    Fix: Analyze performance and learned-policy values separately.

Check Your Understanding

MEDIUM

In the cliff-walking task, explain why Sarsa can prefer the longer upper route even though Q-learning learns the edge route as optimal. Include the role of ε-greedy action selection and explain why Q-learning can still experience a cliff fall during learning.

Hints
  • Start by stating what policy Sarsa learns values for.
  • Then state what policy Q-learning learns values for.
  • Connect ε = 0.1 to the possibility of a random action.
  • Separate the route's learned value from the agent's online behavior.

On-policy means that the learned values correspond to the policy the agent is actually following. Off-policy means that the learned values can correspond to a different target policy from the behavior policy used to collect experience.

The Central Distinction

  1. The cliff-walking task separates the shortest optimal route from the safer route under exploration. Sarsa is on-policy, so it includes the risk of ε-greedy actions and prefers the longer upper route. Q-learning is off-policy, so it learns the optimal policy and prefers the short edge route. Q-learning can nevertheless perform worse during learning because its exploratory behavior can still enter the cliff and receive −100. Always distinguish the policy whose values are learned from the actions the agent actually takes.

Key Takeaways

  • The cliff-walking task places a short route beside a dangerous cliff and a longer route farther from it.
  • Sarsa learns the ε-greedy policy that the agent is actually following, so it accounts for exploration risk and prefers the safer route.
  • Q-learning learns the optimal policy, which favors the shortest edge route when optimal choices are followed.
  • With ε = 0.1, the agent still sometimes selects a random action, so Q-learning can fall into the cliff during exploration.
  • Learned-policy quality and online performance are different: one describes what values are being learned, while the other describes what the exploring agent actually experiences.