Concepts / TD Control Methods

TD Control Methods

The cliff-walking task requires safe movement from S to a goal in a grid world.

  • Programming

The Cliff-Walking Objective

The cliff-walking task places an agent in a grid world. The agent starts at S and must move to a goal while avoiding a cliff that occupies part of the grid. Reaching the goal is not enough by itself: entering the cliff violates the task, even when the action seemed to move toward the goal.

move through gridentering violates taskreach destinationSstart pointSafe cellavailable movement areaClifftask violationGoaldestination
Where are the start, safe cells, cliff, and goal, and what must the agent accomplish?

Route Trade-Offs

Every route must satisfy two requirements at once. It must make progress from S toward the goal, and it must keep the agent out of the cliff. A route that appears attractive because it moves toward the goal can still fail if it enters the cliff. This makes cliff walking useful for comparing how TD control methods behave while learning a route.

may enterreachesShort routedirect progressClifftask violationSafer routeavoids cliffGoaldestination
How do route choices trade off progress toward the goal against the risk of entering the cliff?

Evaluating One Candidate Route

An agent begins at S and considers a route that moves toward the goal but enters the cliff before arrival. Does this route satisfy the task?

Check progress: The route does move in the direction of the goal, so it appears to make progress.

Check safety: The route enters the cliff. The cliff is part of the task's prohibited area.

Judge the route: Because entering the cliff violates the task, the route does not satisfy the objective.

A route must both reach the goal and avoid entering the cliff. Progress toward the goal alone is insufficient.

ε-Greedy Action Selection

The compared TD control methods use an ε-greedy policy with ε = 0.1. At each decision, the policy uses the action currently regarded as best while retaining an exploratory component that can select a random action. Thus, the comparison does not evaluate methods only on a fixed route: the agent continues to balance using what it currently knows with trying alternatives.

select actiongreedy componentexploratory componentchoosechooseCurrent stateagent must actε = 0.1selection ruleBest-known actionuse current knowledgeSelected actionmove through gridRandom actionexplore alternative
How does the agent choose between the currently best-known action and an exploratory action?

Learning Performance Over Time

Performance viewWhat it measuresEvaluation period
Interim performanceBehavior during the first part of learningFirst 100 episodes
Asymptotic performanceAverage behavior over a much longer period after learning has largely developedAverage over 100,000 episodes

Interim performance asks how a method behaves early in training, using the first 100 episodes. Asymptotic performance asks about behavior over a much longer evaluation period, using an average over 100,000 episodes. A method can therefore be discussed in two different ways: by its early learning behavior and by its longer-run average behavior. One performance number should not be treated as the whole story.

The source describes these results as averages across repeated runs. The interim case uses averages across 50,000 runs, while the asymptotic case uses averages across 10 runs. The figure also marks the maximal values for the methods in the relevant performance views.

Reading a Method Comparison

  1. First identify the task objective: move from S to the goal without entering the cliff.
  2. Then check the action-selection condition: the compared methods use an ε-greedy policy with ε = 0.1.
  3. Next identify whether the result describes interim performance or asymptotic performance.
  4. Finally interpret the evaluation period and repeated-run averages before comparing the methods.
  • Treating movement toward the goal as the only requirement.

    Entering the cliff violates the task.

    Fix: Evaluate both progress toward the goal and avoidance of the cliff.

  • Treating interim and asymptotic performance as interchangeable.

    The two measures describe different stages of learning.

    Fix: Label results as interim or asymptotic and use the matching evaluation period.

  • Ignoring the policy used during the comparison.

    The action-selection policy is part of the comparison setup.

    Fix: Record the ε-greedy condition before interpreting performance.

Check Your Understanding

MEDIUM

Explain why a route that moves toward the goal can still fail the cliff-walking task. Then state how interim performance differs from asymptotic performance in the comparison.

Hints
  • Mention what happens when the agent enters the cliff.
  • Use the first 100 episodes for interim performance.
  • Use the 100,000-episode average for asymptotic performance.

Key Takeaways

  1. Cliff walking asks an agent to move from S to a goal without entering the cliff. A successful route must balance progress with safety. The compared TD control methods use an ε-greedy policy with ε = 0.1, so action selection includes both current knowledge and exploration. Interim performance covers the first 100 episodes, whereas asymptotic performance is an average over 100,000 episodes. These measurements describe different stages of learning and must be interpreted separately.

Key Takeaways

  • The cliff-walking objective is to reach the goal from S without entering the cliff.
  • A route must balance progress toward the goal with avoidance of the cliff.
  • The compared TD control methods use an ε-greedy policy with ε = 0.1.
  • Interim performance uses the first 100 episodes, while asymptotic performance uses an average over 100,000 episodes.
  • Early learning behavior and long-run average behavior are different evaluation views.