Concepts / Episodic Tasks in Reinforcement Learning

Episodic Tasks in Reinforcement Learning

The on-policy distribution describes how state probability is distributed under a policy.

  • Programming

The Question Behind the Distribution

When reinforcement learning evaluates a policy, it does not necessarily treat every state as equally important. The useful question is: under this policy, which states are encountered, and how often? The on-policy distribution answers this question by describing how state probability is distributed under the policy.

The on-policy distribution is about states encountered while following a policy. It is not simply a list of all states in the task.

In an episodic task, this distribution depends on the complete way an episode unfolds. The episode-start rule determines where visits begin, and policy-following transitions determine where later visits occur.

Starting States Shape Later Visits

policy-following transitionspolicy-following transitionsStart rule Ainitial-state distributionLater statesvisits under policyStart rule Bdifferent initial-statedistributionLater statesvisits under same policy
How does changing the distribution of starting states alter the probabilities of visiting later states under the same policy?

Initial-state selection is part of the distribution calculation in an episodic task. Two settings can use the same policy and the same transition behavior but still produce different expected visits if their episode-start rules differ.

A state that receives more episode starts can accumulate more expected visits before normalization. A state that is rarely selected initially may still receive many visits later if policy-following transitions lead into it. Therefore, initial-state choices affect not only the first state in an episode, but potentially the entire on-policy distribution.

Tracing Visits Through an Episode

begin visitsadd visitsnormalizeEpisode startsinitial-state choicesPolicy transitionslater state visitsExpected visitscounts by stateState probabilitiesnormalized distribution
How do episode starts and policy-driven transitions combine to produce the expected number of visits to each state?

A useful way to trace the calculation is to follow the source of every visit. Some visits occur immediately because the episode-start rule places an episode in a state. Other visits occur later because policy-controlled transitions move the agent from states already being visited into additional states.

  1. Identify how episodes begin and which states receive those starts.
  2. Follow the transitions produced while the policy is followed.
  3. Add the contributions from initial visits and later transition-based visits.
  4. Obtain an expected visit count for each state.
  5. Normalize those expected counts to obtain probabilities that sum to one.

This order matters. First determine how many visits each state is expected to receive. Only after that should you convert those counts into a probability distribution.

Episodes and Their Objective

An episode is a sequence of interactions between the agent and the environment that starts and ends at specific points. Episodic tasks are composed of multiple episodes.

Organizing a task into episodes gives the agent repeated sequences with defined beginnings and endings. Each episode supplies a path through the environment: it starts according to the initial-state rule, continues through interactions under the policy, and ends at the task's specified endpoint.

The objective of an episodic task is to maximize the expected total reward. Rewards guide the agent toward the desired objective, so the reward system must communicate what achievement is wanted.

In the maze-running example, the reward is positive one for escaping the maze and zero at all other times. If the agent does not improve at escaping, the reward system should be investigated to determine whether it has effectively communicated the desired achievement.

A Worked Visit Calculation

Separate expected visits from probabilities

Suppose an episodic calculation produces expected visit counts of 2 for state A, 1 for state B, and 1 for state C. Determine the resulting on-policy state probabilities.

Record the visit counts: The expected visits are A: 2, B: 1, and C: 1. These are counts of expected visits, not yet probabilities.

Find the total expected visits: The counts together represent 4 expected visits.

Normalize each count: Convert each state's expected visit count into its share of the total: state A receives one-half of the total, while states B and C each receive one-quarter.

Check the distribution: The resulting state probabilities add to one, as a probability distribution should.

The normalized on-policy distribution assigns one-half to state A, one-quarter to state B, and one-quarter to state C.

The same separation applies when initial-state choices are involved. First include the visits supplied by episode starts and the additional visits supplied by policy-following transitions. Then normalize the resulting expected counts.

Mistakes in Distribution Reasoning

  • Treating every state as equally important

    The on-policy distribution describes which states are encountered, and how often, under a particular policy.

    Fix: Trace episode starts and policy-following transitions to determine expected visits.

  • Ignoring the initial-state distribution

    Initial-state selection is part of the distribution calculation for episodic tasks.

    Fix: Include the episode-start rule before accounting for later transitions.

  • Stopping at expected visit counts

    Expected visits are counts. They do not yet form probabilities that sum to one.

    Fix: Normalize the expected visit counts.

  • Mixing the visit and normalization stages

    This makes it difficult to tell whether an error came from modeling visits or from converting counts into probabilities.

    Fix: Keep the calculation in two layers: expected visits first, normalization second.

  • Assuming a weak reward system will still clearly guide learning

    Ineffective reward systems can hinder learning and may not communicate the desired objective effectively.

    Fix: Investigate whether the reward system clearly guides the agent toward the task objective.

When solving an episodic on-policy distribution problem, write two separate parts of the reasoning. In the first part, account for initial-state probabilities and policy-following transitions. In the second part, normalize the resulting expected visit counts. This makes the source of each result easier to inspect.

Practice the Two-Layer Method

MEDIUM

A task uses multiple episodes. You are given an initial-state rule and a policy that determines later transitions. Explain, in words, how you would calculate the on-policy distribution without treating all states as equally important.

Hints
  • Begin with the states selected when episodes start.
  • Add visits produced by transitions while following the policy.
  • Do not call the result a probability distribution until the expected visit counts have been normalized.

A strong answer should distinguish three ideas: where visits begin, how later visits arise, and how expected counts become probabilities. It should also recognize that changing the initial-state rule can change the final on-policy distribution even when the policy and transition behavior remain the same.

Key Takeaways

  1. The on-policy distribution describes how state probability is distributed under a policy. In episodic tasks, its calculation includes both the initial-state distribution and the transitions that occur while following the policy. Expected visits must be calculated before normalization, because normalization converts visit counts into probabilities that sum to one. Episodes have defined starts and ends, their objective is to maximize expected total reward, and rewards guide the agent toward that objective.

Key Takeaways

  • The on-policy distribution tells us how often states are encountered under a policy.
  • Episode-start choices influence expected visits and can change the final distribution.
  • Expected visits come from both episode starts and policy-following transitions.
  • Normalize expected visit counts to obtain probabilities that sum to one.
  • An episodic task has defined starts and ends, seeks to maximize expected total reward, and relies on rewards to guide learning.