Concepts / State Visitation in Reinforcement Learning

State Visitation in Reinforcement Learning

An episodic on-policy distribution depends on both episode starting states and transitions within episodes.

  • Programming

From Episodes to State Frequencies

When reinforcement learning uses an on-policy distribution, the central question is how often the learning process encounters each state while following its policy. In an episodic task, that frequency is shaped by the states from which episodes begin and by the transitions that occur afterward. The starting-state probabilities therefore matter even when the policy and transitions remain unchanged.

How Visits Accumulate

A useful way to organize the calculation is to keep an accounting record for every state. A state can receive expected visits in two ways: an episode can start there, or the process can transition into it during the episode. The expected visit count combines both contributions. Once all expected visits have been collected, they can be compared across states.

begins episodesadds state timenormalizeStartingprobabilitiesepisode entryTransitionswithin episodeExpected visitsη(s)On-policydistributionnormalized shares
How does probability flow from episode starting states through transitions and accumulate into visits to each state?

η(s) is the average number of time steps spent in state s during one episode. It is an expected visit count, so it does not have to be a whole number. Across many possible episodes, a state might be visited once in some episodes and twice in others; the average can then be fractional.

η(s) describes the state’s expected contribution within an episode rather than one particular episode. It also differs from a probability of being in that state at one particular time step. The count accumulates time across the episode, so repeated visits contribute repeatedly to the expected count.

A Visit-Count Calculation

Normalizing three expected visit counts

Suppose a constructed episodic example has three states: A, B, and C. Their expected visit counts are η(A) = 1, η(B) = 2, and η(C) = 1. Find the on-policy distribution.

Add the expected visits: The total expected episode time represented by the record is 1 + 2 + 1 = 4 time steps.

Preserve each state's share: State A contributes 1 of the 4 expected visits, state B contributes 2 of the 4, and state C contributes 1 of the 4.

Normalize: Divide each expected visit count by the total of 4. The resulting shares are A: 1/4, B: 2/4, and C: 1/4.

Check the result: The normalized values add to 1/4 + 2/4 + 1/4 = 1, so they form a distribution.

The on-policy distribution is A: 0.25, B: 0.50, and C: 0.25.

sumnormalizing valuecountsnormalizedη(A), η(B), η(C)1, 2, 1Total expected visits4Divide by totaleach count ÷ 4On-policy shares0.25, 0.50, 0.25
How do raw expected visit counts become probabilities that sum to 1?

The important separation is between counting and normalizing. The values 1, 2, and 1 retain the scale of expected episode time. The values 0.25, 0.50, and 0.25 retain only each state’s share of the total. Normalization does not change which state has the largest share; it converts the accounting record into probabilities whose total is one.

Why Starting Probabilities Matter

Changing the probabilities of episode starting states can change the expected visit count for every state reached later. This is because starting probabilities determine which parts of the transition process are sampled more often. The policy and transitions may stay the same, yet the episodic on-policy distribution can change because the mixture of starting episodes has changed.

contributescontributespropagates timepropagates timeStart mixture 1more episodes from AStart mixture 2more episodes from BVisit record 1A, B, C sharesVisit record 2A, B, C sharesSame transitionsunchanged
How can changing the probabilities of starting states change the final state-visitation distribution even when the policy and transitions stay the same?

When analyzing an episodic on-policy distribution, write down the starting-state contribution and the transition contribution separately before combining them. This prevents the starting probabilities from being mistaken for the final distribution.

Mistakes in Visit Accounting

  • Using starting-state probabilities as the final on-policy distribution.

    Expected visits include both episodes that start in a state and visits produced by transitions into states during the episode.

    Fix: Compute the expected visit count for every state first, then normalize the counts by their total.

  • Treating η(s) as a one-time-step probability.

    η(s) is the average number of time steps spent in state s during one episode, so it can include accumulated or repeated visits and can be fractional as an average.

    Fix: Interpret η(s) as an expected episode-level visit count.

  • Assuming expected visit counts already form a distribution.

    Their total is 4 rather than 1.

    Fix: Divide every expected visit count by the sum of all expected visit counts.

  • Expecting every η(s) value to be a whole number.

    An average across possible episodes can be fractional.

    Fix: Allow η(s) to be a fractional expected value.

Practice the Conversion

EASY

A constructed episodic task has expected visit counts η(X) = 3, η(Y) = 1, and η(Z) = 2. Calculate the total expected visits and then determine the normalized on-policy share for each state.

Hints
  • Add the three expected visit counts to obtain the total.
  • Divide each state’s count by that total.
  • Check that the three normalized shares add to 1.

What do you think happens?

In the practice task, which state receives the largest normalized share?

  • X
  • Y
  • Z
  • All three are equal
Reveal answer

Answer: X

Normalization divides every count by the same total, so the state with the largest expected visit count keeps the largest share. Here, X has 3 expected visits, compared with 2 for Z and 1 for Y.

The Complete Reasoning Path

  1. Identify how often episodes begin in each state.
  2. Account for visits produced by those starting states and by transitions during the episode.
  3. Record the resulting expected visit count η(s) for every state.
  4. Add all expected visit counts to obtain the total expected episode time represented by the record.
  5. Divide each state’s expected visit count by that total to obtain probabilities that sum to one.

Key Takeaways

  • An episodic on-policy distribution depends on both episode starting states and transitions within episodes.
  • η(s) is the average number of time steps spent in state s during one episode.
  • Expected visit counts include visits from starting in a state and visits from transitioning into it.
  • Expected visit counts are not probabilities until each count is divided by the total expected visit count.
  • The normalized values preserve each state’s share of expected episode time and sum to one.