Concepts / Episodes and Time Steps in Reinforcement Learning

Episodes and Time Steps in Reinforcement Learning

MDP notation distinguishes individual objects from sets of possible objects.

  • Programming

A Symbol Can Describe One Thing or Many

Markov Decision Processes use a small set of symbols to describe what exists in the process and when it occurs. The main groups are states, actions, rewards, and time steps. The first question to ask about any symbol is whether it describes one item or a set of possible items. The second is whether it identifies a particular time.

A lowercase symbol often identifies an individual item, while an uppercase symbol can describe a set of possible items. A subscript such as t identifies a particular time step.

individual versus setindividual versus setindividual versus setsone stateSnonterminal statesaone actionS+inclusive state setrone rewardA(s)actions available in sRpossible rewards
How does the notation distinguish one state, action, or reward from a set of possible states, actions, or rewards?

The Main MDP Notation

IdeaIndividual notationSet or context notationMeaning
States or s′S and S+A state notation; S describes nonterminal states and S+ describes the inclusive state set.
ActionaA(s)An action, or the set of actions available in state s.
Action at a timeA_t—The action identified at time step t.
RewardrRAn individual reward, or the set of all possible rewards.
State at a timeS_t—The state identified at time step t.

Use the notation table by checking both the object type and any time subscript.

The symbols s and s′ refer to state notation for individual states. The symbols S and S+ describe state sets: S describes nonterminal states, while S+ describes the inclusive state set. For actions, a is an individual action and A(s) is the action set associated with state s. Rewards use r for an individual reward and R for the set of all possible rewards.

Following Time Through an Episode

A time subscript tells you which occurrence is being discussed. In S_t, t identifies the state associated with time step t. In A_t, t identifies the action associated with time step t. The subscript does not turn the symbol into a set; it adds a time reference to the notation.

identifiesidentifieslocatesadvances toidentifiesttime stepS_tstate at tA_taction at tt+1later time steprindividual rewardS_t+1state at t+1
What happens to the notation when the time-step index changes from t to t+1?

Reading a Time-Indexed Description

A description says: At time step 4, the agent takes action a and the state being discussed is S_4. Translate the time references into MDP notation.

Find the time index: The number 4 identifies the particular time step.

Find the state notation: S_4 means the state identified at time step 4.

Find the action notation: The action is written as a. If the notation emphasizes that it occurs at time step 4, it can be written as A_4.

The time-indexed forms are S_4 for the state at time step 4 and A_4 for the action at time step 4.

State-Specific Action Sets

A(s) describes the actions available in a particular state s. This is different from a, which denotes one action. The state inside the parentheses matters: A(s) connects the available-action set to the state being considered.

determines the context forcontains an available actionsone stateA(s)actions available in saone action
How is a state-specific action set related to one state and one individual action?

Reading an Interaction Sequence

The notation can be read as a sequence of information about an MDP interaction: identify the state being considered, identify an action, identify the reward, and attach a time index when the particular step matters. The symbols do different jobs, so keep the object type and the time reference separate while reading.

read action at the same timeread reward notationcontinue to later indexed stateS_tstate at tA_taction at trindividual rewardS_t+1state at t+1
How can the main MDP symbols be read in an ordered interaction description?

Translating Plain Language

From Words to Symbols

Translate this generated description: The process is considering one state at time step 2. The action taken at that time is one action from the actions available in that state. The result includes one reward.

Identify the state: One state is represented with state notation. Because the state is tied to time step 2, write S_2.

Identify the action: One action is represented by a. Because it is tied to time step 2, write A_2.

Identify the available-action set: The actions available in the state are represented by A(s).

Identify the reward: One reward is represented by r. The set of all possible rewards is represented by R, but the description asks for one reward.

The key symbols are S_2, A_2, A(s), and r.

translatetranslatetranslatetranslateone statestate objectS_2state at time 2one actionaction objectA_2action at time 2available actionsstate-specific setA(s)actions in state sone rewardreward objectrindividual reward
Which parts of a plain-language description correspond to the state, action, reward, time step, and possible-value set?

Common Notation Mistakes

  • Treating S as one particular state.

    S describes the nonterminal state set, not one individually time-indexed state.

    Fix: Use S_t for a state identified at time t, and use S when discussing the nonterminal state set.

  • Treating A(s) as one action.

    A(s) describes the actions available in state s; a denotes one action.

    Fix: Use a for one action and A(s) for the state-specific action set.

  • Ignoring the time subscript.

    The subscript identifies which time step the action belongs to.

    Fix: Read the subscript as part of the meaning: A_2 is the action at time step 2, while A_5 is the action at time step 5.

  • Using R when the description names one reward.

    R denotes the set of all possible rewards, while r denotes an individual reward.

    Fix: Use r for one reward and R for the possible-reward set.

For every symbol, ask two questions in order: Does it describe one item or a set of possible items? Does it identify a particular time? This simple check separates s from S, a from A(s), r from R, and unindexed notation from forms such as S_t and A_t.

Practice: Decode the Notation

EASY

For each item, identify whether it represents one item, a set of possible items, or an item identified at a particular time: s, S+, A(s), A_7, r, and R. Then write the symbol for the state at time step 3 and the action at time step 3.

Hints
  • Check capitalization and parentheses first.
  • A subscript identifies a time step.
  • Use the notation table to distinguish an individual reward from the reward set.

What do you think happens?

Which symbol best represents the action at time step 7: a, A(s), or A_7?

  • a
  • A(s)
  • A_7
Reveal answer

Answer: A_7

a identifies an individual action without a time reference, A(s) identifies the actions available in state s, and A_7 identifies the action at time step 7.

The Notation Checklist

  1. Use s or s′ for individual state notation; use S for nonterminal states and S+ for the inclusive state set.
  2. Use a for one action and A(s) for the actions available in state s.
  3. Use r for one reward and R for the set of all possible rewards.
  4. Use a subscript such as t to identify the time step associated with a state or action, as in S_t and A_t.
  5. Translate descriptions by checking both dimensions: individual versus set, and unindexed versus time-indexed.

Key Takeaways

  • MDP notation describes states, actions, rewards, and when they occur.
  • Individual objects and sets of possible objects use different notation.
  • A(s) is the set of actions available in state s, while a is one action.
  • Subscripts such as t identify a particular time step.
  • A reliable translation method is to classify each symbol by object type and time reference.