States, Actions, and Rewards in Reinforcement Learning
MDP notation distinguishes individual objects from sets of possible objects.
Reading MDP Symbols
A Markov Decision Process uses a small set of symbols to describe what exists in the process and when it occurs. The main groups are states, actions, rewards, and time steps. A reliable way to read any symbol is to ask two questions: does it describe one item or a set of possible items, and does it identify a particular time?
Capital and lowercase letters often separate a collection of possible objects from one individual object. A subscript such as t identifies a particular time.
States and State Sets
A state is one individual condition represented with a lowercase symbol such as s. The notation s′ is also used for a state. A state set describes possible states rather than one particular state. The source notation uses S for the set of nonterminal states and S+ for the inclusive state set. Therefore, s can refer to one state, while S refers to the collection of nonterminal states in which a state may occur.
Separating one state from the state set
A description refers to one particular state and also to all nonterminal states that may occur. Which symbols represent the two ideas?
Identify the individual: Use s for one state. The notation s′ is another notation used for a state.
Identify the collection: Use S for the set of nonterminal states.
Check inclusiveness: Use S+ when the notation refers to the inclusive state set.
One state is represented by s or s′; the nonterminal state set is S; the inclusive state set is S+.
Actions and Availability
An individual action is written as a. The notation A(s) represents the action set associated with state s: the actions available in that state. A_t identifies an action at time t. These symbols describe different levels of information. The lowercase a names one action, A(s) names the available-action set for a state, and A_t names the action selected or considered at a particular time.
Reading action notation
A plain-language description says that a state has a collection of available actions and that one action occurs at time t. Translate the two ideas into notation.
Name the collection: Use A(s) for the actions available in state s.
Name one action: Use a for an individual action.
Add timing: Use A_t when the action is identified as occurring at time t.
The available-action collection is A(s), one action is a, and a time-indexed action is A_t.
Time-Indexed Notation
A subscript t tells you that the symbol is being identified at a particular time step. Thus S_t refers to a state at time t, and A_t refers to an action at time t. The subscript does not turn the symbol into a set. It adds timing information to the object being described.
What do you think happens?
What changes when the subscript t is added to a state symbol?
Reveal answer
Answer: It identifies a state at a particular time.
The subscript supplies timing information. It does not change an individual object into a set.
The MDP Event Sequence
The notation can also be read as a time-ordered description. At a time step t, the process has a state S_t, an action A_t is associated with that time, and a reward R_t can be identified with the step. A later state can be written with a later time index, such as S_{t+1}. The important reading habit is to use the subscript to keep the objects tied to their time steps.
This sequence is a notation-reading aid: the subscripts show which objects belong to time t and which state belongs to the later time t+1.
Rewards and Reward Sets
An individual reward is written as r. R represents the set of all possible rewards. As with states and actions, the notation separates one object from a collection of possible objects.
Reading reward notation
A description refers first to one reward and then to every reward that could occur. Which symbols match the two descriptions?
Identify one reward: Use lowercase r for an individual reward.
Identify the collection: Use uppercase R for the set of all possible rewards.
Check the pattern: The same individual-versus-set distinction appears in the state and action notation.
One reward is r, while the set of all possible rewards is R.
Common Notation Mistakes
Treating S as one particular state
S describes the set of nonterminal states, while s or s′ describes an individual state.
Fix:
Use lowercase s or s′ for one state and uppercase S for the nonterminal state set.Treating A(s) as one action
A(s) describes the action set associated with state s.
Fix:
Use a for one action and A(s) for the actions available in state s.Ignoring the subscript
The subscript t identifies the action at a particular time.
Fix:
Read the base symbol and the time index separately: A_t is an action identified at time t.Confusing r with R
r denotes an individual reward, while R denotes the set of all possible rewards.
Fix:
Match lowercase r to one reward and uppercase R to the full possible-reward set.
Notation Translation Practice
Translate each description into the notation taught in this article: one state; the nonterminal state set; the inclusive state set; one action; the actions available in state s; one action at time t; one reward; and the set of all possible rewards.
Hints
- Use lowercase notation for individual states, actions, and rewards.
- Use S, S+, A(s), and R for the specified sets.
- Use the subscript t when the description identifies time t.
Answer key
Translate the eight descriptions into MDP notation.
State entries: One state is s or s′; the nonterminal state set is S; the inclusive state set is S+.
Action entries: One action is a; the actions available in state s are A(s); an action at time t is A_t.
Reward entries: One reward is r; the set of all possible rewards is R.
s or s′; S; S+; a; A(s); A_t; r; R.
Key Takeaways
- Lowercase symbols such as s, s′, a, and r represent individual states, actions, or rewards.
- Uppercase or functional notation can represent sets: S is the nonterminal state set, S+ is the inclusive state set, A(s) is the action set for state s, and R is the set of all possible rewards.
- A subscript such as t identifies a particular time step, as in S_t and A_t.
- To translate MDP notation, first decide whether the description refers to one item or a set, then check whether a time index is present.