Agent-Environment Interaction Basics
The agent-environment interaction is a repeating sequence of state, action, reward, and next state.
One Decision Cycle
Reinforcement learning can be understood as a repeating interaction between an agent and an environment. The agent receives information describing the current state, chooses an available action, and then encounters the result of that action as a reward and a new state. The cycle then repeats from the new state.
The central sequence is S_t, A_t, R_t+1, S_t+1. The agent acts using the current state, and the reward and next state describe what follows that action.
Tracing the State Change
A Single Abstract Interaction
Trace one interaction when the current state is S_t and the agent must choose from the actions available in that state.
Observe: At time t, the agent receives information representing the current environment state, written as S_t.
Choose: The agent selects an available action, written as A_t. The action must belong to the set of actions available in the current state, represented by A_t ∈ A(S_t).
Respond: After the action, the environment provides a numerical reward and the agent encounters a new state.
Label the result: The reward is written R_t+1 and the new state is written S_t+1. Both describe what follows the action selected at time t.
The complete interaction is S_t, then A_t, then R_t+1 and S_t+1. The next cycle can use S_t+1 as its current state.
A state is a representation of the environment, not necessarily the entire environment. This matters because the agent makes its decision from the information represented by S_t. The action also depends on the current state: A_t must come from the actions available in that state. The notation A_t ∈ A(S_t) records this availability constraint.
How a Policy Chooses
A policy describes action probabilities conditioned on the current state. In other words, the policy connects the state the agent currently observes with probabilities for selecting the actions available in that state.
The policy does not describe one fixed action independently of the state. Different states can be associated with different action probabilities. The probabilities shown in an illustration are only an example of the policy idea; the important structure is that the current state conditions the action probabilities.
Generated illustration: suppose a current state makes two actions available, Action 1 and Action 2. A policy could associate one probability with Action 1 and another probability with Action 2 for that state. If the state changes, the policy can associate different probabilities with the available actions. The exact numerical probabilities are not supplied by the source; they are useful only for illustrating the mapping structure.
Reading the Time Indices
| Notation | Meaning in the interaction |
|---|---|
| S_t | The current state received at time t. |
| A_t | The action selected at time t. |
| R_t+1 | The numerical reward that follows the action selected at time t. |
| S_t+1 | The new state encountered after the action selected at time t. |
The notation separates the decision time from the resulting reward and state.
The index on A_t identifies the time at which the agent chooses the action. The indices on R_t+1 and S_t+1 emphasize that the reward and new state follow that action one step later in this model. Therefore, R_t+1 is not labeled with t merely because the agent chose A_t; its notation highlights the result after the action.
Why Steps Are Discrete
This interaction model organizes reinforcement learning into separate time steps. At one step, the agent receives a state and selects an action. The resulting reward and new state belong to the following step. Using discrete time makes the interaction easier to describe because each decision and its consequences have a clear position in the sequence.
Discrete time is a simplifying model for this presentation, not a claim that every reinforcement-learning situation must be represented this way. The source notes that many ideas can also be extended to continuous time. Here, separate steps provide a clear framework for tracing the repeating decision cycle.
Treating the reward as if it were selected before the action.
The action A_t is selected at time t, while R_t+1 describes what follows that action one step later.
Fix:
Trace the decision first: S_t, then A_t, then R_t+1 and S_t+1.Assuming the policy chooses the same action independently of the state.
A policy connects the current state with probabilities for selecting available actions, and different states can have different action probabilities.
Fix:
Ask which state is being considered, then describe the probabilities for actions available in that state.Assuming a state must contain the entire environment.
A state is a representation of the environment and is not necessarily the entire environment.
Fix:
Interpret S_t as the information representation used in the interaction model.Ignoring action availability.
The action must come from the set available in the current state, represented by A_t ∈ A(S_t).
Fix:
Check the current state before identifying the actions the agent can select.
Check the Sequence
What do you think happens?
An agent has just observed S_t and selected A_t. Which notation identifies the reward and new state that follow this action?
Reveal answer
Answer: R_t+1 and S_t+1
The action is selected at time t. The reward and new state that follow it are written R_t+1 and S_t+1, emphasizing the next step in the interaction.
Describe one complete interaction in your own words using all four symbols S_t, A_t, R_t+1, and S_t+1. Then explain what the policy contributes at the point where A_t is selected.
Hints
- Begin with the state the agent currently receives.
- Mention that A_t must be available in S_t.
- End with the reward and new state that follow the action.
- Explain that the policy supplies action probabilities conditioned on the current state.
Interaction Cycle Summary
- The repeating cycle is current state, selected action, resulting reward, and next state.
- S_t is the current state, A_t is the action selected at time t, R_t+1 is the following reward, and S_t+1 is the following state.
- A policy describes probabilities for selecting available actions conditioned on the current state.
- The constraint A_t ∈ A(S_t) records that the selected action must be available in the current state.
- Discrete time separates the interaction into clear steps, while many ideas can also be extended to continuous time.
Key Takeaways
- A reinforcement-learning interaction repeats the sequence S_t, A_t, R_t+1, S_t+1.
- The agent uses the current state to select an available action, and the environment then provides a reward and a new state.
- A policy maps a current state to probabilities for selecting available actions.
- The time indices distinguish the decision at time t from the reward and new state that follow one step later.
- Discrete time makes the interaction easier to describe, although many ideas can be extended to continuous time.