Episodes and Transitions in Reinforcement Learning
Models let reinforcement-learning systems simulate experience rather than relying only on directly observed experience.
From Direct Experience to Simulation
A reinforcement-learning system does not always need to obtain every learning experience directly from the environment. A model can imitate the environment well enough to produce possible experience for inspection or planning. This changes the central question from “Can the system produce an outcome?” to “What kind of outcome does the model produce?”
One Simulated Transition
A transition is one simulated step from a starting state after an action is supplied. The input for this step is a state together with an action. The model then produces the possible result of that step, including the next state and reward information represented by the transition.
Reading a Transition Request
A model receives a state and an action. What level of simulation is being requested?
Identify the input: The input contains a state and an action.
Identify the scale: A state and action describe one step, so the request is transition-level rather than episode-level.
Identify the output meaning: A sample model gives one possible way that step could unfold. A distribution model gives all possible transitions together with their probabilities.
The request is for a transition. The model type determines whether the result is one possible transition or the full probability-weighted collection of transitions.
Sample and Distribution Models
| Model type | Input scale | Result |
|---|---|---|
| Sample model | State and action for a transition, or state and policy for an episode | One possible transition or episode |
| Distribution model | State and action for a transition, or state and policy for an episode | All possible transitions or episodes together with their probabilities |
Generating a Complete Episode
The same distinction extends beyond one step. To generate an episode, begin with a starting state and use a policy. The policy supplies the actions used as the simulation proceeds. A sample model produces one possible sequence of experience. A distribution model represents all possible episodes and the probabilities of those episodes.
From a Starting State to an Episode
A simulation begins with a starting state and a policy. What kind of result does each model type produce?
Start the simulation: Use the starting state as the beginning of the simulated experience.
Use the policy: The policy determines the actions used as the simulation proceeds.
Extend the experience: The model supplies possible next steps and rewards as the sequence continues.
Interpret the result: A sample model produces one possible episode. A distribution model produces all possible episodes together with the probabilities of those episodes.
The input determines the episode-level scale, while the model type determines whether the output is one possible episode or a probability-weighted collection of all possible episodes.
Episode Boundaries and Task Objectives
An episodic task is organized into multiple episodes. Each episode is a sequence of interactions between the agent and the environment that starts and ends at specific points. The objective of an episodic task is to maximize expected total reward across the episode.
In the maze-running example, the reward is +1 for escaping the maze and zero at all other times. This reward communicates that escaping is the desired achievement. If the agent does not improve at escaping, the reward system should be investigated to determine whether it has communicated the objective effectively.
Common Interpretation Mistakes
Treating every model output as a complete episode
A state and action describe one step, so the request is transition-level.
Fix:
Classify the input first. State plus action means transition; state plus policy means episode.Confusing one possible result with all possible results
Sample models return one possible transition or episode. Distribution models return all possibilities together with their probabilities.
Fix:
Identify whether the model is sample-based or distribution-based before interpreting its output.Ignoring the role of rewards in the task objective
The objective of an episodic task is to maximize expected total reward, and rewards guide the agent toward the desired objective.
Fix:
Interpret the episode together with its rewards and ask whether the reward system communicates the desired achievement.
Check Your Reasoning
A reinforcement-learning system is given a starting state and a policy. It uses a sample model. Predict the scale of the result and the kind of output the system should expect.
Hints
- A starting state plus a policy identifies the scale of the simulation.
- A sample model returns one possible result rather than all possible results with probabilities.
What do you think happens?
A model receives a state and an action. Is the requested result a transition or an episode?
Reveal answer
Answer: A transition
A state and action describe one simulated step. An episode requires a starting state together with a policy.
Compare these two requests: a starting state with an action, and a starting state with a policy. For each request, state whether it is transition-level or episode-level, then state what a distribution model would return.
Hints
- Use the input to determine the scale.
- For a distribution model, include all possible results and their probabilities.
Key Takeaways
- A model can imitate an environment's experience so a reinforcement-learning system can simulate experience for inspection or planning.
- A state and action describe one transition, while a starting state and policy describe an episode.
- A sample model produces one possible transition or episode.
- A distribution model represents all possible transitions or episodes together with their probabilities.
- Episodic tasks contain sequences of interactions, and their objective is to maximize expected total reward.
Key Takeaways
- Models allow reinforcement-learning systems to simulate possible experience rather than relying only on direct environment interaction.
- The input determines whether the simulation concerns one transition or a complete episode.
- The model type determines whether the result is one possible outcome or all possible outcomes with probabilities.
- Episodes organize interactions into bounded sequences, and rewards define the objective of maximizing expected total reward.