On-Policy Control with Sarsa
Sarsa controls the acrobot through repeated online interaction rather than through access to a simulator model.
Learning Without a Simulator Model
An acrobot is a two-link robotic arm. In the swing-up task, a reinforcement learning agent applies torques until the goal is reached. The important feature is how the agent learns: it controls the simulated arm through repeated interaction without access to the simulator itself. This makes the task an online, model-free control problem, treated as though the agent were interacting with a physical acrobot.
Sarsa controls the acrobot by repeatedly observing the current situation, choosing a torque, applying it, and continuing the interaction until the goal is reached.
The Four-Part Acrobot State
The controller needs a description of the acrobot's current condition. The direct representation contains four continuous quantities: the position of the first link, the velocity of the first link, the position of the second link, and the velocity of the second link. They are written as θ1, θ̇1, θ2, and θ̇2.
| Quantity | Symbol | What it describes |
|---|---|---|
| First-link position | θ1 | Position of the first link |
| First-link velocity | θ̇1 | Velocity of the first link |
| Second-link position | θ2 | Position of the second link |
| Second-link velocity | θ̇2 | Velocity of the second link |
The four continuous quantities in the direct acrobot state
From Continuous State to Action Value
The controller combines tile coding, linear function approximation, Sarsa(λ), and replacing traces. Tile coding supplies a representation for the continuous four-dimensional state space. Linear function approximation then uses that representation to approximate the value function rather than storing a separate exact value for every possible point in the region.
Representing One Controller Situation
Suppose the controller receives one current acrobot situation described by θ1, θ̇1, θ2, and θ̇2. How does that situation move through the learning configuration?
Describe: The current situation is described by the four direct state variables: the two link positions and the two link velocities.
Represent: Tile coding represents the point in the continuous four-dimensional state space.
Evaluate: Linear function approximation uses the tile representation to approximate the value associated with an action.
Control: The Sarsa controller uses the configured value representation while selecting and applying a torque.
The controller does not need a separate exact table entry for every possible point in the continuous state region.
Action Choice and Episode Reset
Every episode begins with both links hanging straight down and at rest. The learning agent then applies torques. Interaction continues until the goal is reached. At that point, the acrobot returns to the same initial resting position and another episode begins.
The action set is small and discrete, so the setup uses a separate set of tilings for each action. This means the representation is not only a shared description of the physical state. Each available action has its own tiling set associated with evaluating that action in the state.
Traces in the Learning Configuration
Replacing traces are another component of the stated Sarsa(λ) configuration. Their role should be kept distinct from the role of tile coding. Tile coding represents the continuous state, action-specific tilings distinguish the discrete actions, and replacing traces belong to the learning configuration used alongside linear function approximation and Sarsa(λ).
| Component | Primary role |
|---|---|
| Direct state | Describes the current acrobot condition with four continuous quantities |
| Tile coding | Represents the continuous four-dimensional state space |
| Action-specific tilings | Associates a separate tiling set with each discrete action |
| Linear function approximation | Approximates the value function using the tile representation |
| Replacing traces | Forms part of the Sarsa(λ) learning configuration |
Check Your Understanding
Explain why the acrobot task is described as online and model-free. Then list the four direct state variables and state the role of action-specific tilings.
Hints
- Online refers to learning through repeated interaction with the arm.
- Model-free refers to controlling the arm without access to the simulator itself.
- The direct state contains two positions and two velocities.
- Each discrete action has its own tiling set.
What do you think happens?
After the acrobot reaches the goal, what happens before the next episode?
Reveal answer
Answer: It returns to the initial resting position.
Each episode ends when the goal is reached. The acrobot is then returned to the same position in which the episode began: both links hanging straight down and at rest.
Treating the acrobot state as a small list of individually enumerated states.
The direct state is continuous, and its possible values form a bounded region in four dimensions.
Fix:
Use tile coding to represent the continuous state space and linear function approximation to approximate values.Assuming one shared tiling set is the complete action representation.
The stated setup uses a separate set of tilings for each discrete action.
Fix:
Keep the physical state description distinct from the action-specific tiling sets used for action evaluation.Confusing replacing traces with tile coding.
Tile coding represents the continuous state; replacing traces are part of the Sarsa(λ) learning configuration.
Fix:
Assign representation to tile coding and learning configuration to replacing traces.Forgetting the reset between episodes.
The acrobot is returned to its initial resting position after the goal is reached.
Fix:
Trace each episode as resting start, repeated torque application, goal reached, and return to the initial resting position.
The Complete Design
- The acrobot swing-up task is online and model-free because the agent learns through interaction without access to the simulator itself.
- The direct state contains θ1, θ̇1, θ2, and θ̇2: the positions and velocities of the two links.
- Tile coding represents the continuous four-dimensional state space, while linear function approximation approximates the value function.
- Each discrete action has its own tiling set, allowing the controller to distinguish action-specific value estimates.
- Replacing traces are included in the Sarsa(λ) learning configuration, and each episode resets the acrobot after the goal is reached.
Key Takeaways
- Sarsa controls the acrobot through repeated online interaction rather than through a simulator model.
- The direct acrobot state has four continuous variables: two link positions and two link velocities.
- Tile coding and linear function approximation make the continuous state space usable without storing a separate exact value for every point.
- Separate tiling sets are used for the small discrete action set, while replacing traces form part of the Sarsa(λ) learning configuration.
- An episode starts with both links hanging down at rest, continues through torque applications until the goal, and then resets to the same initial position.