Tile Coding for Continuous State Spaces
Sarsa controls the acrobot through repeated online interaction rather than through access to a simulator model.
Learning Without a Simulator Model
An acrobot is a two-link robotic arm. In the swing-up task, a reinforcement learning agent applies torques until the goal is reached. The distinctive feature is how the agent learns: it controls the simulated arm through interaction without access to the simulator itself. The controller therefore treats the task as online, model-free control rather than as a problem solved from a known transition model.
The interaction cycle is direct. The agent receives the current situation, chooses a torque, and observes what happens next. It does not need a separate model that predicts the simulator's transition before acting. This repeated interaction is the setting in which Sarsa controls the acrobot.
The Four-Part State
To choose actions, the controller needs a description of the acrobot's current condition. The direct state representation uses four continuous quantities: the position of the first link, the velocity of the first link, the position of the second link, and the velocity of the second link. They are written as θ1, θ̇1, θ2, and θ̇2.
The direct state is a four-dimensional continuous description: two quantities describe link positions and two describe link velocities. It describes both the acrobot's configuration and its motion at the current point in the episode.
The two angles are constrained by the acrobot's physics, and the angles are naturally restricted to the range from 0 to 2π. Consequently, the complete state space is a bounded rectangular region in four dimensions. That is substantially different from a small table whose entries could list every state individually. The learning method therefore needs a function approximation approach.
| Variable | Role in the state description |
|---|---|
| θ1 | Position of the first link |
| θ̇1 | Velocity of the first link |
| θ2 | Position of the second link |
| θ̇2 | Velocity of the second link |
The four continuous quantities in the direct acrobot state representation.
From State to Sarsa Update
The controller combines several pieces, each with a different job. Tile coding represents the continuous four-dimensional state. Linear function approximation uses that representation to approximate the value function instead of storing a separate exact value for every possible point in the region. Sarsa(λ) supplies the learning-control configuration, and replacing traces are included as part of that configuration.
Reading one controller cycle
Trace what happens when the controller receives one current acrobot state during an episode.
Describe: The current situation is described by θ1, θ̇1, θ2, and θ̇2.
Represent: Tile coding represents that continuous state. Because the action set is small and discrete, the setup associates a separate set of tilings with each action.
Estimate: Linear function approximation uses the tile representation to approximate the value function for evaluating actions.
Learn: Sarsa(λ) is the learning-control method. Replacing traces are included in the stated learning configuration.
Act: The controller applies the chosen torque, and the interaction continues with the resulting acrobot situation.
The four continuous state variables are transformed into a representation that supports action evaluation and Sarsa learning without requiring a separately stored exact value for every point in the four-dimensional state region.
The Episode Loop
Every acrobot episode follows a repeating pattern. The arm begins with both links hanging straight down and at rest. The learning agent then applies torques. Interaction continues until the goal is reached. Once the goal is reached, the acrobot is returned to the same initial resting position, and another episode begins.
Tracing an episode
Follow the episode structure from its initial condition through repeated interaction and back to a new episode.
Start: Both links are hanging straight down and the acrobot is at rest.
Observe: The controller describes the current condition with the four direct state variables.
Interact: The agent applies torques and continues interacting with the acrobot.
Continue: The control-and-learning cycle continues while the goal has not yet been reached.
Reset: After the goal is reached, the acrobot returns to the same initial resting position.
The reset is not the end of the overall learning process. It is the starting condition for another episode.
Action Tilings and Replacing Traces
Two parts of the setup are easy to confuse because both belong to the learning system but perform different jobs. The small action set is discrete, so the controller uses a separate set of tilings for each action. This means the representation associated with evaluating one action is distinguished from the representation associated with evaluating another action.
Replacing traces are different. They are identified by the source as a component of the Sarsa(λ) learning configuration, alongside linear function approximation and tile coding. In the division of responsibilities, action-specific tilings distinguish action representations, while replacing traces belong to the learning configuration used by Sarsa(λ).
| Component | Primary role | What it should not be confused with |
|---|---|---|
| Separate tiling set for each action | Distinguishes the state representation used for evaluating each discrete action | It is not the complete learning algorithm by itself |
| Replacing traces | Forms part of the Sarsa(λ) learning configuration | It is not the action-specific tiling arrangement |
Mistakes in Reading the Setup
Treating the acrobot as if the controller starts with a simulator model.
The task is described as online, model-free control. The agent learns through interaction without access to the simulator itself.
Fix:
Describe the controller as choosing actions and observing the resulting interaction.Leaving out the velocities from the state.
The direct representation contains four quantities: two positions and two velocities.
Fix:
Include θ1, θ̇1, θ2, and θ̇2.Using one shared tiling arrangement without noting the action distinction.
The setup uses a separate set of tilings for each member of the small discrete action set.
Fix:
State that each action has its own associated tiling set.Describing tile coding as the whole learning method.
Tile coding represents the continuous state, while linear function approximation and Sarsa(λ) handle value approximation and learning.
Fix:
Explain the representation, approximation, and learning roles separately.Stopping the process permanently when an episode reaches the goal.
After the goal is reached, the acrobot is returned to its initial resting position and another episode begins.
Fix:
Describe goal achievement as the transition to reset and a new episode.
Check Your Understanding
Explain the complete role of each item in this chain: four direct state variables, tile coding, action-specific tilings, linear function approximation, Sarsa(λ), and replacing traces. Then describe what happens after the acrobot reaches the goal.
Hints
- Start with what describes the acrobot's current condition.
- Separate state representation from value approximation and learning.
- Remember that a separate tiling set is associated with each discrete action.
- End by describing the reset to the initial resting position.
What do you think happens?
After the acrobot reaches the goal, does the overall interaction stop permanently, or does another episode begin?
Reveal answer
Answer: The acrobot is reset and another episode begins.
Each episode ends with the acrobot being returned to the same initial resting position, after which the repeating episode pattern starts again.
Key Takeaways
- The acrobot swing-up task is online and model-free: the agent learns by interacting with the acrobot rather than using access to the simulator's model.
- The direct state contains four continuous quantities: θ1, θ̇1, θ2, and θ̇2.
- Tile coding represents the continuous four-dimensional state, while linear function approximation approximates values without storing an exact value for every possible state point.
- Because the action set is small and discrete, each action has its own associated tiling set.
- Sarsa(λ) and replacing traces belong to the learning configuration, and each episode returns to the initial resting state after the goal is reached.
Key Takeaways
- The acrobot controller learns online through repeated interaction and does not rely on access to a simulator model.
- Its direct state representation has two link positions and two link velocities.
- Tile coding converts the continuous state into a useful representation for linear function approximation.
- Separate tiling sets distinguish the discrete actions, while replacing traces are part of the Sarsa(λ) learning configuration.
- An episode begins with both links hanging down at rest, continues through torque application until the goal is reached, and then resets to begin another episode.