Concepts / Tile Coding for Continuous State Spaces

Tile Coding for Continuous State Spaces

Sarsa controls the acrobot through repeated online interaction rather than through access to a simulator model.

  • Programming

Learning Without a Simulator Model

An acrobot is a two-link robotic arm. In the swing-up task, a reinforcement learning agent applies torques until the goal is reached. The distinctive feature is how the agent learns: it controls the simulated arm through interaction without access to the simulator itself. The controller therefore treats the task as online, model-free control rather than as a problem solved from a known transition model.

torquenext state and rewardobserveLearning agentAcrobotState and reward
How does the controller learn by repeatedly choosing actions and observing rewards and next states without using a transition model to predict the dynamics?

The interaction cycle is direct. The agent receives the current situation, chooses a torque, and observes what happens next. It does not need a separate model that predicts the simulator's transition before acting. This repeated interaction is the setting in which Sarsa controls the acrobot.

The Four-Part State

To choose actions, the controller needs a description of the acrobot's current condition. The direct state representation uses four continuous quantities: the position of the first link, the velocity of the first link, the position of the second link, and the velocity of the second link. They are written as θ1, θ̇1, θ2, and θ̇2.

containscontainscontainscontainsAcrobot stateθ1first-link positionθ̇1first-link velocityθ2second-link positionθ̇2second-link velocity
What are the four state variables, and how do the two joint angles and two angular velocities together describe the acrobot's current configuration and motion?

The direct state is a four-dimensional continuous description: two quantities describe link positions and two describe link velocities. It describes both the acrobot's configuration and its motion at the current point in the episode.

The two angles are constrained by the acrobot's physics, and the angles are naturally restricted to the range from 0 to 2π. Consequently, the complete state space is a bounded rectangular region in four dimensions. That is substantially different from a small table whose entries could list every state individually. The learning method therefore needs a function approximation approach.

VariableRole in the state description
θ1Position of the first link
θ̇1Velocity of the first link
θ2Position of the second link
θ̇2Velocity of the second link

The four continuous quantities in the direct acrobot state representation.

From State to Sarsa Update

The controller combines several pieces, each with a different job. Tile coding represents the continuous four-dimensional state. Linear function approximation uses that representation to approximate the value function instead of storing a separate exact value for every possible point in the region. Sarsa(λ) supplies the learning-control configuration, and replacing traces are included as part of that configuration.

representassociate with actionapproximateuse in learningselectsupport updateFour-dimensionalstateθ1, θ̇1, θ2, θ̇2Active tilesAction featuresAction-value estimateSarsa(λ)Next actionUpdated estimate
How does a continuous acrobot state become active tiles, how are those tile features used to estimate action values, and how does Sarsa update the estimate after the next action?

Reading one controller cycle

Trace what happens when the controller receives one current acrobot state during an episode.

Describe: The current situation is described by θ1, θ̇1, θ2, and θ̇2.

Represent: Tile coding represents that continuous state. Because the action set is small and discrete, the setup associates a separate set of tilings with each action.

Estimate: Linear function approximation uses the tile representation to approximate the value function for evaluating actions.

Learn: Sarsa(λ) is the learning-control method. Replacing traces are included in the stated learning configuration.

Act: The controller applies the chosen torque, and the interaction continues with the resulting acrobot situation.

The four continuous state variables are transformed into a representation that supports action evaluation and Sarsa learning without requiring a separately stored exact value for every point in the four-dimensional state region.

The Episode Loop

Every acrobot episode follows a repeating pattern. The arm begins with both links hanging straight down and at rest. The learning agent then applies torques. Interaction continues until the goal is reached. Once the goal is reached, the acrobot is returned to the same initial resting position, and another episode begins.

begincurrent situationchosen torquenew interaction resultcheck progressnot reachedreachednext episodeInitial restingstatelinks hanging downObserve statefour variablesSelect actionApply torqueSarsa learningGoal reachedEpisode resetlinks hanging down
What happens next at each step of an acrobot episode, and how does the interaction return to a reset state after the goal is reached?

Tracing an episode

Follow the episode structure from its initial condition through repeated interaction and back to a new episode.

Start: Both links are hanging straight down and the acrobot is at rest.

Observe: The controller describes the current condition with the four direct state variables.

Interact: The agent applies torques and continues interacting with the acrobot.

Continue: The control-and-learning cycle continues while the goal has not yet been reached.

Reset: After the goal is reached, the acrobot returns to the same initial resting position.

The reset is not the end of the overall learning process. It is the starting condition for another episode.

Action Tilings and Replacing Traces

Two parts of the setup are easy to confuse because both belong to the learning system but perform different jobs. The small action set is discrete, so the controller uses a separate set of tilings for each action. This means the representation associated with evaluating one action is distinguished from the representation associated with evaluating another action.

Replacing traces are different. They are identified by the source as a component of the Sarsa(λ) learning configuration, alongside linear function approximation and tile coding. In the division of responsibilities, action-specific tilings distinguish action representations, while replacing traces belong to the learning configuration used by Sarsa(λ).

representform part ofAction-specifictilingsone set per actionReplacing tracesSarsa(λ) componentAction evaluationLearningconfiguration
What is the difference between storing separate tile features for each action and including replacing traces in the Sarsa learning configuration?
ComponentPrimary roleWhat it should not be confused with
Separate tiling set for each actionDistinguishes the state representation used for evaluating each discrete actionIt is not the complete learning algorithm by itself
Replacing tracesForms part of the Sarsa(λ) learning configurationIt is not the action-specific tiling arrangement

Mistakes in Reading the Setup

  • Treating the acrobot as if the controller starts with a simulator model.

    The task is described as online, model-free control. The agent learns through interaction without access to the simulator itself.

    Fix: Describe the controller as choosing actions and observing the resulting interaction.

  • Leaving out the velocities from the state.

    The direct representation contains four quantities: two positions and two velocities.

    Fix: Include θ1, θ̇1, θ2, and θ̇2.

  • Using one shared tiling arrangement without noting the action distinction.

    The setup uses a separate set of tilings for each member of the small discrete action set.

    Fix: State that each action has its own associated tiling set.

  • Describing tile coding as the whole learning method.

    Tile coding represents the continuous state, while linear function approximation and Sarsa(λ) handle value approximation and learning.

    Fix: Explain the representation, approximation, and learning roles separately.

  • Stopping the process permanently when an episode reaches the goal.

    After the goal is reached, the acrobot is returned to its initial resting position and another episode begins.

    Fix: Describe goal achievement as the transition to reset and a new episode.

Check Your Understanding

MEDIUM

Explain the complete role of each item in this chain: four direct state variables, tile coding, action-specific tilings, linear function approximation, Sarsa(λ), and replacing traces. Then describe what happens after the acrobot reaches the goal.

Hints
  • Start with what describes the acrobot's current condition.
  • Separate state representation from value approximation and learning.
  • Remember that a separate tiling set is associated with each discrete action.
  • End by describing the reset to the initial resting position.

What do you think happens?

After the acrobot reaches the goal, does the overall interaction stop permanently, or does another episode begin?

  • The interaction stops permanently
  • The acrobot is reset and another episode begins
Reveal answer

Answer: The acrobot is reset and another episode begins.

Each episode ends with the acrobot being returned to the same initial resting position, after which the repeating episode pattern starts again.

Key Takeaways

  1. The acrobot swing-up task is online and model-free: the agent learns by interacting with the acrobot rather than using access to the simulator's model.
  2. The direct state contains four continuous quantities: θ1, θ̇1, θ2, and θ̇2.
  3. Tile coding represents the continuous four-dimensional state, while linear function approximation approximates values without storing an exact value for every possible state point.
  4. Because the action set is small and discrete, each action has its own associated tiling set.
  5. Sarsa(λ) and replacing traces belong to the learning configuration, and each episode returns to the initial resting state after the goal is reached.

Key Takeaways

  • The acrobot controller learns online through repeated interaction and does not rely on access to a simulator model.
  • Its direct state representation has two link positions and two link velocities.
  • Tile coding converts the continuous state into a useful representation for linear function approximation.
  • Separate tiling sets distinguish the discrete actions, while replacing traces are part of the Sarsa(λ) learning configuration.
  • An episode begins with both links hanging down at rest, continues through torque application until the goal is reached, and then resets to begin another episode.