Concepts / On-Policy Control with Sarsa

On-Policy Control with Sarsa

Sarsa controls the acrobot through repeated online interaction rather than through access to a simulator model.

  • Programming

Learning Without a Simulator Model

An acrobot is a two-link robotic arm. In the swing-up task, a reinforcement learning agent applies torques until the goal is reached. The important feature is how the agent learns: it controls the simulated arm through repeated interaction without access to the simulator itself. This makes the task an online, model-free control problem, treated as though the agent were interacting with a physical acrobot.

Sarsa controls the acrobot by repeatedly observing the current situation, choosing a torque, applying it, and continuing the interaction until the goal is reached.

representevaluate and chooseapply torquecontinue learningnext situationCurrent stateθ1, θ̇1, θ2, θ̇2Tile featuresaction-specificrepresentationChosen torquediscrete actionAcrobot transitioninteraction continuesSarsa controllerlinear approximation andtraces
What cycle does the controller repeat while it learns to control the acrobot?

The Four-Part Acrobot State

The controller needs a description of the acrobot's current condition. The direct representation contains four continuous quantities: the position of the first link, the velocity of the first link, the position of the second link, and the velocity of the second link. They are written as θ1, θ̇1, θ2, and θ̇2.

QuantitySymbolWhat it describes
First-link positionθ1Position of the first link
First-link velocityθ̇1Velocity of the first link
Second-link positionθ2Position of the second link
Second-link velocityθ̇2Velocity of the second link

The four continuous quantities in the direct acrobot state

part of statepart of statepart of stateθ1first-link positionθ̇1first-link velocityθ2second-link positionθ̇2second-link velocity
What four variables are presented to the controller as the direct state?

From Continuous State to Action Value

The controller combines tile coding, linear function approximation, Sarsa(λ), and replacing traces. Tile coding supplies a representation for the continuous four-dimensional state space. Linear function approximation then uses that representation to approximate the value function rather than storing a separate exact value for every possible point in the region.

encodecombineapproximateContinuous stateθ1, θ̇1, θ2, θ̇2Active tilesstate representationLearned weightslinear approximationAction-value estimatevalue used by Sarsa
How does the continuous acrobot state become a value estimate used by the controller?

Representing One Controller Situation

Suppose the controller receives one current acrobot situation described by θ1, θ̇1, θ2, and θ̇2. How does that situation move through the learning configuration?

Describe: The current situation is described by the four direct state variables: the two link positions and the two link velocities.

Represent: Tile coding represents the point in the continuous four-dimensional state space.

Evaluate: Linear function approximation uses the tile representation to approximate the value associated with an action.

Control: The Sarsa controller uses the configured value representation while selecting and applying a torque.

The controller does not need a separate exact table entry for every possible point in the continuous state region.

Action Choice and Episode Reset

beginchooseapplycontinue until goalterminate and returnstart another episodeResting startlinks hanging straight downCurrent statefour direct variablesApplied torqueagent actionContinued interactionrepeated actions andtransitionsGoal reachedepisode terminationInitial restingpositionnext episode begins
How does an episode proceed from initialization through interaction to reset?

Every episode begins with both links hanging straight down and at rest. The learning agent then applies torques. Interaction continues until the goal is reached. At that point, the acrobot returns to the same initial resting position and another episode begins.

associateassociateassociateAcrobot statesame four variablesAction 1own tiling setAction 2own tiling setAction Nown tiling set
How are the same physical state features associated with different candidate actions?

The action set is small and discrete, so the setup uses a separate set of tilings for each action. This means the representation is not only a shared description of the physical state. Each available action has its own tiling set associated with evaluating that action in the state.

Traces in the Learning Configuration

Replacing traces are another component of the stated Sarsa(λ) configuration. Their role should be kept distinct from the role of tile coding. Tile coding represents the continuous state, action-specific tilings distinguish the discrete actions, and replacing traces belong to the learning configuration used alongside linear function approximation and Sarsa(λ).

record for learningincluded inActive featuresfrom tile codingFeature tracesreplacing-trace componentSarsa(λ)learning configuration
Where do replacing traces fit in the overall Sarsa(λ) setup?
ComponentPrimary role
Direct stateDescribes the current acrobot condition with four continuous quantities
Tile codingRepresents the continuous four-dimensional state space
Action-specific tilingsAssociates a separate tiling set with each discrete action
Linear function approximationApproximates the value function using the tile representation
Replacing tracesForms part of the Sarsa(λ) learning configuration

Check Your Understanding

EASY

Explain why the acrobot task is described as online and model-free. Then list the four direct state variables and state the role of action-specific tilings.

Hints
  • Online refers to learning through repeated interaction with the arm.
  • Model-free refers to controlling the arm without access to the simulator itself.
  • The direct state contains two positions and two velocities.
  • Each discrete action has its own tiling set.

What do you think happens?

After the acrobot reaches the goal, what happens before the next episode?

  • It remains at the goal indefinitely
  • It returns to the initial resting position
  • The four state variables are removed
  • The action set becomes continuous
Reveal answer

Answer: It returns to the initial resting position.

Each episode ends when the goal is reached. The acrobot is then returned to the same position in which the episode began: both links hanging straight down and at rest.

  • Treating the acrobot state as a small list of individually enumerated states.

    The direct state is continuous, and its possible values form a bounded region in four dimensions.

    Fix: Use tile coding to represent the continuous state space and linear function approximation to approximate values.

  • Assuming one shared tiling set is the complete action representation.

    The stated setup uses a separate set of tilings for each discrete action.

    Fix: Keep the physical state description distinct from the action-specific tiling sets used for action evaluation.

  • Confusing replacing traces with tile coding.

    Tile coding represents the continuous state; replacing traces are part of the Sarsa(λ) learning configuration.

    Fix: Assign representation to tile coding and learning configuration to replacing traces.

  • Forgetting the reset between episodes.

    The acrobot is returned to its initial resting position after the goal is reached.

    Fix: Trace each episode as resting start, repeated torque application, goal reached, and return to the initial resting position.

The Complete Design

  1. The acrobot swing-up task is online and model-free because the agent learns through interaction without access to the simulator itself.
  2. The direct state contains θ1, θ̇1, θ2, and θ̇2: the positions and velocities of the two links.
  3. Tile coding represents the continuous four-dimensional state space, while linear function approximation approximates the value function.
  4. Each discrete action has its own tiling set, allowing the controller to distinguish action-specific value estimates.
  5. Replacing traces are included in the Sarsa(λ) learning configuration, and each episode resets the acrobot after the goal is reached.

Key Takeaways

  • Sarsa controls the acrobot through repeated online interaction rather than through a simulator model.
  • The direct acrobot state has four continuous variables: two link positions and two link velocities.
  • Tile coding and linear function approximation make the continuous state space usable without storing a separate exact value for every point.
  • Separate tiling sets are used for the small discrete action set, while replacing traces form part of the Sarsa(λ) learning configuration.
  • An episode starts with both links hanging down at rest, continues through torque applications until the goal, and then resets to the same initial position.