Linear Function Approximation in Reinforcement Learning
Sarsa controls the acrobot through repeated online interaction rather than through access to a simulator model.
The Control Challenge
The acrobot swing-up task is a useful setting for studying function approximation in reinforcement learning. An acrobot is a two-link robotic arm. The agent applies torques until the swing-up goal is reached, but it does not receive a simulator model that tells it in advance how the arm will respond. Instead, it learns through repeated interaction. This makes the task an online, model-free control problem, treated as though the agent were controlling a physical acrobot.
The phrase model-free matters here: the learning agent controls the simulated arm without access to the simulator itself. It must use interaction rather than a complete advance description of the environment.
The Four-Part State
To choose an action, the controller needs a description of the acrobot's current condition. The direct state representation contains four continuous quantities: the position of the first link, the velocity of the first link, the position of the second link, and the velocity of the second link. These are written as θ1, θ̇1, θ2, and θ̇2.
Reading one direct state
Suppose the controller receives one current acrobot state represented by θ1, θ̇1, θ2, and θ̇2. What information does that state provide?
Step 1: Read θ1 and θ2 as the positions of the first and second links.
Step 2: Read θ̇1 and θ̇2 as the velocities of the first and second links.
Step 3: Treat all four quantities together as one description of the acrobot's current condition, rather than as four separate states.
The controller's input is one four-variable continuous state describing both link positions and both link velocities.
From Continuous State to Action Values
Because the state space is continuous and four-dimensional, the controller does not store a separate exact value for every possible point. Tile coding represents the continuous state space, and linear function approximation uses that representation to approximate the value function. Together, they allow the controller to represent values across the state region instead of treating every possible continuous point as an unrelated table entry.
The small discrete action set changes how tile coding is organized. The setup uses a separate set of tilings for each action. Thus, the state is not represented only by one shared tiling arrangement. Each available action has its own tiling set associated with evaluating that action.
Tile coding answers the representation question: how should a continuous state be described using features? Linear function approximation answers the value question: how should those features represent an estimated value? Sarsa uses the resulting estimates while controlling the acrobot.
Action Tilings and Traces
The stated learning configuration combines Sarsa(λ), linear function approximation, tile coding, and replacing traces. These components have different responsibilities. The acrobot supplies the continuous four-variable state. Tile coding represents that state. Action-specific tilings distinguish the discrete actions. Linear approximation represents the value function used by Sarsa. Replacing traces are included as part of the learning configuration to mark and update recently active features over time.
| Component | Role in the setup |
|---|---|
| Direct state representation | Describes the acrobot using two positions and two velocities |
| Tile coding | Represents the continuous state space with features |
| Action-specific tilings | Associates a separate tiling set with each discrete action |
| Linear function approximation | Represents an approximation of the value function |
| Sarsa(λ) | Controls the acrobot using the estimated values |
| Replacing traces | Forms part of the learning configuration for recently active features |
Function Approximation from Evidence
In reinforcement learning, function approximation uses available examples from a desired function to construct a broader representation of that function. The desired function may be a value function. The examples are the evidence available to the learner; the approximation is the generalized representation built from that evidence.
The desired function, the examples, and the approximation are not the same thing. The desired function is the target relationship. The examples are the observations available to the learner. The approximation is the representation constructed from those observations. Its purpose is not merely to store the examples, but to generalize from them and represent the desired function more broadly.
Function approximation is an instance of supervised learning. Supervised learning is also studied in machine learning, artificial neural networks, pattern recognition, and statistical curve fitting. This connection allows reinforcement learning to combine its own learning methods with generalization methods developed in those related fields.
Separating target from approximation
A reinforcement learning system has some examples related to a value function but not a complete description of that function. What is being learned?
Identify the target: The desired function is the value function the learner would like to represent.
Identify the evidence: The available examples are the information the learner has received about that desired function.
Identify the constructed object: The function approximator builds a broader representation from the examples instead of simply memorizing each example separately.
Connect it to reinforcement learning: The resulting approximation can take the role of a value-function approximator inside a reinforcement learning algorithm.
Function approximation fills the gap between incomplete examples and a broader representation of the desired value function.
Common Conceptual Mistakes
Treating the acrobot task as model-based control
The setup describes the agent as controlling the simulated arm without access to the simulator itself.
Fix:
Describe the task as online and model-free: the agent learns through repeated interaction.Calling the four state variables four separate states
The four quantities together describe one current acrobot condition.
Fix:
Treat θ1, θ̇1, θ2, and θ̇2 as the components of one continuous state.Confusing tile coding with the value function
Tile coding supplies a representation of the continuous state space, while linear function approximation represents the value function using that representation.
Fix:
Keep representation and value estimation distinct.Assuming one shared tiling set is used for every action
The described setup uses a separate set of tilings for each action.
Fix:
Associate each available action with its own tiling set.Equating examples with the desired function
Examples are evidence about the desired function, whereas the approximation is the broader representation constructed from that evidence.
Fix:
Distinguish the target function, the observed examples, and the generalized approximation.
Check Your Understanding
Explain the complete learning setup in your own words. Start with the four quantities in the acrobot state. Then describe why the continuous state motivates function approximation, how tile coding represents it, why each action has its own tiling set, and where linear approximation, Sarsa(λ), and replacing traces fit.
Hints
- Name both positions and both velocities.
- Separate state representation from value-function approximation.
- Mention that the agent learns through interaction without access to the simulator model.
- Explain that the episode returns to the initial resting position after the goal is reached.
What do you think happens?
Before reading the reveal, decide whether the following statement is correct: the available examples, the desired function, and the approximation are three names for the same object.
Reveal answer
Answer: Incorrect
The desired function is the target, the examples are the available evidence, and the approximation is the broader representation constructed from that evidence.
Key Takeaways
- The acrobot swing-up task is online and model-free because the agent learns through interaction without access to the simulator model.
- The direct acrobot state contains θ1, θ̇1, θ2, and θ̇2: two link positions and two link velocities.
- Tile coding represents the continuous four-dimensional state space, while linear function approximation represents the value function used by Sarsa.
- Each discrete action has its own tiling set, and replacing traces are part of the Sarsa(λ) learning configuration.
- Function approximation generalizes from available examples of a desired function, and this supervised-learning perspective allows reinforcement learning to use generalization methods from related fields.
Key Takeaways
- The acrobot is controlled through repeated online interaction rather than through a simulator model.
- Its direct state is represented by two positions and two velocities.
- Tile coding and action-specific tilings represent the continuous state for discrete actions, while linear approximation estimates values for Sarsa.
- Replacing traces are part of the stated Sarsa(λ) learning configuration.
- Function approximation uses examples to construct a broader representation of a desired function and connects reinforcement learning with supervised learning.