Concepts / Model-Free Reinforcement Learning

Model-Free Reinforcement Learning

A model gives an agent a way to infer how the environment may respond.

  • Programming

Before the Environment Responds

An agent does not always have to choose an action and then wait for the real environment to respond before considering what might happen. Some reinforcement learning systems include a model of the environment. The model supports inferences about how the environment may behave, allowing the agent to consider possible consequences before those situations are actually experienced.

The central question is whether the agent has a model that can infer what may happen next, or whether it learns directly from trial-and-error interaction without such a model.

The Model Query

The basic model query pairs a current state with an action. The model receives that state-action pair and produces a prediction of the next state and the next reward. In symbols from the source example, an agent in state S1 considers action A1; the model predicts that the environment may move to state S2 and provide reward R1.

paired withpaired withpredictspredictsState S1Environment modelinferenceState S2predicted next stateAction A1Reward R1predicted next reward
How do a state and action pass through an environment model to produce a predicted next state and reward?

Reading a Model Query

An agent is currently in state S1 and is considering action A1. What does the model provide?

Form the query: The model receives the pair consisting of state S1 and action A1.

Read the prediction: The model predicts that the environment may move to state S2 and provide reward R1.

Interpret the result: The prediction describes a possible consequence. The real environment has not been moved by this query.

The model maps the state-action pair S1 and A1 to a predicted next state S2 and predicted next reward R1.

Planning Through Possibilities

Planning means considering possible future situations before they are actually experienced. A model makes this possible by supplying predictions about what may follow a state and action. The agent can use those possible future situations when deciding on a course of action, rather than relying only on a response that has already occurred in the real environment.

considerconsidermodel predictsmodel predictsCurrent stateS1Action A1Possible outcomeS2 and R1Action A2Possible outcomeanother predicted state andreward
How does an agent explore possible future states and use their predicted rewards to choose an action?

The tree represents alternatives, not a record of events that have already happened. Each branch begins with a possible action and leads to a possible future situation supplied by the model. Planning uses these possibilities to reason about which course of action to take.

What do you think happens?

An agent asks a model what may happen after taking A1 in state S1. Has the real environment already moved to the predicted next state?

  • Yes, because the model produced a next state
  • No, because the model only supplied an inference
  • Only if the predicted reward is positive
Reveal answer

Answer: No, because the model only supplied an inference.

The source distinguishes a model prediction from a real environment transition. The model has not caused the environment to move.

Two Learning Families

AspectModel-based methodsModel-free methods
Environment modelUse a model of the environmentDo not use a model of the environment
Decision approachUse planning and possible future situationsLearn through trial and error
Relationship to interactionCan consider consequences before those situations are experiencedLearn through interaction rather than first using a model to reason about possible futures
useslearns throughModel-basedmodel and planningPossible futuresconsidered beforeexperienceModel-freetrial and errorInteractionexperience supplieslearning
What is the difference between an agent that predicts future outcomes with a model and one that learns from trial-and-error experience?

Model-free does not mean that the agent cannot learn. It means that the agent is described as learning through trial and error without a model of the environment. The contrast is with model-based methods, whose decisions involve a model and planning.

Trial-and-Error Experience

In a model-free method, the agent does not first use an environment model to reason about possible futures. Instead, it learns through interaction with the environment: it takes actions and learns from the resulting experience. This is why model-free methods are described as almost the opposite of planning.

selectsinteracts withprovidesbecomes part ofinformsreturns to action selectionStateActionEnvironmentRewardExperiencetrial and errorFuture action choice
How does experience flow from taking an action and receiving a reward into future action choices without an environment model?

Common Classification Mistakes

  • Treating a model prediction as a real environment transition.

    The source states that the model has not caused the real environment to move. It has only supplied an inference about what could happen.

    Fix: Describe S2 as a predicted next state until the real environment actually responds.

  • Defining model-free learning as learning without rewards.

    The model-free distinction concerns the absence of an environment model and the use of trial-and-error learning, not the absence of rewards.

    Fix: Remember that the source's model query includes a predicted next reward, while model-free methods are characterized by learning through interaction without a model.

  • Calling any action selection planning.

    Planning specifically means considering possible future situations before they are actually experienced.

    Fix: Use planning for the model-based process of considering possible futures, and use trial and error for learning through interaction without first using a model.

Check Your Understanding

EASY

An agent is in state S1 and considers action A1. A component predicts a possible next state S2 and reward R1, but the real environment has not yet responded. Identify the component's role and classify the decision process as model-based or model-free.

Hints
  • Look for the component that receives a state-action pair and produces a predicted next state and next reward.
  • Ask whether the agent is considering a possible consequence before the situation is actually experienced.

Practice Solution

Classify the process in which a component predicts S2 and R1 from S1 and A1 before the real environment responds.

Identify the component: A component that receives a state and action and predicts a next state and reward is an environment model.

Identify the reasoning pattern: The agent is considering a possible future situation before it is experienced, which is planning.

Classify the method: Using a model and planning identifies the process as model-based rather than model-free.

This is a model-based process: the model predicts a possible outcome, and the agent can use that prediction for planning.

Key Takeaways

  1. An environment model gives an agent a way to infer how the environment may respond.
  2. A basic model query pairs a state and action with a predicted next state and next reward.
  3. A model prediction is an inference; it does not itself cause the real environment to move.
  4. Planning considers possible future situations before they are actually experienced.
  5. Model-based methods use models and planning, while model-free methods learn through trial and error without a model.

Key Takeaways

  • A model predicts how the environment may respond to a state-action pair.
  • The predicted result includes a possible next state and next reward.
  • Planning uses possible future situations before they are experienced.
  • Model-free reinforcement learning learns through trial and error without an environment model.
  • The presence or absence of a model separates model-based and model-free methods.