Concepts / Reinforcement Learning Problems

Reinforcement Learning Problems

The central question is what model of the environment is available to the agent.

  • Programming

The Agent's Starting Knowledge

Two reinforcement learning problems may involve similar agents and environments but still require different approaches. The decisive question is not only what the environment is like. It is what the agent already knows about how that environment behaves.

The central classification question is: what model of the environment is available to the agent?

availableunavailableRL problemComplete modelComplete knowledgeNo complete modelIncomplete knowledge
How does the information available to the agent determine whether the problem has complete or incomplete knowledge?

Reading an Environment Dynamics Model

A model of an environment's dynamics describes what can happen after the agent takes an allowable action in a state. In an MDP, the model contains two kinds of information for all states and allowable actions: one-step transition probabilities and expected rewards.

The model therefore connects a current state and an allowable action with information about the next state and the reward expected from that one-step interaction. To count as complete knowledge, this information must cover the relevant states and allowable actions and must be complete and accurate.

inputinputprovidesprovidesStateDynamics modelTransitionprobabilitiesOne-stepAllowable actionExpected rewardsOne-step
Given a state and an allowable action, what does the environment dynamics model tell the agent about the next state and reward?

Complete Knowledge

A complete-knowledge problem provides the agent with a complete and accurate dynamics model of the environment.

For an MDP, completeness means that the model covers all states and allowable actions relevant to the problem. It supplies both one-step transition probabilities and expected rewards. Accuracy matters as well: a model that is available but not correct does not satisfy the requirement for complete knowledge.

coverscoverscontainscontainsmust beAvailable modelAll relevant statesAll allowable actionsTransitionprobabilitiesOne-stepExpected rewardsOne-stepAccurate model
What information is available to the agent when its knowledge is complete?

Classifying a Supplied Model

An agent begins with a model that covers every relevant state and allowable action. For each one-step interaction, the model provides transition probabilities and expected rewards, and the model is accurate. What kind of reinforcement learning problem is this?

Check coverage: The model covers the relevant states and allowable actions.

Check information: The model includes one-step transition probabilities and expected rewards.

Check accuracy: The model is described as accurate, so the available information correctly represents the environment's dynamics.

Classify: Because the complete and accurate dynamics model is available from the beginning, this is a complete-knowledge problem.

Complete knowledge

Incomplete Knowledge

An incomplete-knowledge problem lacks a complete and perfect model of the environment's dynamics.

The model may be unavailable as a complete source of information about the relevant states, allowable actions, transition probabilities, and expected rewards. The key point is the absence of a complete and perfect model, not simply the fact that the agent knows little in a general sense.

may lackmay lackmay lackmay lackIncomplete modelRelevant statesNot fully coveredAllowable actionsNot fully coveredTransitionprobabilitiesIncomplete or unavailableExpected rewardsIncomplete or unavailable
Which parts of the environment's dynamics model are unknown or unavailable in an incomplete-knowledge problem?

Finding the Missing Model

An agent begins in an environment where the complete and perfect dynamics model is not available. The agent does not initially have complete transition-probability and expected-reward information for all relevant states and allowable actions. How should the problem be classified?

Inspect the initial information: The agent does not begin with a complete and perfect model.

Compare with the dividing line: Complete knowledge requires a complete and accurate dynamics model.

Classify: Because that complete model is unavailable, the problem has incomplete knowledge.

Incomplete knowledge

Classifying by Initial Information

To classify a problem, inspect what the agent knows at the start. Do not classify it only from the environment's description. The same kind of environment can belong to different categories if one agent begins with a complete and accurate model while another does not.

used byused byclassified asclassified asEnvironmentSimilar settingAgent AComplete model initiallyavailableComplete knowledgeAgent BNo complete model initiallyavailableIncomplete knowledge
How can two otherwise similar reinforcement learning problems receive different classifications?
EASY

A problem description says that an agent begins with a model covering all relevant states and allowable actions, including one-step transition probabilities and expected rewards. The model is accurate. Classify the problem and name the evidence that supports your answer.

Hints
  • Check whether the model is complete.
  • Check whether the model is accurate.
  • Look for both transition-probability and expected-reward information.
MEDIUM

A second problem uses a similar environment, but the agent begins without a complete and perfect model of the environment's dynamics. Classify the problem. Explain why the environment alone is not enough to determine the classification.

Hints
  • The dividing line concerns the model available to the agent.
  • Compare the agent's initial information with the requirements for complete knowledge.
  • Treating any model as complete knowledge.

    Complete knowledge requires a complete and accurate model covering the relevant states and allowable actions and supplying transition-probability and expected-reward information.

    Fix: Check completeness and accuracy rather than merely checking whether some model information exists.

  • Classifying the problem from the environment alone.

    The classification depends on what the agent initially knows about the environment's behavior.

    Fix: Inspect the model available to the agent at the beginning of the problem.

  • Ignoring expected rewards.

    In an MDP, the model contains both one-step transition probabilities and expected rewards for all states and allowable actions.

    Fix: Verify that both kinds of information are included.

  • Ignoring model accuracy.

    Complete knowledge requires a complete and accurate dynamics model.

    Fix: Treat accuracy as a required part of the classification.

Classification Checklist

  1. Identify the model of the environment's dynamics available to the agent at the start.
  2. Check whether the model covers the relevant states and allowable actions.
  3. Check whether it provides one-step transition probabilities and expected rewards.
  4. Check whether the model is complete and accurate.
  5. Classify the problem as complete knowledge only when the complete and accurate model is available; otherwise classify it as incomplete knowledge.
  1. Complete knowledge means that the agent begins with a complete and accurate dynamics model. In an MDP, that model includes one-step transition probabilities and expected rewards for all states and allowable actions. Incomplete knowledge means that a complete and perfect model is unavailable. The classification depends on the agent's initial information, not merely on the environment itself.

Key Takeaways

  • The central question in a reinforcement learning problem is what model of the environment is available to the agent.
  • Complete knowledge requires a complete and accurate dynamics model.
  • In an MDP, the model includes one-step transition probabilities and expected rewards for all states and allowable actions.
  • Incomplete knowledge means that a complete and perfect model is unavailable.
  • The classification is determined by the agent's initial information, so similar environments can produce different problem classifications.