Environment Dynamics
The central question is what model of the environment is available to the agent.
The Information Question
Environment dynamics is best understood by asking one question: what model of the environment is available to the agent? Two reinforcement learning problems may involve similar agents and environments but still require different approaches because the agent may begin with different information about how the environment behaves.
The classification depends on what the agent already knows about the environment's behavior, not simply on what the environment is like.
Tracing a Model
A model of environment dynamics describes how states, allowable actions, transitions, and rewards are connected. In an MDP, the model contains two important kinds of information for all states and allowable actions: one-step transition probabilities and expected rewards.
A complete and accurate dynamics model covers the relevant states and allowable actions and supplies both transition-probability information and expected-reward information.
Complete Knowledge
A complete-knowledge problem provides the agent with a complete and accurate dynamics model of the environment.
The model must cover the relevant states and allowable actions. In an MDP, it must also contain one-step transition probabilities and expected rewards for all of those states and actions. Because the model is complete and accurate, it supplies the information needed to describe how each allowable action is connected to possible next states and expected rewards.
Classifying a Complete-Knowledge Description
A problem description says that the agent begins with a complete and accurate model covering the relevant states and allowable actions, including one-step transition probabilities and expected rewards. How should the problem be classified?
Check model availability: The description says that a model is available to the agent.
Check completeness: The model covers the relevant states and allowable actions.
Check accuracy and contents: The description identifies the model as accurate and includes transition probabilities and expected rewards.
Classify the problem: These conditions match the definition of a complete-knowledge problem.
The problem has complete knowledge.
Incomplete Knowledge
An incomplete-knowledge problem lacks a complete and perfect model of the environment's dynamics.
The key point is the absence of the complete and accurate model required for complete knowledge. A problem belongs to the incomplete-knowledge category when that complete and perfect model is unavailable to the agent.
Reading Problem Descriptions
To classify a reinforcement learning problem, inspect the information available to the agent at the beginning. Do not classify it only from the environment's apparent complexity or from the fact that the agent has some information. Look specifically for a complete and accurate model of the environment's dynamics.
| Description of the agent's initial information | Classification |
|---|---|
| A complete and accurate dynamics model is available. | Complete knowledge |
| A complete and perfect dynamics model is unavailable. | Incomplete knowledge |
| The description only says that the agent has some information. | Not enough information to establish complete knowledge |
The decisive question is whether a complete and accurate model is available.
Classifying an Incomplete-Knowledge Description
A problem description says that the agent and environment are similar to those in another problem, but the agent does not have a complete and perfect model of the environment's dynamics. How should the problem be classified?
Ignore superficial similarity: Similar agents and environments do not guarantee the same knowledge classification.
Inspect the initial model: The description says that a complete and perfect model is unavailable.
Classify the problem: The absence of a complete and perfect model places the problem in the incomplete-knowledge category.
The problem has incomplete knowledge.
Agent and Environment
The model can be viewed as the link between the agent's available knowledge and the environment's dynamics. For a state and allowable action, the model contains one-step transition probabilities and expected rewards. This is the information that distinguishes a complete model from an incomplete one.
Common Classification Mistakes
Classifying the problem from the environment alone.
The decisive issue is what model is available to the agent, not simply what the environment is like.
Fix:
Check whether the agent has a complete and accurate dynamics model.Treating any model as complete knowledge.
Complete knowledge requires a complete and accurate model.
Fix:
Check coverage of relevant states and allowable actions, along with transition probabilities and expected rewards.Ignoring accuracy.
The definition requires both completeness and accuracy.
Fix:
Classify the problem as complete-knowledge only when the model is complete and accurate.Assuming similar problems must have the same classification.
Different initial information can make the problems different reinforcement learning problems.
Fix:
Compare what each agent initially knows about the environment's behavior.
Classification Practice
A problem description says that the agent begins with a model containing expected rewards but does not provide a complete and accurate set of one-step transition probabilities for all relevant states and allowable actions. Is this complete knowledge or incomplete knowledge? Explain which requirement is missing.
Hints
- Complete knowledge requires more than one type of information.
- Check whether the model is complete and accurate for the relevant states and allowable actions.
- The source identifies both transition probabilities and expected rewards as model contents in an MDP.
Two reinforcement learning problems use similar environments. In the first, the agent has a complete and accurate dynamics model. In the second, that complete and perfect model is unavailable. Classify both problems and state the evidence for each classification.
Hints
- Focus on the model initially available to each agent.
- Do not classify the problems only from their environments.
- Compare each description with the definitions of complete and incomplete knowledge.
Key Takeaways
- Environment dynamics is analyzed by asking what model of the environment is available to the agent.
- Complete knowledge means that a complete and accurate dynamics model is available.
- In an MDP, the model includes one-step transition probabilities and expected rewards for all states and allowable actions.
- Incomplete knowledge means that a complete and perfect model is unavailable.
- To classify a problem, inspect the agent's initial information rather than judging only the environment or the agent's general similarity to another problem.
Key Takeaways
- The central question in environment dynamics is what model is available to the agent.
- A complete-knowledge problem provides a complete and accurate model covering relevant states and allowable actions.
- For an MDP, that model contains one-step transition probabilities and expected rewards.
- An incomplete-knowledge problem lacks a complete and perfect model.
- The classification is determined by the agent's initial information, not merely by the nature of the environment.