Concepts / Policies and Sequential Decision Problems

Policies and Sequential Decision Problems

Artificial neural networks are function representations that can be used within reinforcement learning.

  • Programming

From Situations to Decisions

A reinforcement-learning agent must often respond to many possible situations. Storing a separate answer for every situation can be impractical, so the agent can use a function that receives information about a situation and produces an estimate or decision. An artificial neural network provides one way to represent such a function.

stateactionreward and next stateAgentEnvironment
How does information and control move from the current state to an action, reward, and next state over repeated decisions?

The loop shows why the problem is sequential. The agent receives information about a current situation, produces a decision, and then receives reinforcement-learning information connected to what happened next. A neural network can help produce the decision or estimate the usefulness of the situation and decision.

Neural Networks as Function Approximators

An artificial neural network is a function representation that can be used within reinforcement learning. Information describing a situation enters the network, passes through one or more layers, and produces an output.

information entersfunction outputSituationinformationinputArtificial neuralnetworkconnected computationalunitsDecision or estimateapproximation
How does a neural network transform many possible state inputs into an approximation of an otherwise difficult-to-store function?

The network does not need to store a separate answer for every possible situation. Instead, its connected computational units and weights represent a function that maps situation information to an output. Learning changes the weights so that the output becomes more useful for the task.

Replacing a Separate Answer Table with an Approximation

An agent encounters many possible situations in a sequential decision problem. How can a neural network help it respond without storing a separate answer for every situation?

Receive information: Information describing the current situation enters the artificial neural network.

Process information: The information passes through the network's connected computational units and layers.

Produce an output: The network produces either a decision rule output or an estimate associated with the situation.

Improve the representation: Learning changes the network's weights so that later outputs become more useful for the reinforcement-learning task.

The network serves as a function representation: it maps situation information to an approximate decision or estimate rather than requiring a separately stored answer for every situation.

Single-Layer and Multi-Layer Roles

The important structural distinction is whether the network has only one layer of neurons or several. A single-layer artificial neural network is the simpler arrangement. A multi-layer network adds intermediate layers, whose role can include learning a suitable representation before the final output is produced.

direct processingintermediate processingfinal processingInputInputOutputsimpler arrangementIntermediate layerslearn a representationOutputcomplex representation
What changes in the path from inputs to outputs when hidden layers are added, and what kinds of functions can each network represent?
StructureRole described in the sourceReinforcement-learning significance
Single-layer networkA simpler arrangement of neuronsProvides a basic network structure for representing a function
Multi-layer networkAdds intermediate layers that can learn a suitable representationIncreases the ability to learn complex representations and nonlinear policies

The historical progression moved from the abstract model neuron proposed by McCulloch and Pitts in 1943, through single-layer systems such as the Perceptron and ADALINE, to multi-layer networks trained with error backpropagation. The significance of adding layers is not merely that the network becomes larger. Intermediate layers make it possible to learn more complex representations and nonlinear control policies.

Forward Output and Backward Error

Training a multi-layer neural network involves comparing what the network produces with the output required for training, measuring the error, and using that error to adjust the network. Error backpropagation is the method associated with this process.

forwardforwardcompare with required outputbackward error informationguide adjustmentSituationinformationforward inputIntermediate layerforward computationNetwork outputproduced decision orestimateError informationcomparison resultNetwork weightsadjusted for learning
How does the output error travel backward through the layers to change the network's weights?

The defining direction is two-part. Information first travels forward through the layers to produce an output. Information about the error then travels backward through the network to guide learning. The purpose of this backward pass is to adjust the network's weights, including weights associated with intermediate layers, so that future outputs become more useful.

Tracing One Training Adjustment

A multi-layer network produces an output that differs from the output required for training. What happens next?

Forward pass: Situation information moves through the network and produces an output.

Error measurement: The produced output is compared with the output required for training, and the difference provides error information.

Backward pass: The error information is propagated backward through the network.

Weight adjustment: The backward information guides changes to the network's weights.

The next output can become more useful because the network has used error information to adjust its internal parameters.

Policies and Value Functions

In a sequential decision problem, a neural network can approximate more than one kind of function. A policy is a mapping used to choose behavior. A value function provides an estimate associated with situations or decisions. Both can be represented by artificial neural networks, while reinforcement-learning information is used to improve the approximations.

state enterschoose behaviorsituation entersestimateState informationinputState or decisioninformationinputPolicy networkfunction approximationValue networkfunction approximationAction choicebehaviorValue estimateevaluation
How are the inputs and outputs different when a network selects an action using a policy versus estimates future return using a value function?

The distinction is about the network's role in the agent's interaction with the task. A policy approximation is connected to behavior: it helps determine what the agent does. A value-function approximation is connected to evaluation: it estimates something associated with situations or decisions. These roles can be represented separately or used together.

The source discusses actor-critic systems to illustrate the two roles. In reported pole-balancing work, both an actor and a critic were represented by artificial neural networks. The actor concerned behavior, while the critic concerned evaluation. The example shows that neural networks can participate in more than one part of a reinforcement-learning method.

A Complete Decision Trace

Consider a generated abstract example. At one point in a sequential task, the agent receives information describing its current situation. A policy-approximating network processes that information and produces a behavior choice. The environment then supplies reinforcement-learning information connected to the result, including a reward and a next situation. A value-approximating network can provide an estimate associated with the situation or decision. The process repeats as the agent encounters the next situation.

inputinputchoose behaviorevaluateCurrent situationstate informationPolicy approximationbehavior functionValue approximationevaluation functionActionagent behaviorValue estimatesituation or decision
How can the same situation information support both action selection and evaluation in a sequential decision problem?

The policy and value roles should not be confused. The policy is associated with choosing behavior, whereas the value function is associated with estimating or evaluating situations or decisions.

What do you think happens?

An artificial neural network receives information about a situation and produces an estimate rather than an action choice. Which role does this output most directly represent?

  • A policy approximation
  • A value-function approximation
  • The error-backpropagation process
Reveal answer

Answer: A value-function approximation

The source distinguishes a policy as a mapping used to choose behavior from a value function as an estimate associated with situations or decisions. Backpropagation is the training process used to adjust a multi-layer network; it is not the output role.

Mistakes About Network Roles

  • Treating a neural network as a fixed lookup table with one separately stored answer for every situation.

    The source presents the network as a function representation that receives situation information and produces an estimate or decision.

    Fix: Describe the network as an approximation of a function whose weights are changed through learning.

  • Assuming that adding layers only increases the number of computational units without changing representation ability.

    Intermediate layers can learn a suitable representation, and the move to multi-layer networks increased the ability to learn complex representations and nonlinear policies.

    Fix: Explain the representational role of the intermediate layers.

  • Calling the forward production of an output backpropagation.

    The source distinguishes the forward movement of information from the backward propagation of error information.

    Fix: Use backpropagation for the backward error-guided training process that adjusts weights.

  • Using policy and value function as interchangeable terms.

    A policy is used to choose behavior, while a value function provides an estimate associated with situations or decisions.

    Fix: Identify whether the network output directs behavior or evaluates a situation or decision.

Check Your Understanding

MEDIUM

Explain, in your own words, why a neural network can be useful when a reinforcement-learning agent must respond to many possible situations. Then distinguish the roles of a policy approximation, a value-function approximation, and error backpropagation.

Hints
  • Start with the idea of representing a function instead of storing a separate answer for every situation.
  • For the policy, focus on behavior and action choice.
  • For the value function, focus on estimation or evaluation.
  • For backpropagation, focus on error information traveling backward to guide weight adjustment.
EASY

Create a short trace for a sequential decision problem using these terms in order: current situation, network output, action or estimate, error information, weight adjustment, next situation.

Hints
  • A policy output is connected to behavior.
  • A value output is connected to evaluation.
  • Error information is used during training rather than being the policy itself.

Key Takeaways

  1. Artificial neural networks can represent functions that map situation information to decisions or estimates in reinforcement learning.
  2. Single-layer networks are simpler arrangements, while multi-layer networks add intermediate layers that can learn more complex representations and nonlinear policies.
  3. Backpropagation sends error information backward through a multi-layer network to guide changes to its weights.
  4. A policy approximation is connected to choosing behavior, while a value-function approximation is connected to estimating situations or decisions.
  5. Actor-critic systems illustrate how separate neural-network roles can support behavior and evaluation within one reinforcement-learning method.

Key Takeaways

  • Neural networks provide function representations for handling many possible situations in reinforcement learning.
  • Adding intermediate layers enables learning of more complex representations and nonlinear control policies.
  • Backpropagation uses error information to adjust the weights of a multi-layer network.
  • Networks can approximate policies that guide behavior and value functions that evaluate situations or decisions.