Concepts / Semi-gradient prediction methods

Semi-gradient prediction methods

Action-value prediction changes the training input from a state to a state-action pair.

  • Programming

The Input Changes

Semi-gradient prediction can be extended from predicting a value for a state to predicting an action value for a state-action pair. The central change is the training input. A state-value example uses a state alone, while an action-value example uses the state together with the action taken in that state.

predicts frompredicts fromState Ststate-only inputState valueprediction for StState St, action Atstate-action inputAction valueprediction for St, At
How does the training input change when prediction moves from a state value to an action value?

Action-Value Inputs

A state can be associated with an estimated value, but action-value prediction must distinguish among actions available in that state. For this reason, the method receives both St and At as its example input. The action is not an optional label added after prediction; it is part of the input used to estimate the action value.

Changing One Training Example

Represent the training input for state-only prediction and then represent the corresponding input for action-value prediction.

State-only example: The example is represented by the state St. The predictor is asked to estimate a value associated with that state.

Action-value example: The example is represented by the pair St, At. The predictor is asked to estimate the value associated with taking At in St.

What changed: The state remains part of the example, but the action is now included in the input because the prediction concerns an action value.

The extension changes the example input from St to the pair St, At.

supplies example inputparameterizesState St, action Attraining inputApproximate actionvalueq-hatWeight vector thetaparameters
How do a state-action pair and the weight vector combine to produce the estimated action value?

The Approximate Predictor

The approximate action-value function, written as q-hat, is represented in parameterized form with a weight vector theta. The state-action pair is supplied to this parameterized predictor, which produces an estimated action value.

The weight vector is the adjustable part of the predictor. The general gradient-descent update uses the state-action example and its target to adjust this parameterized action-value predictor. The supplied source identifies this update structure but does not provide its full equation, so the essential information to track is the example, the target, the predictor, and the weight vector being adjusted.

What Ut Represents

In action-value prediction, Ut is the target associated with the training example St, At. It is a backed-up approximation of the action value at that state-action pair. The target may be the full Monte Carlo return, an n-step Sarsa return, or another approximation of q-pi at the state-action pair.

is associated withprovidesState St, action Atobserved exampleBacked-upapproximationreturn or otherapproximationTarget Utdesired prediction for thepair
What information does Ut represent, and how does it connect the observed state-action pair to the desired prediction?
Target optionRole in the training example
Full Monte Carlo returnA possible target Ut for the state-action pair
n-step Sarsa returnA possible target Ut for the state-action pair
Another approximation of q-piA possible target Ut for the state-action pair

Following the Update

A useful way to follow the general update is to track the information it receives. First, identify the state St and the action At that form the training input. Next, identify Ut, the target associated with that pair. The parameterized approximate action-value predictor uses this example and target in a gradient-descent update that adjusts its weight vector theta.

Tracing One Action-Value Example

A training example contains a state St, an action At, and a target Ut. Trace the role of each item without supplying a numerical update.

Identify the input: Combine St and At into the state-action example used by the action-value predictor.

Identify the target: Treat Ut as the backed-up approximation of the action value for that state-action pair.

Apply the general update: Use the state-action example and Ut to adjust the parameterized approximate action-value predictor.

Track the adjustable object: The adjustment is made to the predictor's weight vector theta.

The update is organized around the pair St, At, its target Ut, the approximate action-value function, and its weight vector theta.

example inputtarget informationupdate adjustsState St, action Attraining exampleAction-valuepredictorparameterized q-hatWeight vector thetaadjusted parametersTarget Utbacked-up approximation
How does the action-value training information flow into the parameter adjustment?

Common Misreadings

  • Using only St as the input for an action-value example.

    Action-value prediction must distinguish the action taken in the state, so the training input is the pair St, At.

    Fix: Record both the state and the action whenever the prediction concerns an action value.

  • Treating Ut as the state-action input.

    The pair is the training input, while Ut is the target associated with that pair.

    Fix: Keep the roles separate: St, At is the example input and Ut is its backed-up target.

  • Assuming Ut must always be one specific kind of return.

    The source identifies the full Monte Carlo return, an n-step Sarsa return, and another approximation of q-pi as possible targets.

    Fix: Recognize Ut as a target that may take one of the source-identified forms.

  • Confusing the weight vector with the training example.

    Theta represents the weights of the parameterized approximate action-value function.

    Fix: Identify St, At as the input and theta as the adjustable weight vector.

  • Writing a full update equation that is not supplied by the source.

    The supplied source names the general update structure but does not include its full equation.

    Fix: Explain the information flow: state-action example and target are used to adjust the parameterized predictor and its weights.

Check Your Understanding

EASY

A learner says: An action-value training example is St mapped to Ut. Correct the statement by naming the complete input and explaining the role of Ut.

Hints
  • Ask what must be added when prediction moves from a state value to an action value.
  • Separate the example input from the target.
  • Mention the parameterized approximate action-value predictor and its weight vector.

What do you think happens?

If a prediction method changes from state values to action values, what is the main change in one training example?

  • The input changes from St to St, At
  • The target must always become a full Monte Carlo return
  • The weight vector becomes the training input
  • The state is removed from the example
Reveal answer

Answer: The input changes from St to St, At.

Action-value prediction uses the state together with the action taken in that state. The target remains a backed-up approximation of the action value, and the parameterized predictor continues to use its weight vector.

paired withpaired withState StinputTargetvalue targetState St, action AtinputTarget Utaction-value target
What does each training example contain in state-value prediction versus action-value prediction?

Key Takeaways

  • Action-value prediction extends state-based prediction by changing the input from St to the state-action pair St, At.
  • The approximate action-value function q-hat is parameterized by a weight vector theta.
  • Ut is the target associated with St, At and may be a full Monte Carlo return, an n-step Sarsa return, or another approximation of q-pi.
  • The general gradient-descent update uses the state-action example and its target to adjust the parameterized action-value predictor.
  • The state-action pair, the target, the predictor, and the weight vector have separate roles and should not be conflated.