Concepts / Policy Evaluation in Reinforcement Learning

Policy Evaluation in Reinforcement Learning

The optimal state-value function v∗(s) gives the maximum expected return achievable from a state.

  • Programming

From Situations to Decisions

Reinforcement learning agents must evaluate both situations and decisions. A state describes where the agent is, while an action describes what the agent does from that state. The optimal value functions organize these two perspectives: v*(s) evaluates a state, and q*(s, a) evaluates a particular action in a particular state.

The central distinction is the information being conditioned on. v*(s) is indexed by a state. q*(s, a) is indexed by both a state and an action.

evaluatesconditions onconditions onv*(s)state-level expected returnq*(s, a)state-action expectedreturnstate swhere the agent isaction awhat the agent does
What information does v*(s) contain about a state, and what additional information does q*(s, a) contain about each possible action from that state?

The Optimal State Value

The optimal state-value function v*(s) gives the maximum expected return achievable from state s. It asks: how much return can be achieved from this state when the agent follows an optimal policy?

The word maximum is important. A state may be evaluated under different policies, and those policies can produce different expected returns from that same starting state. v*(s) identifies the largest expected return achievable from the state. It therefore describes the best state-level outcome, rather than the outcome of one arbitrarily selected policy.

evaluateevaluateevaluateselect largest returnstate ssame starting statepolicy Aexpected returnv*(s)maximum expected returnpolicy Bexpected returnoptimal policylargest achievable return
How do the expected returns of different policies from the same state compare, and how is the largest one selected as v*(s)?

Reading v*(s)

Interpret the question answered by v*(s) for a state s.

Identify the input: The input is the state s, which describes where the agent is.

Consider policies: Consider the expected returns available from that state under possible policies.

Select the optimal result: The optimal state-value function focuses on the maximum expected return achievable from the state.

v*(s) is a state-level evaluation: it describes the maximum expected return achievable from state s when an optimal policy is followed.

The Action-Value Perspective

The optimal action-value function q*(s, a) evaluates a particular action a in a particular state s, followed by an optimal policy. It asks: what return is expected if the agent first takes this particular action and then follows an optimal policy?

q*(s, a) contains more specific information about the initial decision than v*(s) does. v*(s) describes the best expected return associated with the state as a whole. q*(s, a) keeps the selected initial action visible, so it evaluates a state-action pair rather than only the starting state.

FunctionIndexed byQuestion answered
v*(s)State sWhat maximum expected return is achievable from this state?
q*(s, a)State s and action aWhat expected return follows from taking this action first and then following an optimal policy?

Both functions evaluate expected return, but q*(s, a) also identifies the initial action being evaluated.

takeleads tothen followevaluate expected returnstate sstarting situationaction acommitted initial actionresulting situationafter the initial actionoptimal policysubsequent decisionsq*(s, a)expected return
What happens after an agent takes action a in state s, and how does it then follow optimal actions to accumulate return?

Connecting the Two Views

A useful mental sequence is to identify the state, identify a candidate action, evaluate that state-action pair with q*(s, a), and then relate the action-specific evaluation to the optimal state-level evaluation v*(s). The functions answer related questions, but they do not preserve the same amount of decision detail.

Comparing the Questions

An agent is in state s and is considering action a. Decide whether each question is about v*(s) or q*(s, a).

Question about the state: A question asking for the maximum expected return achievable from s without naming an initial action is about v*(s).

Question about the decision: A question asking what return follows after the agent first takes action a is about q*(s, a).

Follow-up behavior: For q*(s, a), the initial action is fixed for the evaluation, and the subsequent behavior follows an optimal policy.

v*(s) evaluates the state-level opportunity; q*(s, a) evaluates one initial decision within that opportunity.

Why Exact Tables Become Impractical

Reinforcement-learning environments may contain too many states for exact tabular storage. An agent may therefore be unable to store one exact entry for every state. The limitation is not only what the agent knows about the environment; it is also what the agent can store and compute.

Memory is needed for state information as well as for approximations of value functions, policies, and models. Computation can also be limiting: even with a complete and accurate environment model, an agent might not have enough time to perform all necessary computations at every time step.

createmotivatemotivatemotivateexact tableone entry for each statemany statestoo many for exact storagememory andcomputation limitsstorage and processingconstraintsapproximated valuefunctionestimated state informationapproximated policyestimated decisionsapproximated modelestimated environmentinformation
How does an agent represent or estimate information about many states when it cannot store exact values for every state?

When discussing a reinforcement-learning solution, separate the ideal question from the representation available to the agent. First ask what the optimal value or policy would be. Then ask whether memory and computation permit the agent to represent or calculate it exactly.

Ideal Versus Achievable Performance

Optimality is useful as a theoretical target even when an agent cannot reach an optimal solution exactly. The optimal value functions describe the best expected returns associated with states and state-action choices. They provide a reference point for understanding what an agent is trying to approximate.

In practice, limited memory and computation can prevent exact storage or complete calculation. The resulting value function, policy, or model may therefore be an approximation of the optimal one. Aiming for optimality and achieving an approximation are different claims: the first describes the objective, while the second describes the result that resource limits may permit.

target fortarget forinfluencesinfluencesoptimal valuetheoretical targetoptimal policybest expected returnapproximate valueresource-limited estimateapproximate policyresource-limited resultmemory andcomputationlimits exact achievement
What is the difference between the ideal optimal value or policy and the approximate result an agent can actually compute?

Common Interpretation Mistakes

  • Treating v*(s) as an evaluation of one named action

    v*(s) is indexed only by the state and describes the maximum expected return achievable from that state.

    Fix: Use q*(s, a) when the initial action is part of the question.

  • Treating q*(s, a) as a state-only value

    q*(s, a) evaluates a particular action in a particular state.

    Fix: Keep both the state and the initial action in the interpretation.

  • Assuming optimality means exact implementation is always possible

    Environments may contain too many states for exact tabular storage, and computation may also be limited.

    Fix: Treat optimality as a theoretical target and distinguish it from an approximate value, policy, or model.

  • Ignoring the follow-up policy in q*(s, a)

    The optimal action-value function evaluates the initial action followed by an optimal policy.

    Fix: Include both stages: commit to the initial action, then follow an optimal policy.

Apply the Distinction

EASY

For each question, decide whether it refers to v*(s) or q*(s, a): (1) the maximum expected return achievable from a state, (2) the expected return after first taking a named action and then following an optimal policy, and (3) an evaluation that must preserve which action was selected.

Hints
  • Look at whether the question is indexed by a state only or by a state and an action.
  • The phrase first takes a particular action points to the action-value perspective.
  • The phrase maximum expected return from a state points to the state-value perspective.
MEDIUM

Explain in your own words why an agent might use an approximation even when the optimal value function or optimal policy is the desired target.

Hints
  • Consider how many states the environment may contain.
  • Consider the memory needed for state information, value functions, policies, and models.
  • Consider the computation required even when a complete and accurate model is available.

Key Takeaways

  1. v*(s) gives the maximum expected return achievable from state s under an optimal policy.
  2. q*(s, a) evaluates a particular initial action in state s, followed by an optimal policy.
  3. v*(s) provides a state-level view, while q*(s, a) preserves action-specific information.
  4. Large state spaces, memory limits, and computation limits make exact storage and calculation impractical in some environments.
  5. Optimality is a theoretical target; an agent may need to use approximate value functions, policies, or models.

Key Takeaways

  • v*(s) is the maximum expected return achievable from a state.
  • q*(s, a) evaluates a specific initial action followed by an optimal policy.
  • The state-value function answers a state-level question; the action-value function answers a state-and-action question.
  • Memory and computation limits often require approximate representations of value functions, policies, and models.
  • Optimality remains a useful theoretical ideal even when an agent can achieve only an approximation.