Concepts / Policies in Markov Decision Processes

Policies in Markov Decision Processes

A state-value function evaluates a state under a policy.

  • Programming

One State, Two Evaluations

In a Markov decision process, a value function describes how valuable a situation or decision is under a policy. The key question is whether we are evaluating only the current state or evaluating one particular action chosen in that state. That single difference changes what the value function is describing.

start atthen followevaluateState sstarting stateAction aselected firstPolicy πfollowed afterwardExpected return
What is held fixed when evaluating qπ(s, a), and what happens after the selected action?

State Value Under a Policy

A state-value function evaluates a state under a policy. In the expression for state value, the focus is the state s and the policy π. The function asks how valuable that state is when considered under that policy. It does not single out one particular initial action in the description.

considerguidescontributes toState sevaluate herePolicy πunder this policyFuture actionsselected by πState valuevalue of s under π
Starting from a state, how do the policy and later outcomes determine the state’s value?

Reading a state-value description

A learner says, “The value of state s under policy π evaluates one particular action taken in s.” Is this an accurate description of state value?

Identify the object being evaluated: The object is state s, not a named action.

Identify the policy: The state is evaluated under policy π.

Check for an initial action: No specific initial action is included in the state-value description.

The description is not accurate. State value evaluates state s under policy π; evaluating a specific initial action requires the action-value function.

Action Value After the First Move

The action-value function qπ(s, a) evaluates the expected return from starting at state s, taking action a, and then following policy π. Its description fixes the starting state, fixes the action taken first, and specifies that the policy is followed afterward.

The initial action and the later policy have different roles. Action a is selected explicitly at the beginning of the description. After that initial action, policy π governs what is followed afterward. Therefore, qπ(s, a) is more specific than a state-value description: it evaluates one particular action in one particular state while keeping the later policy fixed.

takethen reachπ selectscontinue towardState sfixed startLater stateExpected returnAction afixed first actionPolicy actionselected by π
How does control move from a fixed initial action to actions selected by the policy in later states?

Reading qπ(s, a)

Interpret the statement: qπ(s, a) evaluates the expected return from starting at s, taking a, and thereafter following π.

Fix the starting state: The evaluation begins at state s.

Fix the first action: The action taken first is specifically a.

Apply the policy afterward: After the selected first action, the policy π is followed.

Identify the evaluated result: The function evaluates the expected return for that complete description.

qπ(s, a) evaluates one state-action starting situation under a specified policy for what follows.

State Versus Action Evaluation

QuestionState-value evaluationAction-value evaluation
What is evaluated?A stateA particular action taken in a state
Starting pointState sState s
Is an initial action named?No specific initial action is includedAction a is named and taken first
What policy is involved?State s is evaluated under policy πAfter action a, policy π is followed
Main distinctionFocuses on the state under the policyFocuses on one action in the state, then the policy afterward
evaluatesunderstarts attakesthen followsState valuestate s under πState sevaluatedPolicy πunderlying policyAction valueqπ(s, a)State sstarting stateAction ataken firstPolicy πfollowed afterward
How does evaluating a state under a policy differ from evaluating one specific action chosen in that state?

When answering a question about value functions, state the object being evaluated before discussing the policy. For state value, say that state s is evaluated under policy π. For action value, add the missing specificity: start at s, take a, and thereafter follow π.

Mistakes in Reading qπ(s, a)

  • Treating qπ(s, a) as if it evaluated only state s.

    The expression also fixes action a as the action taken first.

    Fix: Describe qπ(s, a) as the expected return from starting at s, taking a, and then following π.

  • Forgetting the policy after the first action.

    The action-value description includes what happens afterward: policy π is followed.

    Fix: Mention both the initial action and the policy followed afterward.

  • Assuming the initial action is selected by the policy in the same way as every later action.

    The action a is explicitly fixed in the action-value description.

    Fix: Keep action a fixed as the first action, then identify π as the policy followed afterward.

  • Leaving out the starting state.

    The description begins from a particular state s.

    Fix: Name all three parts: starting state s, initial action a, and later policy π.

Check Your Interpretation

EASY

Explain the difference between these two descriptions: evaluating state s under policy π, and evaluating qπ(s, a). Your answer should identify what is evaluated, what is fixed at the start, and what happens after the first action.

Hints
  • State value does not name one specific initial action.
  • Action value names action a as the first action.
  • The policy π is followed after that initial action.

A complete answer

Give a complete explanation of qπ(s, a).

Name the start: Begin from state s.

Name the selected action: Take the specific action a first.

Name what follows: After that action, follow policy π.

Name the evaluation: The function evaluates the expected return for this starting state, first action, and later policy.

qπ(s, a) evaluates the expected return from starting at s, taking a, and thereafter following π.

Key Takeaways

  1. A state-value function evaluates a state under a policy.
  2. The action-value function qπ(s, a) evaluates the expected return from starting at s, taking a, and then following π.
  3. State-value evaluation focuses on a state; action-value evaluation focuses on a specific initial action in that state.
  4. A complete action-value description fixes the starting state, the first action, and the policy followed afterward.
  5. The essential distinction is whether a specific initial action is included.

Key Takeaways

  • State value evaluates state s under policy π.
  • qπ(s, a) evaluates the expected return from starting at s, taking action a, and thereafter following π.
  • Action value is more specific because it names the initial action.
  • To interpret qπ(s, a), track the starting state, the selected first action, and the later policy.