Concepts / Action Selection Methods

Action Selection Methods

Semi-gradient Sarsa with function approximation was first explored by Rummery and Niranjan in 1994.

  • Programming

Why Action Choice Matters

The history of Semi-gradient Sarsa with function approximation is also a history of asking how action selection affects learning behavior. The method was first explored by Rummery and Niranjan in 1994. Later research examined whether particular combinations of function approximation and action selection would converge, and what convergence should mean in each setting.

Do not treat convergence as an unconditional property of Semi-gradient Sarsa. The reported results depend on the action selection setting being discussed.

From First Exploration to Later Studies

The historical sequence begins in 1994, when Rummery and Niranjan first explored Semi-gradient Sarsa with function approximation. Later, Gordon reported a specific convergence concern for linear Semi-gradient Sarsa using ε-greedy action selection. Still later, Precup and Perkins showed convergence in a differentiable action selection setting. These stages should be read as a progression of research questions rather than as one result that applies to every version of the method.

later studystill later1994Rummery and NiranjanGordonlinear and ε-greedyPrecup and Perkinsdifferentiable actionselection
How did research progress from the first exploration of function-approximation-based Semi-gradient Sarsa to later convergence studies?
led to later questionsRummery and Niranjan1994Convergence researchlater studies
When was Semi-gradient Sarsa with function approximation first explored, and where does that event appear in the later research timeline?

The Two Convergence Results

The ε-greedy result is deliberately qualified. For linear Semi-gradient Sarsa with ε-greedy action selection, Gordon reported that the method did not converge in the usual sense. However, it did enter a bounded region near the best solution. This is different from saying that learning becomes unbounded or that the method has no useful behavior.

The later differentiable-action-selection result is also conditional. Precup and Perkins showed convergence in a differentiable action selection setting. The source does not turn this result into a claim that every action selection rule produces convergence, nor does it provide enough detail here to generalize the result beyond the stated setting.

reported behaviorreported resultε-greedylinear Semi-gradient SarsaDifferentiableaction selectionBounded regionnear the best solutionConvergencereported by Precup andPerkins
What is the difference between the reported result for ε-greedy action selection and the result for a differentiable action selection setting?
SettingReported resultHow to interpret it
Linear Semi-gradient Sarsa with ε-greedy action selectionDoes not converge in the usual sense; enters a bounded region near the best solutionA bounded-region result is not the same as ordinary convergence
Differentiable action selectionConvergence was shownThe result is tied to the differentiable action selection setting

The two results are conditional on their stated action selection settings.

Connecting the Four Ideas

Four ideas must be kept together when reading this research history: Semi-gradient Sarsa is the learning method, function approximation describes how the method represents what it learns, action selection specifies the setting in which actions are chosen, and convergence is the behavior being studied. Changing the action selection setting changes which convergence statement can be supported.

studied withcombined withaffects reportedSemi-gradient Sarsalearning methodFunctionapproximationmethod settingAction selectionε-greedy or differentiableConvergenceresearch result
How are the algorithm, function-approximation choice, action-selection method, and convergence result connected?

Reading a conditional claim

Suppose a research note says that one action selection setting gives a bounded-region result and another setting gives convergence. What should a careful reader conclude?

Identify the method: Recognize that both statements concern Semi-gradient Sarsa with function approximation.

Identify the action selection setting: Keep the ε-greedy and differentiable settings separate instead of merging them into one general case.

Match each outcome: Associate the bounded-region result with linear Semi-gradient Sarsa using ε-greedy action selection, and associate the convergence result with the differentiable action selection setting.

Avoid overgeneralizing: Do not conclude that Semi-gradient Sarsa always converges or never converges. The evidence supports conditional claims.

The correct interpretation is setting-specific: the reported outcome depends on the action selection rule and the other conditions named in the result.

Mistakes in Reading the History

  • Saying that linear Semi-gradient Sarsa with ε-greedy action selection converges.

    The reported result says that it does not converge in the usual sense, although it enters a bounded region near the best solution.

    Fix: Describe the bounded-region result and distinguish it from ordinary convergence.

  • Saying that Semi-gradient Sarsa never converges.

    Convergence was later shown in a differentiable action selection setting.

    Fix: State the action selection setting before making a convergence claim.

  • Treating the differentiable-action-selection result as applying to ε-greedy action selection.

    The source reports the results as belonging to different settings.

    Fix: Keep the two settings and their reported outcomes separate.

  • Using an application reference to supply missing convergence details.

    The source states that convergence claims here come from the explicitly reported action selection settings, not from an unstated interpretation of an application reference.

    Fix: Base the claim on the named action selection setting and reported result.

When summarizing a convergence result, name the method, the function-approximation condition, the action selection setting, and the exact strength of the reported conclusion. This prevents a conditional result from becoming an unconditional claim.

Check Your Interpretation

MEDIUM

Write a two-sentence comparison of the ε-greedy and differentiable action selection results. Your first sentence should state the reported outcome for linear Semi-gradient Sarsa with ε-greedy action selection. Your second sentence should state the later result for a differentiable action selection setting without extending it to other settings.

Hints
  • Use the phrase bounded region near the best solution for the ε-greedy result.
  • Use the phrase convergence was shown for the differentiable setting.
  • Mention that the claims are conditional on their action selection settings.

What do you think happens?

A result says that a method enters a bounded region near the best solution but does not converge in the usual sense. Should you summarize that as convergence?

  • Yes, because bounded behavior and convergence mean the same thing.
  • No, because the source distinguishes the bounded-region result from convergence in the usual sense.
Reveal answer

Answer: No, because the source distinguishes the bounded-region result from convergence in the usual sense.

Precise research reading requires preserving the distinction between the reported bounded-region behavior and the later convergence result in a differentiable action selection setting.

Historical Takeaway

The research history moves from initial exploration to increasingly specific convergence questions. Rummery and Niranjan first explored Semi-gradient Sarsa with function approximation in 1994. Gordon later reported that linear Semi-gradient Sarsa with ε-greedy action selection does not converge in the usual sense, while entering a bounded region near the best solution. Precup and Perkins later showed convergence in a differentiable action selection setting. The main lesson is to attach every convergence claim to its stated action selection conditions.

Key Takeaways

  • Semi-gradient Sarsa with function approximation was first explored by Rummery and Niranjan in 1994.
  • Linear Semi-gradient Sarsa with ε-greedy action selection was reported not to converge in the usual sense, although it entered a bounded region near the best solution.
  • Convergence was later shown in a differentiable action selection setting.
  • The evidence supports conditional, setting-specific convergence claims rather than an unconditional claim that the method always converges or never converges.