REINFORCE eligibility vector
A Bernoulli-logistic unit is a stochastic neuron-like unit for binary actions.
From State Features to a Binary Action
A Bernoulli-logistic unit is a stochastic neuron-like unit for choosing between two actions. Its action is either 0 or 1. Instead of storing a separate arbitrary action for every possible state, the unit uses the state's feature vector to produce a probability. That probability determines how likely the unit is to sample action 1; the probability of action 0 is whatever remains.
Preference Difference and Logistic Probability
Let h(s, 0, θ) and h(s, 1, θ) be the preferences for actions 0 and 1. The model does not need to specify both preferences independently. Their difference is defined by the weighted feature sum θᵀφ(s). The logistic function converts this difference into the policy probability for action 1.
π(1 | s, θ) = 1 / (1 + exp(−θᵀφ(s)))The dot product θᵀφ(s) is the preference difference before the logistic transformation. A larger preference difference produces a larger probability for action 1 through the logistic expression. Once that probability is known, the two-action policy is determined: action 1 has probability π(1 | s, θ), and action 0 has the remaining probability.
Reading the Eligibility Vector
For a sampled action A_t in state S_t, the REINFORCE eligibility vector is the gradient of the log policy probability with respect to the parameter vector: e_t = ∇θ log π(A_t | S_t, θ_t). It describes how the log probability of the action that was actually sampled changes when the parameters change.
e_t = [A_t − π(1 | S_t, θ_t)] φ(S_t)
The eligibility vector is not another action and is not the return. It is a vector describing the local sensitivity of the selected action's log probability to the parameters. Each feature contributes a component because each feature participates in the preference difference through the dot product.
A Symbolic REINFORCE Trace
From a sampled action to eligibility
Suppose the current state has feature vector φ(S_t), the policy probability for action 1 is p, and the sampled action is A_t = 1. Determine the eligibility vector in terms of p and φ(S_t).
Identify the sampled action: The sampled action is A_t = 1.
Use the action factor: Substitute A_t = 1 into A_t − π(1 | S_t, θ_t). Since π(1 | S_t, θ_t) is p, the factor becomes 1 − p.
Multiply by the features: The eligibility vector is e_t = (1 − p)φ(S_t). Every feature component is scaled by the same action-dependent factor.
When action 1 is sampled, the eligibility vector is (1 − p)φ(S_t).
The other sampled action
Using the same state features and probability p, suppose the sampled action is A_t = 0. Determine the eligibility vector.
Identify the sampled action: The sampled action is A_t = 0.
Use the action factor: Substitute A_t = 0 into A_t − π(1 | S_t, θ_t). The factor becomes −p.
Multiply by the features: The eligibility vector is e_t = −pφ(S_t). The sign differs from the action-1 case because the sampled action was 0.
When action 0 is sampled, the eligibility vector is −pφ(S_t).
These two cases show why the sampled action matters. The same feature vector can produce different eligibility vectors depending on whether the unit sampled 0 or 1. The current policy probability also matters, because it scales the action-dependent factor.
What Monte-Carlo REINFORCE Changes
In Monte-Carlo REINFORCE, the agent receives a return G_t after the sampled action. The parameter vector changes from θ_t to θ_t+1. The eligibility vector supplies direction-related information for that change, while the return supplies the outcome associated with the sampled action.
The eligibility vector does not by itself specify the complete size of the update. It identifies how the selected log policy probability responds to the parameters. In Monte-Carlo REINFORCE, that direction-related information is combined with the return when the parameter vector is changed.
Common Interpretation Errors
Treating the unit as if it stores one fixed action for each state.
The Bernoulli-logistic unit turns the state's feature vector into a probability and then produces a stochastic binary action.
Fix:
Track the probability of action 1 and remember that the unit samples either action 0 or action 1.Using θᵀφ(s) as the probability itself.
The dot product is the preference difference. The logistic function converts that preference difference into π(1 | s, θ).
Fix:
Separate the two stages: compute the preference difference, then apply the logistic function.Ignoring which action was sampled when computing eligibility.
The eligibility quantity uses the sampled action, the policy probability, and the state features.
Fix:
Use e_t = [A_t − π(1 | S_t, θ_t)]φ(S_t).Confusing the eligibility vector with the return.
The return is received after the sampled action, whereas the eligibility vector describes the log-policy sensitivity to the parameters.
Fix:
Keep their roles separate: the eligibility vector supplies direction-related information, and the return is part of the Monte-Carlo learning signal.Assuming the update leaves the policy representation unchanged.
Monte-Carlo REINFORCE changes the parameter vector from θ_t to θ_t+1.
Fix:
Use the updated parameter vector in later evaluations of the logistic expression.
Check Your Understanding
A Bernoulli-logistic unit has state feature vector φ(S_t), parameter vector θ_t, and sampled action A_t = 0. Write the preference difference, the probability of action 1, and the eligibility vector. Then state which quantity changes from θ_t to θ_t+1 during the Monte-Carlo REINFORCE update.
Hints
- The preference difference is the dot product θ_tᵀφ(S_t).
- Apply the logistic function to obtain π(1 | S_t, θ_t).
- Substitute A_t = 0 into the eligibility expression.
- The parameter vector is the quantity explicitly changed by the update.
Key Takeaways
- A Bernoulli-logistic unit represents a binary action by assigning a probability to action 1 and using the remaining probability for action 0.
- The parameter vector and state feature vector form the preference difference θᵀφ(s), which the logistic function converts into π(1 | s, θ).
- The REINFORCE eligibility vector is the gradient of the log probability of the sampled action with respect to the parameters.
- For this unit, the eligibility vector is [A_t − π(1 | S_t, θ_t)]φ(S_t), so every feature contributes a component.
- Monte-Carlo REINFORCE changes the policy parameter vector from θ_t to θ_t+1 using the return and eligibility information.
Key Takeaways
- A Bernoulli-logistic unit turns state features into a probability distribution over binary actions.
- The preference difference θᵀφ(s) becomes the probability of action 1 through the logistic function.
- The eligibility vector records the parameter sensitivity of the sampled action's log policy probability.
- The sampled action and current policy probability determine the action-dependent factor in the eligibility vector.
- Monte-Carlo REINFORCE updates the parameter vector using the return and the eligibility direction.