REINFORCE Policy Gradient Algorithm
Exploration is necessary because a team needs varied collective actions in order to learn.
Why Collective Learning Needs Variation
A team can learn about its collective behavior only if it produces more than one collective action. If every member always produces exactly the same action, the team repeatedly visits the same combination and has no variation with which to explore other collective behaviors. Exploration is therefore not merely an individual concern. It is a requirement for learning at the team level.
The central strategy is to make exploration local to each agent. Each agent varies its own output independently. Those local variations combine, producing different action combinations for the team as a whole. The team can then receive different common rewards for different collective actions, giving its learning process something to compare.
Tracing a Team Action
Three Local Choices, One Collective Action
Consider three agents whose outputs are binary. What is the team-level action when their outputs are 1, 0, and 1?
Separate outputs: Agent A produces 1, Agent B produces 0, and Agent C produces 1. Each output is an individual action.
Combine outputs: The team action is the combination of all three outputs: 1, 0, 1.
Compare another trial: If the agents later produce 0, 0, and 1, the team has produced a different collective action even though each agent still makes a binary choice.
Connect to learning: Different collective actions allow the team to explore its collective action space. The common reward can then guide learning toward behavior associated with a higher average reward rate.
Independent variation at the unit level becomes variation in the team's combined actions.
Bernoulli-Logistic Output
A Bernoulli-logistic unit produces a binary output, but its output remains probabilistic. Its weighted input sum biases the probability of firing; it does not turn the output into a permanently fixed choice. With the same input condition, the unit can therefore continue to produce variable outputs across trials.
This persistent variability is what makes the unit useful for exploration. The unit is not required to choose a different output on every trial. Instead, its output remains governed by a probability, so repeated trials can provide opportunities for different choices. When several such units operate independently, their output variation creates variation in the team's combined action.
What do you think happens?
If a Bernoulli-logistic unit receives the same weighted input on several trials, must it produce exactly the same binary output every time?
Reveal answer
Answer: No, because the output remains probabilistic.
The weighted input sum biases the probability of firing, but the Bernoulli-logistic output remains variable. That variability supports exploration.
From Local Learners to Team Policy
Each Bernoulli-logistic unit can use a REINFORCE policy gradient algorithm. At the unit level, its weight adjustment aims to maximize the average reward rate experienced by that unit while it stochastically explores its action space. When the units operate as a team, their outputs define collective actions and the team supplies a common reward signal.
The same system can therefore be viewed in two ways. Locally, it is a collection of REINFORCE learners, each varying its own output and adjusting its weights in relation to reward. Globally, it is a policy-gradient system whose actions are collective combinations and whose objective is the average rate of the team's common reward.
Keep exploration and learning conceptually separate. Exploration comes from the agents' variable outputs. Learning comes from changing the agents' weights toward behavior associated with a higher average reward rate. Confusing these roles makes the team-level mechanism harder to understand.
Mistakes About Team Exploration
Assuming that one fixed collective action is enough for team learning.
The team cannot learn about collective behavior without variation across collective actions.
Fix:
Allow team members to vary their outputs so that the team explores different combinations.Treating independent exploration as unrelated individual noise.
The combined outputs define the team's collective action, so local variability can create team-level action diversity.
Fix:
Trace both levels: first the individual outputs, then the collective combination they produce.Assuming that a weighted input fixes a Bernoulli-logistic output.
The weighted input biases firing probability, but the output remains probabilistic.
Fix:
Understand the weighted input as influencing probability while preserving output variability.Describing the system only as separate learners.
At the team level, the outputs form collective actions and the common reward guides the team-level policy gradient.
Fix:
Use both perspectives: local REINFORCE units and one global policy over collective actions.
| Perspective | Action | Reward context | Learning interpretation |
|---|---|---|---|
| Individual unit | The unit's binary output | Reward experienced by that unit | A local REINFORCE learner explores its action space |
| Whole team | The combination of all unit outputs | The team's common reward | A policy-gradient system explores collective actions |
Check Your Understanding
A team contains several Bernoulli-logistic REINFORCE units. Explain why independent variability at the unit level can help the team explore collective actions. Then distinguish what is being optimized from the viewpoint of one unit and from the viewpoint of the whole team.
Hints
- Start with the fact that the team action is the combination of its members' outputs.
- Explain what happens when at least one member varies its output.
- For the final distinction, identify the unit's action, the team's action, and the role of the common reward.
A strong answer should state that local stochastic outputs create different combinations of team actions, and that the common reward guides learning toward a higher average reward rate for the team-level policy over those collective actions.
Key Takeaways
- Collective learning requires variation in the team's combined actions.
- Independent variability by individual agents creates diversity in collective actions.
- A Bernoulli-logistic unit remains probabilistic even when its weighted input biases its firing probability.
- Local REINFORCE units can operate together as a team-level policy-gradient system.
- At the team level, the action is the combination of member outputs and the learning signal is the team's common reward.