Ldim
SOA begins with V_1 = H and maintains the hypotheses consistent with observed labels.
Why SOA Keeps a Version Space
The Standard Optimal Algorithm, abbreviated SOA, makes each prediction by examining the hypotheses that are still consistent with the labels observed so far. It begins with V_1 = H, meaning that the initial consistent class is the full hypothesis class H. As examples arrive, hypotheses that disagree with the true labels are removed. The remaining class is the current version space used for the next prediction.
SOA does not update its version space using the label it predicted. It updates using the true label supplied after the prediction.
One Round of SOA
At round t, SOA receives an input x_t and starts with the current class V_t. It divides V_t into two subclasses. V_t^(0) contains the hypotheses that assign label 0 to x_t, while V_t^(1) contains the hypotheses that assign label 1 to x_t. SOA compares the Ldim of these two subclasses and predicts the label belonging to the subclass with the larger Ldim. If the two Ldim values are equal, SOA predicts 1.
- Initialize the process with V_1 = H.
- At round t, divide V_t into V_t^(0) and V_t^(1) according to the labels assigned to x_t.
- Compare the Ldim of the two candidate subclasses.
- Predict the label whose subclass has larger Ldim; predict 1 if the values are equal.
- After the true label y_t arrives, keep only the hypotheses in V_t that agree with y_t to form V_(t+1).
Comparing Candidate Labels
Choosing Between 0 and 1
Suppose the current class is divided into V_t^(0) and V_t^(1), and the two subclasses have Ldim values 2 and 4 respectively. Which label does SOA predict?
Identify the candidates: The candidate label 0 is represented by V_t^(0), and the candidate label 1 is represented by V_t^(1).
Compare the values: The Ldim of V_t^(1) is 4, which is larger than the Ldim of V_t^(0), which is 2.
Select the prediction: SOA predicts the label associated with the larger Ldim, so it predicts 1.
Separate prediction from update: The later update still depends on the true label y_t. The prediction itself does not determine which hypotheses remain.
SOA predicts 1 because V_t^(1) has the larger Ldim. The next class is formed only after the true label is observed.
Updating the Consistent Class
The current class V_t represents the hypotheses that have survived all previously observed labels. Once the true label y_t is revealed, SOA removes every hypothesis in V_t that disagrees with y_t on x_t. The survivors form V_(t+1). This means that the version space can become smaller as observations rule out hypotheses, but the rule for removal is always agreement with the true label.
A Wrong Prediction Still Uses the True Label
Suppose SOA compares the two candidate subclasses and predicts 0, but the observed true label is 1. Which label controls the construction of V_(t+1)?
Record the prediction: The prediction is p_t = 0 because SOA selected the subclass associated with label 0.
Record the observed label: The true label is y_t = 1, so the prediction was a mistake.
Apply the update rule: The update keeps only hypotheses that assign label 1 to x_t.
Form the next class: Those surviving hypotheses form V_(t+1). Hypotheses are not kept merely because they supported the prediction 0.
The true label y_t = 1 controls the update, even though SOA predicted 0.
Tie-Breaking Rule
When V_t^(0) and V_t^(1) have equal Ldim, neither candidate has a larger value. SOA resolves this tie by predicting 1. The tie rule is part of the prediction procedure, so a trace that predicts 0 in this situation does not follow SOA, even if its later update uses the true label correctly.
Equal Candidate Values
Suppose both V_t^(0) and V_t^(1) have Ldim equal to 3. What does SOA predict?
Compare the candidates: The two candidate subclasses have equal Ldim, so there is no larger candidate.
Apply the tie rule: The specified tie-breaking rule is to predict 1.
SOA predicts 1.
Checking a Prediction Trace
To check whether a prediction trace follows SOA, inspect the events in order. First verify that the trace starts with V_1 = H. Then check that it divides the current class according to the two labels assigned to the current input. Next verify the Ldim comparison and the resulting prediction, including the rule that a tie produces 1. Finally check that the update uses y_t, the true label, rather than p_t, the prediction. The first mismatch in this sequence is where the trace diverges from SOA.
Updating with the prediction instead of the true label
The update is defined by y_t, not p_t.
Fix:
Keep only the hypotheses that agree with the observed true label 1.Choosing the smaller-Ldim subclass
SOA predicts using the candidate class with larger Ldim.
Fix:
Compare both candidate values and select the label associated with the larger value.Breaking an Ldim tie in favor of 0
SOA's specified tie-breaking rule is to predict 1.
Fix:
Predict 1 whenever the two candidate Ldim values are equal.Treating prediction and update as one event
The selected subclass reflects the prediction, while the true label determines the update.
Fix:
Record p_t first, observe y_t separately, and then construct V_(t+1) using y_t.
Mistake Bound
The formal guarantee stated for SOA is that the number of mistakes it makes on hypothesis class H is at most Ldim(H). Thus, Ldim has two roles in the procedure: the current candidate Ldim values guide each prediction, and the Ldim of the original class gives the upper bound on the total number of mistakes.
Mistake bound: SOA makes at most Ldim(H) mistakes on H.
Practice Check
A current class is divided into V_t^(0) and V_t^(1). Their Ldim values are equal. SOA predicts a label, then the true label is revealed as 0. State the prediction and describe which hypotheses belong in V_(t+1).
Hints
- Use the tie-breaking rule for the prediction.
- Use the true label, not the prediction, for the update.
A trace begins correctly with V_1 = H and forms the two candidate subclasses. It then predicts 0 even though the Ldim of V_t^(1) is larger. Later, it updates using the true label. At which checkpoint does the trace first diverge from SOA, and why?
Hints
- Check the Ldim comparison before checking the update.
- The first incorrect checkpoint determines the answer.
- SOA starts with V_1 = H and maintains the hypotheses consistent with the true labels observed so far. At each round, it divides V_t into the two candidate subclasses V_t^(0) and V_t^(1), compares their Ldim values, and predicts the label associated with the larger value. Equal values are resolved by predicting 1. After the true label arrives, SOA forms V_(t+1) using that true label, not the prediction. Its total number of mistakes on H is at most Ldim(H).
Key Takeaways
- SOA begins with V_1 = H and maintains hypotheses consistent with observed true labels.
- It compares the Ldim of V_t^(0) and V_t^(1) and predicts the label with larger Ldim.
- When the candidate Ldim values are equal, SOA predicts 1.
- The update uses y_t, the true label, rather than p_t, the prediction.
- SOA makes at most Ldim(H) mistakes on hypothesis class H.