Conditional Probability and Bayes' Rule
Try it: Conditional Probability and Bayes' Rule
Conditional probability with natural frequencies: a population is split by the prior P(A) and by P(B | A), P(B | not A); the joint table gives P(B) and Bayes' rule P(A | B) = P(B | A)·P(A) / P(B). The Bayes-optimal classifier predicts the label with the larger conditional probability, and its error is the smallest any rule of the observation can reach.
How it works
- Split 10,000 people by the prior: 10,000·P(A) are in A.
- Split each group by B: P(B | A) of A and P(B | not A) of not A have B — four exact joint counts.
- Total probability: P(B) = P(B | A)·P(A) + P(B | not A)·P(not A).
- Bayes' rule: P(A | B) = (count of A and B) / (count of B); likewise P(A | not B).
- Bayes-optimal classifier h*(x): for x = B and x = not B predict A when η(x) = P(A | x) ≥ 1/2, else not A; count its mistakes and compare with other rules.
Default run (10 steps): A population of 10,000 people. P(A) = 0.05, P(B | A) = 0.9, P(B | not A) = 0.1. … Error of h*: it is wrong on 500 of 10000 people = 1/20 (0.05). "Predict A exactly when B" errs on 1000 (0.1); ignoring x and always predicting the majority label errs on 500. No rule that only sees x does better than h*.
Simplified: Probabilities are whole percentages so every natural-frequency count out of 10,000 is exact; one binary event B plays the role of the observed feature x.
Loading the simulation…