Training a Keras-Style Model: History and Early Stopping
Try it: Training a Keras-Style Model: History and Early Stopping
The Keras workflow on a tiny dense network: stack Dense layers (the summary counts units * (inputs + 1) parameters per layer), compile with an optimizer, loss and accuracy metric, fit for a number of epochs in mini-batches with validation data, and read the history. Training loss keeps falling while validation loss reaches its lowest point and then rises: the model starts to overfit, and EarlyStopping with patience can stop there and restore the best weights.
How it works
- Build a Sequential stack of Dense layers ending in Dense(1, sigmoid); each layer has units * (inputs + 1) weights and biases, initialised Glorot-uniform with zero biases.
- compile(): mini-batch SGD with the chosen learning rate, binary cross-entropy (or mean squared error) loss, and accuracy (prediction > 0.5) as the metric.
- fit(): each epoch walks the training set in batches of batch_size; every batch runs a forward pass, backpropagates the mean batch loss and updates every weight w <- w - lr * dL/dw.
- After each epoch the history records the epoch's training loss and accuracy (batch-weighted average of values seen during the epoch) and the val_loss and val_accuracy on the held-out validation points.
- EarlyStopping(monitor="val_loss", patience) keeps the best val_loss; after patience epochs with no improvement it stops training, and restore_best_weights puts back the weights of the best epoch.
- evaluate() reports loss and accuracy on the validation data with the final weights.
Default run (153 steps): Sequential([Dense(16, "relu"), Dense(16, "relu"), Dense(1, "sigmoid")]) on inputs of shape (2,). model.summary(): dense 16 * (2 + 1) = 48; dense_1 16 * (16 + 1) = 272; dense_2 1 * (16 + 1) = 17; total 337 trainable parameters, initialised with Glorot uniform (biases 0). … model.evaluate(x_val, y_val) with the weights from epoch 150: loss 1.2082, accuracy 0.75. Training loss went 0.6964 -> 0.0194. Stopping at epoch 32 would have given val_loss 0.5272 and val_accuracy 0.775.
Simplified: A toy re-implementation in TypeScript, not Keras or TensorFlow: 2 input features, at most 3 hidden layers of up to 32 units, a seeded 2-D dataset (inside/outside a circle with label noise; 40 validation points), plain SGD without momentum, shuffle=False so batch order is fixed, float64 arithmetic, and binary cross-entropy computed from the output logit (same value as on probabilities, without clipping). The results match an independent NumPy implementation of the same procedure, not a real Keras run.
Loading the simulation…