·

by

How machine learning models actually learn

The phrase “the model learns” does a lot of quiet work. It suggests something like studying — a system reading examples, forming an understanding, coming out the other side wiser. What actually happens is closer to a very patient guessing game, played a few million times, in which being slightly less wrong is the only reward on offer.

Start with a bad guess

A model begins life full of essentially random numbers. Show it a photograph and ask whether it contains a cat, and it will answer with all the confidence of a coin toss, because that is what it is. The interesting part is what comes next: you tell it the right answer, measure how far off it was, and nudge every one of those numbers a very small amount in the direction that would have made the guess better.

Then you do it again with the next example. And the next. The nudges are individually meaningless and collectively decisive.

Why the size of the nudge matters

Nudge too gently and training takes forever. Nudge too hard and the model lurches past the answer, overcorrects on the next example, and oscillates without ever settling. Most of the practical craft of training sits in this unglamorous space — how big a step to take, when to make the steps smaller, when to stop.

Training is not the model discovering the truth. It is the model settling into the lowest place it can find in the landscape your data happens to describe.

Memorising is not learning

Given enough capacity, a model will happily memorise its training examples outright. It will score beautifully on everything it has already seen and fall apart on anything it has not — the machine learning equivalent of a student who learned the answer key rather than the subject.

The defences are all variations on making memorisation harder than generalisation:

  • Hold data back. Keep a slice the model never trains on, and judge it only on that.
  • Stop early. Training accuracy keeps climbing long after held-out accuracy has peaked. The peak is where you stop.
  • Add noise on purpose. Crop, rotate and distort the examples so the model cannot rely on incidental detail.
  • Keep it smaller than you want to. A model with less room has to find the pattern, because it cannot store the answers.

What this means in practice

Understanding the mechanism changes the questions worth asking about any model you are handed. Not “is it intelligent” but “what did it see, what was held back, and how do you know it is not simply repeating the answer key?” Those are answerable questions, and a team that cannot answer them does not yet know what they have built.