How a model learns
Twelve shots: tokens, a neuron, layers, the error, training, frozen weights, the random pick and what happens to your data. The numbers in the film are illustrative, a teaching example.
As text
The same as the films: every shot and its text. You can copy the text and give it to your own assistant along with your question.
How a model learns
The answers of a chat assistant are written by a language model, a program that continues text word by word. Inside it there is no reference book and no written rules of language: there are only numbers, billions of them in a large model, and not one of them was written in by a person. How the program picks these numbers by itself explains the model's main property: the answer sounds smooth and confident even when the fact in it is wrong.
A program cannot compute with letters, so the text is first turned into numbers. It is cut into tokens, that is, words or parts of words, and each token has its own number in the model's vocabulary. The vocabulary holds tens or hundreds of thousands of pieces, and any word can be built from them, even a rare one. The sentence “The defendant appealed” fell into four tokens, because the program cut “defendant” in two. The split and all the numbers in the film are illustrative.
With token numbers the model does one thing: from the start of a text it estimates which token comes next. It answers not with one word but with a probability for every token in the vocabulary. A probability is a number from zero to one: the larger it is, the more plausible the continuation, and all the options together add up to one. After “The defendant” the word “appealed” got 0.62, about six chances in ten.
The model computes the probabilities, and the smallest unit of this computation is a neuron, in its simplest form a small formula. Each input number is multiplied by its weight, the results are added, and the sum is squeezed, here into a number between zero and one: 1.28 becomes 0.78. A weight shows how strongly an input affects the result: with a weight of 0.9 it counts for a lot, with 0.2 for almost nothing. In the picture, a weight is the thickness of a wire.
One neuron can do little, so neurons are gathered into layers, and the layers are placed one after another. The output of each neuron becomes an input for every neuron of the next layer, and each such connection has its own weight. This network has three layers of five neurons and fifty weights between them. A large language model is built in a more complex way, but the principle is the same, and it has billions of weights.
At the start all these weights are random, and the network computes nonsense with them. The token numbers of “The defendant” pass layer by layer, and at the output all the options get almost the same probability: “appealed” gets 0.19, no more than any other word. Setting the right weights by hand is impossible: there are too many of them, and nobody knows what they should be. So the program finds the weights itself, and this search is called training.
Training needs only ready-made texts: in each of them the next word is already known. The model is shown the start, “The defendant”, and gives the word “appealed” a probability of 0.19. In the text the next word really is “appealed”, so the correct probability is one. The difference, 0.81, is the error here. Real training uses a different formula, but the idea is the same: the less the correct token got, the larger the error.
The error is passed back through the network, from the output to the input, and each weight is shifted slightly in the direction where the error gets smaller. One such step changes almost nothing, so millions of them are made, each time on a new passage of text. The probability of “appealed” gradually rises from 0.19 to 0.62. It will not reach one, because in other texts the same words are followed by “objected” or “paid”.
When training ends, the weights are fixed, and that makes a version of the model. Of the texts it read, only these numbers remain: patterns of language, not a database where a fact could be found and checked. A chat conversation only reads the weights to compute an answer and changes nothing in them. If an assistant remembers past chats, those are saved notes it adds to a new request, not training.
The next version is trained on a new set of texts, and the provider may add users' conversations and files to it. Whether this applies to you depends on the product, the plan and your account settings. The switch in the settings works only going forward: what has already entered the set cannot be taken back. Even when you turn training off, the provider keeps conversations on its servers for some time, and in some cases its staff read them.
For the same request a finished model computes practically the same probabilities, yet the answers differ. The next token is drawn from them at random, but so that a likelier option comes up more often: “appealed”, with a probability of 0.62, wins most of the time, but not every time. Of three runs, two gave “appealed” and one gave “objected”. The chosen token is added to the text, and everything repeats for the next one until the answer is complete.
The model was trained for one thing: to continue text the way it is usually continued. On a common phrase such a continuation is mostly correct. A rare fact, such as a case number, hardly ever appeared in the texts, yet the model produces something similar in form just the same. It has internal signals of whether a topic is familiar, but they are unreliable and invisible in the answer: the known and the merely plausible sound the same. That is why facts are checked against the source.
Next
- Where the answer comes from · How a language model works