How a language model works
How a model learns, how it writes an answer, where the statements in the answer come from and why the model sometimes makes things up with full confidence. Without this, checking the result turns into guesswork: it is unclear what exactly to check and why.
- 1How a model learns
- 2The model continues text
- 3Where the answer comes from
- 4Why the model makes things up
- 5Check yourself
- 6Practice
How a model learns
A language model is a very large formula with billions of numbers called weights. During training the model reads a huge amount of text and at every step tries to guess the next fragment. When the guess is poor, the weights shift slightly. After millions of such steps the model reproduces the patterns of language, style and content of the texts it learned from quite well.
After training the weights are frozen. A chat conversation only reads them to compute an answer and changes nothing in the model. The model does not remember what exactly it read during training, and it holds no database of facts against which an answer could be checked.
Film. Twelve shots: tokens, a neuron, layers, the error, training, frozen weights, the random pick and what happens to your data. The numbers in the film are illustrative, a teaching example.
More: training, fine-tuning and what follows from them
First the model learns to predict the next fragment on general texts. Then it is fine-tuned on sample dialogues and on human ratings so that it answers questions, follows instructions and declines harmful requests. The assistant you work with is the result of both stages plus the application around it.
Three practical points follow. The model knows only what was in the texts up to a certain date, so it may not know about a new law, ruling or report, or may guess. The model reproduces what was frequent in the texts, so a common mistake has a chance of being repeated. And the model has internal signals of whether a topic is familiar to it, but they are unreliable and invisible in the answer: the known and the merely plausible sound the same.
The model continues text
The model writes an answer one fragment at a time. At every step it computes probabilities for the possible continuations and picks one of them. There is no separate “check the fact” step in this process: the correct number and a wrong one can have similar probabilities.
Products pick continuations with an element of randomness, so the same request gives different text. Two matching answers prove nothing: a repeat compares the answers with each other, not with the source.
The next word
A simplified model with a few continuations at each step. The case memo it works from: Sosna LLC rents an office from sole proprietor Tkach from 1 February 2026 for 12 months, rent UAH 30,000 per month, the lease says nothing about a penalty. Press “Next step” and watch the probabilities.
Generated sentences
Try this: generate a few sentences at medium randomness, then set it to zero. At zero the model always picks the most likely continuation, and the sentence repeats. Even then it is not checked, only the most likely.
More: tokens, the window and the confident tone
The model reads and writes text in fragments called tokens. A common English word is often one token; a rarer word, a name or a word in Ukrainian often takes two to four. The model's window is limited by a number of tokens, and both the input and the answer count against it.
A product passes a long document whole, in parts or only as the fragments a search found. Which of these happens depends on the product, the plan and the size. So for a question about a specific section, ask for a reference to the page or clause: that makes checking a short action.
The tone of the answer is also a continuation chosen by probability. It is only loosely tied to correctness, so confident wording says nothing about whether a fact is right.
Where the answer comes from
An assistant is a model plus an application around it. The application assembles the input: standing instructions, project files, the conversation history, your message, results of search or other tools. The model receives all of it as one text and writes a continuation. You see only your message and the answer.
Read the answer not as a whole text but as a set of separate statements. Some of them rest on the document you passed in. Some change its content. Some have no source in the input: they may have come from training, from search or from an earlier conversation, and the text itself does not show which. All three kinds look equally confident.
Film. What the assistant hides behind the chat window: what the input consists of, how the model picks words and how the answer breaks down into statements of different origin.
Where each statement comes from
A colleague asked an assistant to summarize a case memo. Read the case memo, then mark each statement of the answer.
- Veres LLC signed a lease for a warehouse with sole proprietor Melnyk on 14 March 2026. The lease term is 24 months.
- Rent is UAH 48,000 per month. The lease contains no indexation terms.
- The landlord sent a demand letter about arrears on 2 September 2026. The letter does not state the amount claimed.
More: model, assistant, tool
| What | What it does | What it does not do |
|---|---|---|
| Language model | Writes a continuation of text by probability. | Does not see your files on its own and does not check facts. |
| Assistant | The application around the model: assembles the input from instructions, files and history, shows the answer. | Does not show which of your materials actually reached the model. |
| Tool | Search, file reading, a registry, a calculation that the assistant calls while it works. | It is not always visible whether it was called and what it returned. |
Do not silently delete a statement without a source; move it to open questions: it is something to find out separately.
Why the model makes things up
A hallucination is a statement the model presents as fact although it has no support either in the input or in reality. It is not a malfunction and not a lie: the model does what it always does, continues the text plausibly. When the plausible continuation does not match reality, the result is a hallucination.
Hallucinations appear most often where the answer has a specific form and there is no data: a case number, a decision date, an amount, a quotation, the title of an article, a figure from a report. The model knows what such things look like and fills in the form. A made-up citation to a ruling or a research paper looks exactly like a real one.
| Kind | Example | How to find it |
|---|---|---|
| Invented fact | “The lease gives the right to terminate after 30 days of delay”, although the case memo says nothing about it. | Check the statement against the document. |
| Altered fact | Rent of UAH 35,000 instead of UAH 30,000. | Check numbers, dates and names separately from the text. |
| Invented source | A case or ruling number that does not exist, or a real number about something else. | Open the source in the registry and read it. |
| Inaccurate quotation | Quotation marks around text that does not appear word for word in the document. | Search for the quotation in the original (Ctrl+F). |
| Outdated | A version of a law or a price that has since changed. | Check the version in force on the relevant date. |
True or myth
More: what reduces hallucinations and what does not
Reduces: putting the document itself into the request, not just a question about it; asking for an answer with a reference to the clause or page; turning on search with citations if the product has it; asking for a verbatim quotation for every important statement; allowing the answer “the document does not say”.
Does not remove: the request “do not make things up” on its own; repeating the request; a confident tone; a more expensive plan. They may reduce the frequency but do not replace checking.
Hence the working rule: the more costly the mistake, the closer to the source the check must be. For a note to a colleague, rereading is enough. For a document filed in court, a report for investors or a publication, every citation is opened in the original source.
Check yourself
Practice
- Where each statement comes fromGive an assistant the case memo and ask for a five-sentence summary. Split the answer into statements and fill in the origin table. Without an account, work with the supplied answer. The full exercise: on Toolmap (in Ukrainian).
- The same request twiceAsk the same request again in a new chat. Compare the two answers: which statements stayed, which disappeared, which appeared. Write down what this shows and what it does not.
- Check the citationsAsk the assistant to name three rulings of the Supreme Court of Ukraine, with case numbers, on a narrow topic, for example recovering a contractual penalty under a lease. Find each one in Ukraine's Unified State Register of Court Decisions. For each, write down: exists and is on topic; exists but is about something else; does not exist. An alternative for those who work with research: three papers with authors, year and journal title, checked on the journal's site or in a library catalogue.
- Data settingsFind in your assistant's settings the switch for training on your conversations and the retention period for history. Write down which of this applies to your plan.
What to take away
- The model continues text by probability. It has no “check the fact” step.
- Read the answer as a set of statements, not as a whole text.
- A hallucination looks exactly like a fact. Most often it hides in numbers, dates, amounts and quotations.
- A confident tone and two matching runs prove nothing. Checking against the source does.