Search by a model and limits for bots
An assistant with web search does not know the answer in advance: it searches, opens a few pages and retells them. The film shows this path and the places the assistant cannot reach.
Worth watching first:
- Search looks in an index · Search
- How a model learns · How a language model works
As text
The same as the films: every shot and its text. You can copy the text and give it to your own assistant along with your question.
Search by a model and limits for bots
A person asks the assistant about something recent: court case law from the last six months, a change in the law, yesterday's news. The screen shows “searching the web”, and half a minute later an answer arrives with links. The film shows what happens in that half minute: who actually searches, which pages the assistant reads and which it does not see at all. Knowing this makes it easier to decide what in the answer to check yourself.
The assistant's answer is written by a language model. It knows only what was in the texts it was trained on, and that knowledge ends at the training date. Rulings issued later are not in it, so it cannot answer a question about the Supreme Court's case law from the last six months on its own. Anything newer the assistant takes from outside, and the answer depends on what it manages to find.
On its own the model cannot fetch anything from outside: it is not connected to the internet, it only takes text in and gives text back. In everyday work this is hardly visible, because the assistant is given tools, that is, separate programs it can call: search, opening a page by its address, reading a file. When the assistant “searches”, a tool does the work, and the model gets ready text from it.
The search tool is, in essence, an ordinary search engine, except that the queries are typed in by the assistant, not by a person. From the question it builds several queries and gets back a list of links, each with a snippet of a line or two. At this step the model has not read any of the pages. It judges the content by the address and the snippet, and uses them to decide what to open.
Of the five links the assistant opens three: a second tool downloads these pages and puts their text into the context window, that is, the model's working memory for the length of the conversation. The model writes the answer from this text, and the numbers in brackets show which page each statement came from. Pages that did not get into the window have no effect on the answer, whatever they say.
Not every page opens, though. The tool comes to a site as a visitor without an account, so everything visible only after logging in is closed to it: email, an online bank account, a work drive, a private group. Paid publications and journals usually show it the headline and the first paragraph. Registers where the result appears only after a query in a form mostly do not let it through either.
What a site closes off depends on what it values. Banks and hospitals guard the personal data of clients and patients, because the law requires it. Publishers and paid databases guard the text, because they sell access to it. Social networks and marketplaces mostly limit automated reading so that their content is not copied in bulk. State registers often add an “I am not a robot” check to cope with the load.
There are several means for this. In a robots.txt file the site owner writes where programs should not go; it is a request, not a lock. A rate limit stops anyone who sends requests too often, and a CAPTCHA, the “I am not a robot” check, tells a person from a program. There are also terms of use that forbid automated collection, and paid access. That is why a person opens an address in a browser while the tool gets a refusal at the same address.
The answer usually says nothing about what the tool did not open. It has links and a confident tone, so it looks complete. Yet the model wrote it from only the few pages it managed to read, and closed pages, paid texts, registers and very new publications did not get into it. The answer itself does not show this: the assistant usually gives no list of what stayed outside the search.
The assistant reaches closed data only when it is given a separate entrance. Email, a drive or a database is connected through an API, that is, an entrance made for programs, not for people. You allow the connection yourself, and from then on the assistant reads on your behalf and sees exactly what your account is allowed to see. Without such a connection this data does not exist for it.
In work these means are combined. The model helps draft a plan and a first version, but it can make mistakes. Search finds open sources, though not all of them. A connected database or register gives exact records, but no conclusions. The last step stays with the person: they open and read the source a conclusion rests on themselves, and remember that the assistant did not see the closed sources.
Next
- Kinds of data and how to reach them · Search
- A request has parts · Working with an assistant