Toolkit by Bot&Partners / Search / watch

Search

Almost all work with information starts with a search. Four films show where a search engine gets its pages, how it picks the first ten, what an assistant does when it searches for you, and where the data lies that search does not see.

4 films

Each film is followed by an exercise. The longer text and the files are in the text version.

Part 1 · about 3 minSpiders read the web

Everything on the open web is read first by programs and only then by people. The film shows how these programs work, what they keep, and what actually closes a page to them.

Keep scrolling

    Try it

    A little internet

    Nine pages of a made-up law firm. Run the spider from the home page, then search what it collected.

      Part 2Search looks in an index

      Between the query and the first ten results, a search engine takes several steps. They decide what you see and why “not found” does not yet mean that the document does not exist.

      Keep scrolling

        Try it

        Find the decisions

        A collection of twelve short made-up decisions. Task: find all decisions on terminating a lease for nonpayment. Write a query with the operators AND, OR, NOT, brackets and quotes.

        Examples:

          Part 3Search by a model and limits for bots

          An assistant with web search does not know the answer in advance: it searches, opens a few pages and retells them. The film shows this path and the places the assistant cannot reach.

          Keep scrolling

            Part 4Kinds of data and how to reach them

            Not all data can be found by search. Some sits in registers, some behind a subscription, some is closed by law. The film sorts it into layers and shows the way to each.

            Keep scrolling

              Try it

              Which route

              Check yourself

              As text

              The same as the films: every shot and its text. You can copy the text and give it to your own assistant along with your question.

              Spiders read the web

              1. A spider is a program. It opens a page at an address and reads everything the page shows without a login: every word, the phone number, the address, every link. What it reads, it keeps as a copy.

              2. To a spider, every link on a page is a new address. The spider writes the addresses into a queue, and another spider goes out to each one. On the new pages they find more links, and the queue grows.

              3. Each spider that finds a link starts another one. Six doublings later one spider has become sixty-four, reading twelve pages at once, word by word. Real systems run thousands of spiders, day and night. Not only search engines run them: web archives, monitoring services, companies that train models and open-source researchers do too.

              4. A site owner has three different tools. A robots.txt file asks spiders to stay out of a section: polite spiders listen, others walk in, because it is a request, not a lock. The noindex tag tells a search engine not to show the page in results; to see the tag, the spider has to read the page. Only a login actually keeps a spider out.

              5. The copies are taken apart into words, and the words become an index: a word and the list of pages where it appears. An index is a snapshot taken on the day of the crawl. If a page changes after the crawl, the index keeps the old copy until the spider comes back.

              6. When you search, no spider runs anywhere. The system takes the words of the query, looks them up in the index and shows the pages that have all of them. What is not in the index, such as a page behind a login, search cannot show.

              7. The same goes for everything you post in public. Spiders read the post and carry copies away: into a search engine's index, a web archive, someone else's database. Delete the post, and the copies can stay. So treat anything published without a login as read and saved.

              Search looks in an index

              1. A search engine does not go through the internet at the moment you type a query: that would take hours. It collects pages in advance. A program called a spider does this: it opens a page, saves a copy and follows the links further. The spider learns new addresses from links or from a sitemap the site owner submits. A page that nothing links to and that is not in the sitemap usually gets missed.

              2. The saved copies are taken apart into words. Each word is reduced to its base form, so a query for “contract” also finds “contracts”, and “terminate” finds “terminating”. The words make up an index that works like the index at the back of a book: next to each word stands the list of pages where it appears. That is why search is fast: the system does not reread pages, it looks in a ready index. Recent changes on a site may not be in it yet.

              3. The system splits the query “lease contract termination” into three words and looks each one up in the index. Pages that have at least one of them become candidates. Page 3 has all three words, page 1 has two, page 4 has one, and page 2, the court news, does not make the list. At this step the candidates have no order yet.

              4. There can be thousands of candidates, and they have to be put in order. The system weighs signals, features it can count: query words in the title, links from other sites, the date of the last update, the searcher's location and language. So the page with all the query words is not necessarily first: here the sample contract, which more sites link to, overtakes it. The top ten reflects this order, not whether what is written is true.

              5. Outside the index stay pages behind a login, databases such as the register of court decisions, where only the search form returns a result, pages with a noindex tag, and new pages the spider has not reached yet. A robots.txt block hides the text from the spider, but the address can still appear in results. Google sometimes recognizes scans by itself, but with errors, so do not rely on searching a scan.

              6. “Nothing found” refers only to this index: it has no pages that match the query. The document itself may still exist, for example in a database the spider does not enter. Then you rephrase the query, try another search engine, or search through the database's or register's own form.

              7. When there are too many candidates, you refine the query with operators, special marks the search engine reads as commands. Quotes keep pages with the exact phrase, a minus before a word removes pages that have it, site: limits the search to one site, and filetype:pdf keeps only PDF files. In the example, four operators cut forty candidates down to three, few enough to check by hand.

              Search by a model and limits for bots

              1. A person asks the assistant about something recent: court case law from the last six months, a change in the law, yesterday's news. The screen shows “searching the web”, and half a minute later an answer arrives with links. The film shows what happens in that half minute: who actually searches, which pages the assistant reads and which it does not see at all. Knowing this makes it easier to decide what in the answer to check yourself.

              2. The assistant's answer is written by a language model. It knows only what was in the texts it was trained on, and that knowledge ends at the training date. Rulings issued later are not in it, so it cannot answer a question about the Supreme Court's case law from the last six months on its own. Anything newer the assistant takes from outside, and the answer depends on what it manages to find.

              3. On its own the model cannot fetch anything from outside: it is not connected to the internet, it only takes text in and gives text back. In everyday work this is hardly visible, because the assistant is given tools, that is, separate programs it can call: search, opening a page by its address, reading a file. When the assistant “searches”, a tool does the work, and the model gets ready text from it.

              4. The search tool is, in essence, an ordinary search engine, except that the queries are typed in by the assistant, not by a person. From the question it builds several queries and gets back a list of links, each with a snippet of a line or two. At this step the model has not read any of the pages. It judges the content by the address and the snippet, and uses them to decide what to open.

              5. Of the five links the assistant opens three: a second tool downloads these pages and puts their text into the context window, that is, the model's working memory for the length of the conversation. The model writes the answer from this text, and the numbers in brackets show which page each statement came from. Pages that did not get into the window have no effect on the answer, whatever they say.

              6. Not every page opens, though. The tool comes to a site as a visitor without an account, so everything visible only after logging in is closed to it: email, an online bank account, a work drive, a private group. Paid publications and journals usually show it the headline and the first paragraph. Registers where the result appears only after a query in a form mostly do not let it through either.

              7. What a site closes off depends on what it values. Banks and hospitals guard the personal data of clients and patients, because the law requires it. Publishers and paid databases guard the text, because they sell access to it. Social networks and marketplaces mostly limit automated reading so that their content is not copied in bulk. State registers often add an “I am not a robot” check to cope with the load.

              8. There are several means for this. In a robots.txt file the site owner writes where programs should not go; it is a request, not a lock. A rate limit stops anyone who sends requests too often, and a CAPTCHA, the “I am not a robot” check, tells a person from a program. There are also terms of use that forbid automated collection, and paid access. That is why a person opens an address in a browser while the tool gets a refusal at the same address.

              9. The answer usually says nothing about what the tool did not open. It has links and a confident tone, so it looks complete. Yet the model wrote it from only the few pages it managed to read, and closed pages, paid texts, registers and very new publications did not get into it. The answer itself does not show this: the assistant usually gives no list of what stayed outside the search.

              10. The assistant reaches closed data only when it is given a separate entrance. Email, a drive or a database is connected through an API, that is, an entrance made for programs, not for people. You allow the connection yourself, and from then on the assistant reads on your behalf and sees exactly what your account is allowed to see. Without such a connection this data does not exist for it.

              11. In work these means are combined. The model helps draft a plan and a first version, but it can make mistakes. Search finds open sources, though not all of them. A connected database or register gives exact records, but no conclusions. The last step stays with the person: they open and read the source a conclusion rests on themselves, and remember that the assistant did not see the closed sources.