Search looks in an index
Between the query and the first ten results, a search engine takes several steps. They decide what you see and why “not found” does not yet mean that the document does not exist.
Worth watching first:
- Spiders read the web · Search · about 3 min
As text
The same as the films: every shot and its text. You can copy the text and give it to your own assistant along with your question.
Search looks in an index
A search engine does not go through the internet at the moment you type a query: that would take hours. It collects pages in advance. A program called a spider does this: it opens a page, saves a copy and follows the links further. The spider learns new addresses from links or from a sitemap the site owner submits. A page that nothing links to and that is not in the sitemap usually gets missed.
The saved copies are taken apart into words. Each word is reduced to its base form, so a query for “contract” also finds “contracts”, and “terminate” finds “terminating”. The words make up an index that works like the index at the back of a book: next to each word stands the list of pages where it appears. That is why search is fast: the system does not reread pages, it looks in a ready index. Recent changes on a site may not be in it yet.
The system splits the query “lease contract termination” into three words and looks each one up in the index. Pages that have at least one of them become candidates. Page 3 has all three words, page 1 has two, page 4 has one, and page 2, the court news, does not make the list. At this step the candidates have no order yet.
There can be thousands of candidates, and they have to be put in order. The system weighs signals, features it can count: query words in the title, links from other sites, the date of the last update, the searcher's location and language. So the page with all the query words is not necessarily first: here the sample contract, which more sites link to, overtakes it. The top ten reflects this order, not whether what is written is true.
Outside the index stay pages behind a login, databases such as the register of court decisions, where only the search form returns a result, pages with a noindex tag, and new pages the spider has not reached yet. A robots.txt block hides the text from the spider, but the address can still appear in results. Google sometimes recognizes scans by itself, but with errors, so do not rely on searching a scan.
“Nothing found” refers only to this index: it has no pages that match the query. The document itself may still exist, for example in a database the spider does not enter. Then you rephrase the query, try another search engine, or search through the database's or register's own form.
When there are too many candidates, you refine the query with operators, special marks the search engine reads as commands. Quotes keep pages with the exact phrase, a minus before a word removes pages that have it, site: limits the search to one site, and filetype:pdf keeps only PDF files. In the example, four operators cut forty candidates down to three, few enough to check by hand.
Next
- Search by a model and limits for bots · Search
- Kinds of data and how to reach them · Search