Spiders read the web
Everything on the open web is read first by programs and only then by people. The film shows how these programs work, what they keep, and what actually closes a page to them.
As text
The same as the films: every shot and its text. You can copy the text and give it to your own assistant along with your question.
Spiders read the web
A spider is a program. It opens a page at an address and reads everything the page shows without a login: every word, the phone number, the address, every link. What it reads, it keeps as a copy.
To a spider, every link on a page is a new address. The spider writes the addresses into a queue, and another spider goes out to each one. On the new pages they find more links, and the queue grows.
Each spider that finds a link starts another one. Six doublings later one spider has become sixty-four, reading twelve pages at once, word by word. Real systems run thousands of spiders, day and night. Not only search engines run them: web archives, monitoring services, companies that train models and open-source researchers do too.
A site owner has three different tools. A robots.txt file asks spiders to stay out of a section: polite spiders listen, others walk in, because it is a request, not a lock. The noindex tag tells a search engine not to show the page in results; to see the tag, the spider has to read the page. Only a login actually keeps a spider out.
The copies are taken apart into words, and the words become an index: a word and the list of pages where it appears. An index is a snapshot taken on the day of the crawl. If a page changes after the crawl, the index keeps the old copy until the spider comes back.
When you search, no spider runs anywhere. The system takes the words of the query, looks them up in the index and shows the pages that have all of them. What is not in the index, such as a page behind a login, search cannot show.
The same goes for everything you post in public. Spiders read the post and carry copies away: into a search engine's index, a web archive, someone else's database. Delete the post, and the copies can stay. So treat anything published without a login as read and saved.
Next
- Search looks in an index · Search