Search
Many work tasks start with a search: for a company, a document, a person, a provision, a court decision, statistics. What you find depends not only on the query but also on where you search and what got there in the first place. This practicum covers how search engines, databases and search by a language model work, which data is open and which is not, and how to combine tools to reach what you need. The exercises use legal databases because exact-rule search is easiest to see there; the same techniques work in library catalogues, article databases, media archives and company document stores.
- 1Spiders read the web
- 2Search looks in an index
- 3Queries and Boolean search
- 4Semantic search
- 5Search by a model and limits for bots
- 6Kinds of data
- 7How to reach the data
- 8Check yourself
- 9Practice
Spiders read the web
The web is read mostly by programs, not people. A program that opens a page, saves a copy and follows every link further is called a spider, or crawler. Each new address gets one more spider, so within a few steps one spider becomes a hundred.
A spider reads everything a page shows without a login: text, names, phone numbers, dates, links. Spiders are run by search engines, web archives, monitoring services, companies that train models and open-source researchers. So treat anything published openly as read and saved: if you delete a page or a post, copies may remain.
A site owner has three different tools. robots.txt asks spiders not to enter a section; it is a request, not a lock, and only polite spiders obey it. A noindex tag tells a search engine not to show the page in results. Only a password login actually closes a page.
Film. One spider reads a page, the links become a queue of addresses, the spiders multiply into dozens, and their copies make up the index that a query later searches.
More: what follows for work
For what is published. A phone number in an open listing, a surname in a comment, a draft on a site without a password: all of it can end up in other people's copies before the page is fixed. Confidential material is closed with a login, not with a ban in robots.txt.
For searching. A copy has a date. A search engine shows a page as the spider last saw it, so open the page itself before you cite it. If the page has already been changed or removed, an older copy can sometimes be found in a web archive.
For your own collecting. That a program can technically read a page does not mean it is allowed to: there are the site's terms, copyright and personal data rules. Look for an API or an open dataset first.
Search looks in an index
When you type a query into Google, the system does not go through the internet at that second. It searches an index (in Ukrainian): a list, built in advance, of words with links to the pages where they occur, like the index at the back of a book. The index is filled by a spider program: it follows links from page to page and saves copies.
Hence the main rule of search: you can find only what got into the index. Pages behind a password, results of searching a register through its form, pages with a noindex tag and new pages the spider has not reached yet do not get into the index. A ban in robots.txt is different: it keeps the spider from reading the page, but the address can appear in results without text if other pages link to it. Google recognizes scans without a text layer by itself, but the recognized text has errors, so an exact phrase may not be found. “Nothing found” (in Ukrainian) means “not in this index for this query”, not “does not exist”.
The system orders the pages it finds by many features: where the words stand, how many other sites link to the page, how fresh it is. The first ten results are an order, not a rating of truthfulness.
Film. Spider, index, query and ranking: how a search engine finds pages and why part of the web is invisible to it.
A little internet
Nine pages of a made-up law firm. Run the spider from the home page, then search what it collected.
More: why results differ between people and days
Search engines take into account language, country, device and, for users with an account, search history. The index is updated all the time: pages are added, changed and removed. So the same query gives different lists to two people or on different days.
For work that has to be reproduced, record the query, the date and the search engine, and save the pages you find together with the date you saved them. A page that no longer exists can sometimes be found in web page archives.
Queries and Boolean search
Databases of court decisions, registers and libraries often search by exact rules. Boolean search (in Ukrainian) joins words with operators: AND (both words in the document), OR (at least one), NOT (exclude a word). Quotes search for an exact phrase. Such search is predictable: the same query in the same system on the same date gives the same result, and you can record it and repeat it. When documents are added to the database, the result changes, so the record needs a date.
Search quality is measured with two numbers. Precision (in Ukrainian): what share of what you found actually concerns the task. Recall: what share of everything relevant you found. A narrow query raises precision and loses recall, a broad one does the opposite. For checking a counterparty, collecting case law or reviewing research, recall usually matters more: a missed document is worse than an extra one.
Many databases offer field search (in Ukrainian): in a database of decisions these are court, date, case category, number; in a database of articles, author, journal, year. A field cuts out noise more precisely than a word in the text.
Find the decisions
A collection of twelve short made-up decisions. Task: find all decisions on terminating a lease for nonpayment. Write a query with the operators AND, OR, NOT, brackets and quotes.
Semantic search
Laws and decisions call the same thing by different names: “terminate”, “withdraw from the contract”, “rescind”; “lease” and “tenancy”. Procurement documents do the same: “supplier”, “vendor”, “contractor”. Search by exact words misses synonyms, so for recall the query is widened by hand, with OR.
Semantic search (in Ukrainian) does this differently: a model turns the text of the query and of the documents into sets of numbers that describe their meaning, and looks for documents with close meaning, even without shared words. It finds more of what you need, but also more of what is similar and irrelevant: an employment contract also gets “terminated”. The result cannot be reproduced exactly by a rule, so for a conclusion the results are still checked and selected by hand.
In the exercise above, switch on “Semantic mode” and watch how precision and recall change.
Search by a model and limits for bots
An assistant with search turned on writes several queries by itself, calls a search engine, opens a few pages and writes an answer from what it managed to read, with links. This is handy for a first overview of a topic. But the answer rests on a few fragments it found, not on the whole database, and it can sound complete even when the key decision is not in those fragments.
Part of the web is closed to bots on purpose. Sites limit automated access with rules for robots, request limits, “I am not a robot” checks, terms of use and paid access. The reasons: load on servers, protection of their own content and of personal data. A person in a browser and a bot see different things.
Machines have their own doors. An API returns data in a structured form, with a key and with limits: for example, a register record as a set of fields. Open data can be downloaded as a whole dataset. So a working approach combines tools: a model for the plan and the draft, a search engine for finding pages, a register or its API for facts, and the primary source a person opens and reads themselves.
Film. What an assistant does when it “searches the internet”, what it does not see and why, and which doors are open to machines.
Kinds of data
Data differs by who publishes it and on what terms you can access it. The open web: ordinary pages that search engines index. Open data: datasets that the state or an organization publishes for reuse, for example on Ukraine's open data portal data.gov.ua, in CSV or JSON formats, sometimes with an API. Public registers: the records are open, but you have to search them through the register's own form or API, because a search engine does not see the results.
Paid databases: commercial legal, financial and research databases, register aggregators, with processed data and convenient search, by subscription. Closed data: personal data, bank details, internal documents; access only on a legal basis. Separately stands information that a public authority holds but has not published: you can obtain it with a request for public information, which the holder must answer within the time limit set by law.
Another division concerns form. Structured data has fields: code, date, amount; it is easy to filter and compare. Unstructured data is the text of decisions, letters, contracts: you search it by words or meaning, and the fields have to be extracted separately.
Film. Layers of data from the open web to closed information, and the route to each layer.
How to reach the data
For each need, choose the best route. Often there are two right routes: a quick one first, then an exact one.
Which route
More: downloads, APIs and collecting pages
Downloading a dataset. If the data is published as an open dataset, it is easier to download the whole file and search in a table than to click through pages one by one. Check when the dataset was last updated.
API. Programmatic access with a key returns records as fields. You need it when checks are many or repeated: for example, a hundred counterparties every month. The key belongs to an account, requests have limits, and the terms of use need to be read.
Automated collecting of pages. A program can save a site's pages, but the site may forbid it in its terms or limit it technically, and the pages may contain personal data. Before collecting, look for an API or an open dataset and check the site's terms.
Check yourself
Practice
- Three searches for one companyFind one company in three ways: with a search engine by its name, in the register by its company code (EDRPOU), and with a search engine using the site: operator on the court decisions site. Write down what each way gave and what it did not.
- Precision and recallIn the register of court decisions, find decisions on a narrow topic with two queries: a narrow one and a broad one, with synonyms joined by OR. Count how many each found and how many of them are actually about your topic.
- Assistant versus registerAsk an assistant with search to find three decisions on the same topic. Compare with what you found yourself in the register: what matched, what the assistant did not find, what it named inaccurately.
- Open dataFind a dataset on data.gov.ua that is useful for your work. Download it, open it in a spreadsheet and write down who publishes it, when it was updated and what fields it has.
What to take away
- Spiders read the open web. Treat anything published without a login as read and saved.
- Search looks in an index. “Nothing found” means “not in this index for this query”.
- A narrow query is more precise, a broad one has better recall. In checks and reviews, a missed document often costs more than an extra one.
- Semantic search finds synonyms, but also similar irrelevant material. The results are selected by hand.
- An assistant with search reads a few fragments, not the whole database.
- Each layer of data has its own route: search engine, download, register or API, subscription, legal basis, request.
- Record the query, the source and the date, so the search can be repeated.