The real material is rarely on a web page
Ask a business where the exact answer lives and the honest answer is usually a file. The website carries the summary, the reassurance and the button. The fee schedule, the eligibility criteria, the installation instructions and the terms of business are documents, because they were made to be printed, signed, posted or attached to an email.
There are good reasons for this and they are not going away. A document has a version number and a date on its front page. It looks the same on every screen and out of every printer. It was approved once and it cannot be quietly edited by whoever last had access to the website. In several trades it is the artefact with standing, and the web page is a description of it.
So when somebody sets out to make their material answerable, the most useful documents in the business turn out to be the ones in the worst shape for the job. The first instinct, which is to upload the lot and move on, is exactly the one that produces the strangest failures, and it produces them quietly.
This format does not contain paragraphs
It is worth knowing what the file actually is, because every failure below follows from it. At the level that matters here it is close to a set of drawing instructions: place this character at this position on this page, in this font, at this size. It does not, in the general case, record that these characters form a word, that these words form a sentence, or that this sentence belongs to the paragraph above rather than the one in the next column.
Extraction software reconstructs all of that by inference. It groups marks that sit close together into words, groups words into lines, and then guesses the order the lines should be read in. On a single column of body text with generous margins that guess is almost always right. On anything else it is a guess, and it is made without any way of telling you how confident it was.
Which is why the same tool produces a clean transcript of one document and nonsense from another that looks, on screen, almost identical. Nothing about the appearance tells you which one you are holding. The only way to find out is to pull the text out and read what comes.
Two columns become alternating lines
A layout with columns is the most common cause of text that is technically present and practically useless. A reader working line by line across the page joins the first line of the left column to the first line of the right, then the second to the second, all the way down. The result is grammatical fragments spliced into each other, and it is obvious once you see it.
The version that is not obvious is worse. A fee list or a contents page set as columns interleaves cleanly enough to read like prose. The service name from one row lands next to the amount from another. Nothing looks broken. A passage like that will match a question about your fees perfectly, and the answer it supports will be wrong in a way nobody can spot without the original in front of them.
Sidebars, pull quotes and boxed notes do the same at a smaller scale, dropping a marketing sentence into the middle of a clause about your notice period. The detection method is the same for all of it and takes five minutes: copy the entire document into a plain text editor and read the first two pages of what appears.
The scan with nothing underneath it
The single most common failure is a document that contains no text at all. Anything printed, signed and scanned back in, anything produced by a copier, anything photographed, is a picture of a page. Extraction returns nothing, or a few stray characters lifted from a stamp.
The dangerous part is that this is silent. The upload succeeds. The file appears in the list with a name, a size and a date, and looks exactly like the ones that worked. It simply contains nothing, so the question it was meant to answer gets refused as though you had never provided it, or answered from some other document that was less relevant and merely legible. The test takes a second: open the file and try to select a line of text with the cursor. If nothing highlights, there is nothing there.
Text recognition is the obvious remedy and it is a partial one. Error rates that are tolerable on prose are not tolerable on figures, and a single misread digit in a fee table produces an answer that looks completely normal. If the document is mostly numbers, the honest move is to retype the numbers onto a page and check them against the original, rather than trusting a conversion nobody proofread.
The furniture that repeats on every page
Running headers and footers are designed to be ignored by a human eye and are not ignored by anything else. The company name, the document title, a page number, a confidentiality line, a web address, sometimes a draft marking. Extraction usually keeps them in the flow, which means they appear in the middle of the running text once per page.
Two things go wrong. The first is dilution: a passage that should be about your cancellation terms is partly about your company name and a page number, which weakens the match and wastes the room. Across a whole corpus, the words that repeat most often become the ones carrying no information at all.
The second is more specific and more expensive. Sentences that run across a page break get a header and a number inserted into the middle of them, and the break tends to land where the qualifying clause sits, because that is where sentences get long. And a footer reading internal use only, or draft, or a superseded document reference, becomes ordinary text that can be quoted back at a customer in a sentence you never wrote.
The revision date on page one that never travels
This is the failure that costs the most and gets noticed the least. Documents that matter announce their version at the front: issued in a given month, edition four, effective from a date, supersedes the previous version. A person reading page nine still has all of that, because they saw page one and because they are holding one physical object.
A retrieved passage from page nine has none of it. It arrives as a paragraph about your cancellation fee, undated, indistinguishable from a paragraph written this morning. If an older edition is also indexed, or still sitting on the website, or still in a shared drive somebody pointed a crawler at, then two undated passages compete and the winner is whichever happened to match the question's wording more closely.
No setting fixes this, because the missing information is genuinely not in the passage. The fix is in the document. Put the effective date in the same sentence as anything time bound, at least once per section rather than once per document, and unpublish superseded editions instead of filing them tidily beside the current one. A document nobody can date is a document that will eventually be quoted after it stopped being true.
Tables inside a document fail twice
A table in this format has two separate problems stacked on top of each other. The first is that it is a table, which is a shape rather than a sequence of sentences and does not survive being read as one. The second is that in many files there is no table structure at all: there is text positioned in a grid, and if you are fortunate, some lines drawn around it.
The practical consequence is that cells get read in an order set by position rather than by meaning, header rows end up a long way from the rows they label, and a merged cell either repeats itself or disappears. If the figures in that table are ones you would not want misquoted, it is the strongest single argument for rewriting the document as a page rather than indexing it and hoping.
Rewrite it, or index it as it is
The decision is per document rather than per business, and three questions settle almost all of them. How often is this asked about. How often does it change. Does the file itself have to exist as an artefact, or is it just where the information happens to live.
Rewrite it as a page when it is asked about often, changes on your own schedule, and is a description rather than an instrument. That covers fee schedules, service descriptions, delivery terms, eligibility criteria and most manuals. The page becomes the thing you maintain and the thing that gets read, which is better for search, better on a phone, and better for the member of staff who currently has to open a file to answer a question on the telephone.
Keep the file and index it as it is when it extracts cleanly, changes rarely, and is genuinely an artefact: a contract template, a form somebody prints and signs, a manufacturer's datasheet, a regulatory notice. And where the file matters but does not extract, do both. Publish a page carrying every fact the file contains, and link the file from that page as the record. It is more work than uploading, and it is the only arrangement that is reliably answerable and reliably authoritative at the same time.