Metrics and terms

This one belongs to what you wrote, not to what answers

Two businesses on identical software, identical settings and identical traffic will get completely different results, and the difference is almost never the assistant. It is whether the answer to the question was written down anywhere, in a form something could reach and retrieve. That is coverage, and it is the variable with the largest effect and the least attention.

not measured here

No coverage figure is computed. The sources screen shows how many documents each source holds and whether it is ready, which counts what was taken in. The insights page shows the share of assistant messages that were not refusals, labelled "Deflection" and hinted "answered, not refused", and a list headed "Unanswered questions" with a best match percentage against each row. Those describe the traffic that arrived. None of them is a share of what could be asked, and nothing anywhere compares your material against a set of questions you did not receive.

What it means

Source coverage is the share of what visitors want to know that some piece of your indexed material could answer. Stating it that way makes the problem obvious: the denominator is everything anybody might reasonably ask, and you can never see that. You see the questions people typed, which is already filtered by who bothered, who found the widget, and who had not already given up. Coverage against asked questions is knowable after the fact and is what the refusal figures describe. Coverage against askable questions is what actually determines whether the assistant is any good, and it cannot be computed by anybody, here or elsewhere. What can be done is to enumerate the limits on it, because each one is specific, each one is fixable, and together they explain most of the difference between an assistant that works and one that does not.

How it is actually calculated

The four limits, in the order they bite

Whether it is written down at all. A policy that lives in the head of the person who has been there longest cannot be covered by anything. This is the largest gap in most small businesses and it is the one nobody thinks of as a software problem, because it is not one.

Whether a source can reach it. Material behind a login, in a folder nobody connected, inside an email thread, or on a page nothing links to is outside every crawl and every connection. It exists, it is written down, and it is not coverable in its current location.

Whether there was room. A crawl takes up to the pages your plan allows, less whatever is already indexed, and stops. A large site under a modest allowance is covered in part, and which part depends on where the crawl got to rather than on what mattered.

Whether a retriever can find the answer inside what was taken in. A sentence about refunds buried in a nine thousand word general page is indexed and effectively uncovered, because the passage that wins a match is about six things and commits to none of them.

Why a document count is not a coverage figure

The number of documents indexed is shown per source, and it is a count of what was taken in. It says nothing about what fraction of your visitors' questions has an answer somewhere inside them, and the two can move in opposite directions.

Adding a large undifferentiated document set raises the count and can lower real coverage, because everything now matches everything weakly. A short set of focused pages, one topic each, can cover far more of what is asked with a tenth of the volume. If you are going to watch one thing about your material, watch whether each document answers a question somebody would ask.

The evidence you do have, and what it can carry

The refused questions are the honest part. Each one is a question that arrived and found nothing close enough, with a figure showing how near the best passage came, and together they describe your coverage against the traffic you actually received.

The share of answers that were not refusals is the same information as a level rather than a list. Read as coverage it is badly behaved: it moves with the mix of questions in a given week, and it moves with the threshold setting, neither of which is a fact about your documents.

The measurement that works is one you construct. Write down the twenty things a customer most often needs to know, ask each in the widget in a customer's words, and record whether the answer was right and which document it cited. That is coverage measured against a denominator you chose deliberately rather than one your traffic chose for you, and it is repeatable next quarter.

How the number gets moved without anything improving

How coverage appears to improve while less gets answered

Upload everything. The document count climbs, the sources screen looks substantial, and matches get vaguer across the board because there is now always something loosely relevant. Refusals fall, which reads as coverage, and the answers that replaced them are built on the least specific material you own.

Report one minus the refusal share and call it coverage. It is a figure about your visitors rather than your documents: an easy week reads as good coverage and a week where somebody linked to you from a new place reads as bad. Nothing about the material moved either time.

Loosen the threshold and the same questions that were uncovered yesterday are covered today. This is the cleanest demonstration that the figure people use as coverage is a setting rather than an observation.

What to look at instead, or alongside

  • Twenty questions you wrote down in advance, asked in the widget and scored by hand, which is coverage against a denominator you can defend.
  • The refused questions with their best match figures, which separate a missing topic from a wording mismatch.
  • Whatever people search for on your own site, which covers the visitors who never opened the widget at all.
  • An honest list of what is not written down anywhere, since no amount of indexing reaches material that does not exist.

Questions

Does uploading more material always improve coverage?
No, and past a point it reliably makes things worse. What helps is material that answers one question clearly. What hurts is length without focus, because a long general passage matches many questions and satisfies none of them, and it displaces the specific page that would have answered properly.
Can the page allowance on my plan limit coverage?
Yes, directly. A crawl takes up to the pages your plan allows minus whatever is already indexed and stops there, and when there is no room left the source fails with a message naming the limit. On a large site that is a real ceiling, and the sensible response is usually to crawl the sections that answer questions rather than the whole thing.
What number do I put in a report?
The result of your own check. Twenty questions, how many were answered correctly, and how many cited the document you expected. Anybody can repeat it, the method is inspectable, and it does not pretend to know about the questions nobody asked.

Keep reading

Try it on your own material

Upload a document or point it at your site, paste one line of HTML, then ask it something only your business could answer.