Metrics and terms

A content backlog your customers wrote for you

This is the most useful output the system produces and the one people look at least. Every row is a question somebody actually typed that nothing in your material covered. It is a list of things to write, ordered by the people who wanted them, which is not a list any internal meeting has ever produced accurately.

shown in the product

The insights page shows a list headed "Unanswered questions", described as "What visitors asked that your knowledge base could not answer. Each one is a gap you can close by adding a document." Each row shows the question, when it was last asked, and a "best match" percentage. When there is nothing to show, the page says every question so far was answered from your knowledge.

What it means

A coverage gap is a question a visitor asked for which the indexed material contained no passage above the match threshold. It is not a model failure and not a retrieval bug: the retriever looked, found the best thing it had, and the best thing it had was not close enough. Treated as an aggregate it is a rate, which is the fallback rate. Treated as a list of texts it is something considerably more valuable, because each entry names a specific missing thing in the visitor's own vocabulary. That vocabulary is half the value. Gaps come in two shapes that look identical in the list and need opposite fixes: the topic is genuinely undocumented, or the topic is documented in words nobody outside the company uses.

How it is actually calculated

How a row is built

For each refused answer, the system takes the visitor message immediately preceding it and treats that as the question that failed. Rows are then grouped by the question text and counted, so the same phrasing asked eleven times is one row showing eleven.

Grouping is by text, not by meaning. The same gap phrased three ways produces three rows, so read the list expecting to merge things yourself rather than expecting it to be pre-summarised.

What each row carries

The question, when it was last asked, and a best match percentage, which is how close the nearest passage came to the threshold. That third column is the one that changes what you do about the row.

A row with a very low best match is a genuinely absent topic. A row with a best match just under the bar is usually a vocabulary problem: something relevant exists and did not quite win. The first needs writing, the second needs the visitor's own words added to material you already have.

Window, ordering and limits

The list covers a recent window and is capped at a sensible number of rows, so on a busy assistant you are reading the top of a longer list rather than all of it. Working through the highest counts first is the right instinct, since those are the gaps most visitors hit.

An empty list is genuinely good news and the page says so plainly rather than showing an empty table. On a new assistant with little traffic it is also uninformative, so read the conversation count next to it before celebrating.

How the number gets moved without anything improving

How the list empties without the gaps closing

Lower the threshold and rows stop appearing, because questions that would have been refused now get answered from whatever weak passage was nearest. The gaps are still gaps. They have simply been converted from a visible refusal into an invisible thin answer, which is a strictly worse position to be in.

Adding one long document that mentions everything has the same effect for the same reason, and it feels like content work while being closer to the opposite.

The accidental version: closing rows by writing a document that repeats the visitor's phrasing without answering the question. The row disappears because retrieval now matches, and the answers built on it say nothing. Check a closed gap by asking the question yourself and reading what comes back, which takes about a minute per row.

What to look at instead, or alongside

  • Read the rows themselves rather than the count of rows: the text is the whole value and a total tells you nothing.
  • Group by meaning by hand for the first few weeks, because the natural groups are fewer than the raw list suggests.
  • Pair each row with its best match percentage to decide between writing new material and rewording existing material.
  • After publishing, ask the question in the widget yourself and read the answer, since a closed row is not the same as a covered topic.

Questions

Why is the same question listed several times?
Because rows are grouped by the exact text. Two visitors asking the same thing in different words produce two rows. Merging them is a job for a person, and doing it by hand for a few weeks teaches you which groups actually exist on your site.
Does a row mean the assistant gave a bad answer?
No. It means the assistant refused, which is the honest outcome. Bad answers do not appear in this list at all, and that is the main reason this list cannot be your only review: it shows you the failures the system admitted to.
How many should we work through in a week?
One or two writing tasks a week is a rate you can sustain for a year, and sustaining it is what closes the backlog. A quarterly effort to clear the whole list produces a lot of documents written in a hurry, which tends to create the vague material that makes the rest of these numbers untrustworthy.

Keep reading

Try it on your own material

Upload a document or point it at your site, paste one line of HTML, then ask it something only your business could answer.