Written, 26 May 2026

Four numbers that mean something, and three that mislead

The reporting that ships with support assistants is almost entirely activity reporting: how many, how often, how fast. Activity is easy to count and tells you nothing about whether the thing is working. The useful numbers are all harder to produce and all point at a specific action you could take this week.

Start from the decision, not the dashboard

Before adding a metric, write down the decision it would change. If you cannot name one, the metric is decoration, and decoration in a dashboard is worse than nothing because it consumes the attention that should go to the numbers that matter.

Support assistants generate a small set of real decisions. What to write next. What to fix in existing material. Where to loosen or tighten the refusal threshold. Which topics should stop being automated at all. Whether the thing is worth keeping. Four metrics cover those decisions between them, and everything below is chosen because it maps onto one of them.

The corollary is that a metric which cannot go in a bad direction is not a metric. If a number only ever goes up as you use the product more, it is measuring your usage, not its performance.

Refusal rate, broken down by topic

The overall refusal rate is nearly meaningless. It moves with traffic mix, with seasonality, and with whatever campaign sent people to the site this week. Broken down by topic it becomes the single most actionable number you have.

A topic with a high refusal rate is telling you one of two things, and they need different fixes. Either the material does not cover that topic, which is a writing job, or it covers it in language nobody uses, which is a rewriting job. You can usually tell which by reading a handful of the refused questions: if the vocabulary in the questions does not appear in your material at all, it is the second case.

Watch the direction rather than the level. A refusal rate that climbs in a topic you have not touched usually means the world moved: a policy changed, a product launched, a price went up, and the material is now behind. That signal arrives weeks before anybody would have noticed the drift by reading the pages.

Questions with nothing behind them

This is the highest value output of the whole system and most teams never look at it. Every question that retrieved nothing above threshold is a question a real visitor asked that your published material does not answer. That is a content backlog written by your customers, sorted by their priorities rather than yours.

Group these by meaning rather than by string, because the same gap arrives in a dozen phrasings and looking at raw text makes it appear to be a dozen problems. Grouped properly, a month of unmatched questions usually collapses into a short list of genuinely missing topics, and that list is what you write next.

It is also a check on your assumptions. Teams are consistently surprised by what appears here, because internal ideas about what customers want to know are formed from the questions that reached the team, and those are filtered by whatever the website already answers badly.

Handovers per topic

Total handovers is a workload number. Handovers per topic is a design number, and the two are frequently confused.

Read it two ways. A topic with more handovers than you expected is either genuinely too complex to automate, in which case accept it and make the handover good, or is failing for a fixable reason, in which case the conversations will show you which. A topic with fewer handovers than you expected is more interesting and more dangerous: it may mean the answers are working, or it may mean the assistant is answering confidently in a place where it should be stopping.

Cross that with the topic level refusal rate. High refusals and low handovers in the same topic is the worst combination on the board, because it means people are being turned away and not offered a route out. That pairing is where the handover design is broken rather than the content.

The same question asked twice in one conversation

When a visitor rephrases their question and asks again, they are telling you the first answer did not work. This is the most honest satisfaction signal available and nobody has to press a button to produce it.

It is also more informative than a rating, because the rephrasing itself tells you what was wrong. People rephrase towards what they actually meant. Reading the pair, first attempt and second attempt, shows you the gap between the words your material uses and the words your customers use, which is exactly the thing that determines whether retrieval works.

Count them as a rate, and read a sample every week. If the rate rises after you publish something new, you have probably introduced material that competes with existing pages and is pulling the wrong passage into answers.

The numbers that mislead

Raw conversation counts. This measures how many people found the widget, which is a function of where the launcher sits on the page and how much traffic you had. It moves when marketing runs a campaign. It has never once indicated that answers got better.

Satisfaction on a two button widget. The response rate is low and it is not low at random: people rate when they are delighted or furious, and almost never in between. What you get is a mood reading from an unrepresentative slice, presented as a percentage, which invites exactly the false precision that makes it dangerous. Read the comments attached to the ratings, which are genuinely useful, and treat the score itself as noise.

Anything called containment. The name refers to the share of conversations that ended without a human being involved, and it is trivially gamed by making it harder to reach a human being. It is worth its own article, and it has one.

Average response time from the assistant. It is fast. It will always be fast. It is a property of the software, not an achievement, and putting it on a dashboard next to real metrics gives it a status it has not earned.

A review rhythm that fits in an hour

Weekly, look at two things: the unmatched questions from the last seven days, grouped, and a sample of conversations that ended in a refusal. That is enough to produce one or two writing tasks, which is a sustainable rate.

Monthly, look at refusal rate and handover rate by topic, and compare them to last month. You are looking for direction, not level. Anything that moved noticeably deserves a read of the underlying conversations before you touch a setting.

Quarterly, ask the harder question: which topics should not be automated at all. This is the review that never happens because there is no dashboard tile for it, and it is the one that most improves an assistant after the first few months. Some topics get worse the more effort you put into them, and the correct fix is to route them to a person and stop.

If you take one thing away

The one thing
Track refusal rate and handovers by topic, the questions nothing matched, and how often a visitor asks the same thing twice, and ignore anything that only measures how much the widget was used.

Everything above is the reasoning. This is the part that changes what you do on Monday.

Questions

What is a good refusal rate?
There is no such number, and anybody quoting one is quoting a number from a different business with different material and different traffic. The useful comparison is against your own previous month, per topic. A refusal rate is only bad if the material should have covered the question.
Should we turn off ratings entirely?
No, but demote the score and keep the comments. A free text comment attached to a negative rating tells you what went wrong and is worth reading individually. The aggregate percentage is a low response rate mood reading and should not appear next to metrics you would act on.
How do we group unmatched questions without doing it by hand every week?
Start by hand, because the first few weeks teach you what the natural groups are, and there are usually fewer than you expect. Once the groups are stable you can sort new questions into them quickly. The reading is the part that produces the insight, so do not automate that away too early.

Keep reading

Try it on your own material

Upload a document or point it at your site, paste one line of HTML, then ask it something only your business could answer.