Start from the decision, not the dashboard
Before adding a metric, write down the decision it would change. If you cannot name one, the metric is decoration, and decoration in a dashboard is worse than nothing because it consumes the attention that should go to the numbers that matter.
Support assistants generate a small set of real decisions. What to write next. What to fix in existing material. Where to loosen or tighten the refusal threshold. Which topics should stop being automated at all. Whether the thing is worth keeping. Four metrics cover those decisions between them, and everything below is chosen because it maps onto one of them.
The corollary is that a metric which cannot go in a bad direction is not a metric. If a number only ever goes up as you use the product more, it is measuring your usage, not its performance.
Refusal rate, broken down by topic
The overall refusal rate is nearly meaningless. It moves with traffic mix, with seasonality, and with whatever campaign sent people to the site this week. Broken down by topic it becomes the single most actionable number you have.
A topic with a high refusal rate is telling you one of two things, and they need different fixes. Either the material does not cover that topic, which is a writing job, or it covers it in language nobody uses, which is a rewriting job. You can usually tell which by reading a handful of the refused questions: if the vocabulary in the questions does not appear in your material at all, it is the second case.
Watch the direction rather than the level. A refusal rate that climbs in a topic you have not touched usually means the world moved: a policy changed, a product launched, a price went up, and the material is now behind. That signal arrives weeks before anybody would have noticed the drift by reading the pages.
Questions with nothing behind them
This is the highest value output of the whole system and most teams never look at it. Every question that retrieved nothing above threshold is a question a real visitor asked that your published material does not answer. That is a content backlog written by your customers, sorted by their priorities rather than yours.
Group these by meaning rather than by string, because the same gap arrives in a dozen phrasings and looking at raw text makes it appear to be a dozen problems. Grouped properly, a month of unmatched questions usually collapses into a short list of genuinely missing topics, and that list is what you write next.
It is also a check on your assumptions. Teams are consistently surprised by what appears here, because internal ideas about what customers want to know are formed from the questions that reached the team, and those are filtered by whatever the website already answers badly.
Handovers per topic
Total handovers is a workload number. Handovers per topic is a design number, and the two are frequently confused.
Read it two ways. A topic with more handovers than you expected is either genuinely too complex to automate, in which case accept it and make the handover good, or is failing for a fixable reason, in which case the conversations will show you which. A topic with fewer handovers than you expected is more interesting and more dangerous: it may mean the answers are working, or it may mean the assistant is answering confidently in a place where it should be stopping.
Cross that with the topic level refusal rate. High refusals and low handovers in the same topic is the worst combination on the board, because it means people are being turned away and not offered a route out. That pairing is where the handover design is broken rather than the content.
The same question asked twice in one conversation
When a visitor rephrases their question and asks again, they are telling you the first answer did not work. This is the most honest satisfaction signal available and nobody has to press a button to produce it.
It is also more informative than a rating, because the rephrasing itself tells you what was wrong. People rephrase towards what they actually meant. Reading the pair, first attempt and second attempt, shows you the gap between the words your material uses and the words your customers use, which is exactly the thing that determines whether retrieval works.
Count them as a rate, and read a sample every week. If the rate rises after you publish something new, you have probably introduced material that competes with existing pages and is pulling the wrong passage into answers.
The numbers that mislead
Raw conversation counts. This measures how many people found the widget, which is a function of where the launcher sits on the page and how much traffic you had. It moves when marketing runs a campaign. It has never once indicated that answers got better.
Satisfaction on a two button widget. The response rate is low and it is not low at random: people rate when they are delighted or furious, and almost never in between. What you get is a mood reading from an unrepresentative slice, presented as a percentage, which invites exactly the false precision that makes it dangerous. Read the comments attached to the ratings, which are genuinely useful, and treat the score itself as noise.
Anything called containment. The name refers to the share of conversations that ended without a human being involved, and it is trivially gamed by making it harder to reach a human being. It is worth its own article, and it has one.
Average response time from the assistant. It is fast. It will always be fast. It is a property of the software, not an achievement, and putting it on a dashboard next to real metrics gives it a status it has not earned.
A review rhythm that fits in an hour
Weekly, look at two things: the unmatched questions from the last seven days, grouped, and a sample of conversations that ended in a refusal. That is enough to produce one or two writing tasks, which is a sustainable rate.
Monthly, look at refusal rate and handover rate by topic, and compare them to last month. You are looking for direction, not level. Anything that moved noticeably deserves a read of the underlying conversations before you touch a setting.
Quarterly, ask the harder question: which topics should not be automated at all. This is the review that never happens because there is no dashboard tile for it, and it is the one that most improves an assistant after the first few months. Some topics get worse the more effort you put into them, and the correct fix is to route them to a person and stop.