Metrics and terms

One word, two numbers, and the distance between them

Read this before you quote our percentage anywhere. The tile on the insights page is a narrow, checkable thing: it counts how often the assistant answered rather than refused. The word deflection, everywhere else in support, means something considerably grander. If you take one for the other you will report a result you did not get, and nobody will catch it for months.

shown in the product

The tile is labelled "Deflection" and hinted "answered, not refused". That hint is the accurate definition and it is deliberately printed under the number rather than hidden in documentation. Nothing on that page claims a contact was resolved, avoided, or removed from a queue, because none of that is knowable from inside a widget.

What it means

Start from what a support team means when they say a contact was deflected: somebody had a question, an automated layer dealt with it, and no person spent time on it. That is a claim about work removed from a queue, and proving it requires knowing what happened after the conversation ended. Our figure claims something much smaller. It is the share of assistant messages that were not the refusal message, which is to say the share of times the assistant had material it considered close enough and produced an answer from it. It says nothing about whether the answer was correct, whether the visitor was satisfied, or whether they emailed your team ninety seconds later. Both numbers are legitimate. Only one of them is ours, and it is the modest one.

How it is actually calculated

The formula, exactly

One minus the count of assistant messages that returned the fallback, divided by the count of all assistant messages, over the last thirty days. That is the whole calculation and there is nothing else in it.

Note what is in the denominator: assistant messages, not conversations. A single conversation with six answers and one refusal contributes six sevenths, not a whole conversation marked good or bad. This makes the figure weight talkative visitors more heavily than brief ones, which is a real property of the number rather than a defect, but it is not how most people picture it.

When there are no assistant messages in the window the figure is zero rather than a hundred percent. An empty denominator has no honest answer, and showing a perfect score for an assistant nobody has used would be the least honest of the available choices.

What counts as a refusal

A refusal is the fallback message, the wording the owner wrote, returned because the best retrieved passage scored below the match threshold. It is recorded as a flag on the message, so the count is exact rather than inferred from the text of the answer.

This matters because it excludes a whole category people assume is included. An answer that hedged, waffled, or told the visitor to check elsewhere in its own words is not a refusal by this definition. It reached the model, the model produced prose, and it counts as answered. The number cannot tell a confident correct answer from a confident wrong one, and it was never built to.

The window, and what is missing from it

Thirty rolling days, ending now. There is no comparison against the previous period on that page, no split by topic, and no way to see whether the figure moved after you published something.

So the tile gives you a level and not a direction, and a level on its own is close to uninterpretable. If you want the direction, write the figure down on the same day each month. Two minutes of clerical work turns a decorative number into a usable one, and a day by day series of messages and refusals is available through the interface if you want to be more thorough than that.

The setting hidden inside the number

The refusal decision happens before the model is called at all. The assistant compares the best retrieved passage against a threshold the owner chose, and below it returns the fallback without spending anything. Above it, the model is called and the message counts as answered.

The owner picks that threshold from three options: Cautious at 0.5, Balanced at 0.35 which is the default, and Willing at 0.22. That means part of this percentage is a preference rather than an observation. Two assistants with identical material and identical traffic will show different figures if their owners chose differently, and neither is more correct.

How the number gets moved without anything improving

How this number goes up while nothing improves

Move the threshold to Willing. Questions that would have been refused now reach the model, the model answers from weaker material, and the figure rises immediately. The interface is direct about the trade in its own wording: fewer refusals, and more thin answers. Nothing about the material changed. The number moved because the bar moved.

The accidental version is more common and harder to spot. Publish one long, general document that touches every topic loosely and match scores rise across the board, because there is now always something vaguely relevant to retrieve. Refusals fall, the figure improves, and the answers quietly get worse because the passage that won was general rather than right.

There is a third route that nobody intends: traffic mix. A week of visitors asking the three things your material covers well will produce a better figure than a week of visitors asking about the new product nobody has written up yet. The assistant did not change. The questions did.

What to look at instead, or alongside

  • The refusal share broken down by topic, which points at a specific document to write rather than at a mood.
  • The unanswered questions list, which is the same information expressed as a task list your visitors wrote for you.
  • A hand read sample of answered conversations, because this figure counts an answer and cannot grade one.
  • Total contacts across every channel over time, if what you actually want to know is whether work went away.
  • Your own figure last month, recorded by you, since the page shows a level and not a movement.

Questions

Can I quote our Deflection percentage to a manager who asked about deflection?
Not without the sentence that goes with it. Say it is the share of assistant messages that were answered rather than refused, over thirty days. If you say deflection and stop there, you will be heard as claiming contacts were removed from a queue, and that is a claim this number cannot support.
Why is it not measured per conversation?
Because the refusal is a property of a message, and a conversation containing one refusal and five good answers is not obviously a failure or a success. Counting messages avoids inventing a rule for that, at the cost of weighting long conversations more heavily. Both choices distort something; this one distorts visibly.
What is a good figure?
There is no such number and anyone offering you one is quoting a different business with different material and different visitors. The only comparison worth anything is against your own previous month, and even then the question to ask is which topics moved, not whether the total went up.

Keep reading

Try it on your own material

Upload a document or point it at your site, paste one line of HTML, then ask it something only your business could answer.