Metrics and terms

Two buttons, and the weight they cannot bear

Ratings are the most requested number in support software and among the least informative. The counts exist here. A percentage does not, anywhere in the product, and the reasoning for that is worth reading before you go looking for one.

recorded, not on the insights page

Thumbs up and thumbs down counts are available through the interface, and a visitor can rate a conversation with one or the other. Neither the counts nor any ratio derived from them appear on the insights page, and no satisfaction percentage is computed anywhere in the product.

What it means

Satisfaction, applied to automated replies, is an attempt to find out whether the person on the other end was helped, by asking them. The method has a structural problem that no amount of design fixes: the people who answer are not a sample of the people who were served. They are the delighted and the furious, in unknown proportions, and everybody in the middle closes the panel and gets on with their day. A score built from that population has the shape of a measurement and the content of a mood. It becomes useful in exactly one form, which is individual negative ratings read one at a time next to the conversation that produced them, because at that scale the unrepresentativeness stops mattering and the specific complaint is the point.

How it is actually calculated

What is collected

A visitor can rate a conversation with a thumbs up or a thumbs down, one or the other, at the level of the conversation rather than the individual message.

That granularity is worth noticing. A conversation with four good answers and one bad one gets a single verdict, and the verdict does not say which message provoked it. You get the sentiment without the location.

What is counted

Thumbs up and thumbs down counts are available through the interface over the window. They are counts, not a ratio, and they are not shown on the insights page.

You can of course divide one by the sum yourself. Before you do, note that both counts share a denominator you do not have: the number of people who saw the buttons and ignored them, which on any real assistant is nearly everybody.

Why no percentage is computed

A percentage would imply a representative sample and there is not one. Displaying a satisfaction score drawn from a very low response rate invites a precision the underlying data cannot support, and once it is on a screen somebody will put a target on it.

The counts on their own are harder to over read. A rise in thumbs down is a signal worth investigating; a score that drifted by three points is noise dressed as information.

How the number gets moved without anything improving

How the ratio improves without anybody being happier

Ask for the rating at a moment of relief rather than a moment of judgement. Prompting just after a successful answer and not after a refusal produces a better ratio out of the same assistant, and the choice looks like a small design decision rather than a distortion.

The accidental version is more common: a change in traffic mix. A month of visitors asking the things you document well produces more thumbs up than a month of visitors asking about the new product, with nothing else different.

The structural version needs no action at all. Negative raters are usually more motivated than positive ones, so an assistant that quietly improves can show a worsening ratio simply because the delighted population shrank faster than the furious one.

What to look at instead, or alongside

  • Individual negative ratings read next to their conversations, which is the form in which this data is genuinely useful.
  • Whether a visitor asked the same thing twice in one conversation, which is an honest dissatisfaction signal nobody has to press a button to produce.
  • The unanswered questions list, which records a specific failure rather than a feeling about one.
  • A hand checked sample of answers against the source material, if the question you actually have is whether the assistant is correct.

Questions

Can we get a satisfaction percentage out of this?
You can divide the two counts yourself, and the page will not do it for you on purpose. Whatever you produce that way rests on a response rate you cannot see, so if it has to appear in a report, put the raw counts next to it so the reader can see how thin the base is.
Why rate the conversation rather than each answer?
Rating each message collects more data and irritates more visitors, and the extra data is still drawn from the same unrepresentative population. Conversation level keeps the interruption to one, at the cost of not knowing which reply caused the verdict.
Is a thumbs down worth acting on individually?
Yes, and it is the only use of this data worth much. Somebody bothered to press a button, which means something specific went wrong. Read the conversation, work out what was missing, and you usually end up with a document to write rather than a number to worry about.

Keep reading

Try it on your own material

Upload a document or point it at your site, paste one line of HTML, then ask it something only your business could answer.