Triage and measurement are not the same reading
When you filter for the conversations that went wrong and work through them, you are doing triage. It is the right thing to do, it produces edits, and it is not a measurement of anything, because you deliberately selected the failures. Nothing about the size of the pile tells you what proportion of conversations it represents, and the pile will feel enormous whether the underlying rate is dreadful or excellent.
Measurement is the other job: finding out what typically happens, including in the conversations nobody flagged. It requires the opposite selection rule. You have to take conversations you have no particular reason to take, including boring ones, and especially including the ones you would never have chosen.
Confusing the two produces a specific and very common error. A team spends an hour a week in the failure pile, forms an impression from it, and reports that impression as the state of the assistant. The impression is drawn entirely from a sample they constructed to contain problems. Both activities are worth an hour. They should not be the same hour, and they should not produce the same sentence at the end.
Draw it before you look at it
The rule that makes a sample a sample is that the selection happens before you see what is in it. Decide the rule first, apply it mechanically, and read whatever comes out. The moment you skip one because it looks uninteresting, you have gone back to browsing.
Two rules work well and neither needs any tooling. Take every nth conversation from the period, with n chosen so you end up with the number you can actually read. Or take a contiguous block: everything between two times on one day, which has the advantage of preserving what a real stretch of traffic looks like, including the runs and the quiet hours.
Whichever you pick, write it down and use the same one every time, because comparability across weeks is the entire point. A sample drawn one way in March and another way in April tells you about your sampling method rather than about your assistant. And when your total volume is smaller than the sample you planned, read everything and say so, which is a better position to be in than most people realise.
The unit is the conversation, not the answer
This is the choice that determines what the exercise can tell you. If you grade individual replies, you learn about the replies. If you take the whole conversation as one object, you learn whether a person got what they came for, which is the thing you actually want to know and which no individual reply contains.
The distinction matters because the two often disagree. A conversation can contain three correct answers and still end with somebody who did not get what they wanted, because the correct answers were to adjacent questions. A conversation can contain a refusal, which grades badly as a reply, and still be a complete success, because the refusal was correct and the handover worked.
So read from the first message to the last, in order, and hold your judgement until the end. The question you are answering is the visitor's, not the assistant's: did this person leave with the thing they arrived for, and if not, what was in the way.
The record card
Write down the same fields on every conversation. Five is enough and five is the number people can sustain. The date and the page it started from. What the visitor wanted, in your own words rather than theirs. Whether they got it, recorded as got it, partly, or did not. What happened at the end: they stopped, they asked again, they took a handover, they were refused. And one line naming the single change that would have made this go better, or none if nothing would have.
The fields are the instrument, and their fixedness is what makes them one. Free notes feel more informative while you are writing them and are useless three weeks later, because nothing lines up and nothing can be counted. Five fixed columns can be counted, sorted and compared, which is the difference between a record and a diary.
Resist the temptation to add fields. Every extra column reduces the chance the exercise happens next week, and the marginal column is almost always something you could derive later from the conversation itself. If you find yourself wanting a sixth field repeatedly over a month, that is evidence, and then you can add it deliberately and start the comparison again from that point.
Why the default output is a feeling
Human memory of a reading session is dominated by two or three conversations: the most recent, the most annoying, and the one that confirmed something you already suspected. That is not a failure of discipline, it is how recall works, and it is why an hour of careful reading routinely produces a summary sentence that would have been produced without the reading.
The card is a defence against your own memory rather than a bureaucratic ritual. Once the fields are written down, the count exists independently of how the session felt, and it is common for the count to disagree with the feeling. A session that felt bad turns out to contain two bad conversations out of fifteen, both about the same missing page. A session that felt fine turns out to contain five where the visitor asked again immediately.
There is a second benefit that only appears later. A written card from six weeks ago is checkable in a way that a remembered impression is not. When somebody asks whether the assistant has got better, you can answer with what you wrote at the time rather than with what you now believe you thought then.
A fixed day, a fixed size, a stopping rule
Pick a day and keep it. The specific day does not matter and the fixedness does, because a habit attached to a slot survives and a habit attached to intention does not. Weekly is right for most sites. Fortnightly is fine for a quiet one. Monthly is usually too slow to catch a problem while it is still cheap to fix.
Fix the size too, at whatever you can genuinely read properly in the time you are willing to spend, and hold it constant even in weeks when there is more to look at. A varying sample size makes the counts incomparable and reintroduces exactly the selection effect the method exists to avoid.
Then have a stopping rule for the session, which is simply that you stop when the sample is done rather than when you are satisfied. Reading on because the last few were interesting is how a sample turns into a browse. If the interesting ones need more attention, note them and give them a separate half hour under the triage heading, where selecting for interest is the correct method.
Why the sample beats the dashboard for the first three months
Early on, the numbers are small, and small numbers move for reasons that have nothing to do with your assistant. A week with forty conversations and a week with sixty produce rates that look like a trend and are mostly noise. A sample read by a person does not pretend to be a trend, so it does not mislead you into acting on one.
More importantly, the categories are wrong at the start and you do not yet know how. Every counted metric depends on a definition of what is being counted, and those definitions were written before anybody had seen your traffic. What counts as a resolved conversation on your site, with your visitors and your material, is something you learn by reading, and until you have learned it, a number aggregating that definition is precise about the wrong thing.
And a dashboard cannot see the failure that matters most, which is a fluent, confident, well formed answer to a question the visitor did not ask. Nothing about that conversation looks unusual from the outside. No refusal, no repeat, no handover, an ordinary length. It is only visible to somebody who read it and understood what the person wanted, which is precisely the thing a sample is for.
When to promote the card into a number
The sample is not permanent. Its job is partly to tell you what is worth counting, and you will know it has done that job when the outcome column stops surprising you: the same handful of categories keep recurring, the boundaries between them stop being arguable, and you find yourself filling the card in without hesitating.
That is the point at which those categories can become counted metrics, because now they mean something specific on your site rather than something generic. The definitions are yours, you can explain them, and you know what a movement in them would imply. A dashboard built at that point is genuinely useful in a way that a dashboard configured on day one is not.
Keep reading a small sample afterwards anyway, at a reduced size. The counting will drift away from reality slowly and invisibly as the site changes and the visitors change, and a few conversations a week read by a person is the cheapest available check that the number still describes the thing it was built to describe.