Written, 19 May 2026

Fifteen conversations, the same five columns, every week

Reading conversations to fix specific things and reading conversations to find out how it is going are different jobs with different methods, and doing the first while believing you are doing the second is how teams end up with a confident impression that is mostly made of whatever they read last. The second job needs a sample, and a sample needs rules.

Triage and measurement are not the same reading

When you filter for the conversations that went wrong and work through them, you are doing triage. It is the right thing to do, it produces edits, and it is not a measurement of anything, because you deliberately selected the failures. Nothing about the size of the pile tells you what proportion of conversations it represents, and the pile will feel enormous whether the underlying rate is dreadful or excellent.

Measurement is the other job: finding out what typically happens, including in the conversations nobody flagged. It requires the opposite selection rule. You have to take conversations you have no particular reason to take, including boring ones, and especially including the ones you would never have chosen.

Confusing the two produces a specific and very common error. A team spends an hour a week in the failure pile, forms an impression from it, and reports that impression as the state of the assistant. The impression is drawn entirely from a sample they constructed to contain problems. Both activities are worth an hour. They should not be the same hour, and they should not produce the same sentence at the end.

Draw it before you look at it

The rule that makes a sample a sample is that the selection happens before you see what is in it. Decide the rule first, apply it mechanically, and read whatever comes out. The moment you skip one because it looks uninteresting, you have gone back to browsing.

Two rules work well and neither needs any tooling. Take every nth conversation from the period, with n chosen so you end up with the number you can actually read. Or take a contiguous block: everything between two times on one day, which has the advantage of preserving what a real stretch of traffic looks like, including the runs and the quiet hours.

Whichever you pick, write it down and use the same one every time, because comparability across weeks is the entire point. A sample drawn one way in March and another way in April tells you about your sampling method rather than about your assistant. And when your total volume is smaller than the sample you planned, read everything and say so, which is a better position to be in than most people realise.

The unit is the conversation, not the answer

This is the choice that determines what the exercise can tell you. If you grade individual replies, you learn about the replies. If you take the whole conversation as one object, you learn whether a person got what they came for, which is the thing you actually want to know and which no individual reply contains.

The distinction matters because the two often disagree. A conversation can contain three correct answers and still end with somebody who did not get what they wanted, because the correct answers were to adjacent questions. A conversation can contain a refusal, which grades badly as a reply, and still be a complete success, because the refusal was correct and the handover worked.

So read from the first message to the last, in order, and hold your judgement until the end. The question you are answering is the visitor's, not the assistant's: did this person leave with the thing they arrived for, and if not, what was in the way.

The record card

Write down the same fields on every conversation. Five is enough and five is the number people can sustain. The date and the page it started from. What the visitor wanted, in your own words rather than theirs. Whether they got it, recorded as got it, partly, or did not. What happened at the end: they stopped, they asked again, they took a handover, they were refused. And one line naming the single change that would have made this go better, or none if nothing would have.

The fields are the instrument, and their fixedness is what makes them one. Free notes feel more informative while you are writing them and are useless three weeks later, because nothing lines up and nothing can be counted. Five fixed columns can be counted, sorted and compared, which is the difference between a record and a diary.

Resist the temptation to add fields. Every extra column reduces the chance the exercise happens next week, and the marginal column is almost always something you could derive later from the conversation itself. If you find yourself wanting a sixth field repeatedly over a month, that is evidence, and then you can add it deliberately and start the comparison again from that point.

Why the default output is a feeling

Human memory of a reading session is dominated by two or three conversations: the most recent, the most annoying, and the one that confirmed something you already suspected. That is not a failure of discipline, it is how recall works, and it is why an hour of careful reading routinely produces a summary sentence that would have been produced without the reading.

The card is a defence against your own memory rather than a bureaucratic ritual. Once the fields are written down, the count exists independently of how the session felt, and it is common for the count to disagree with the feeling. A session that felt bad turns out to contain two bad conversations out of fifteen, both about the same missing page. A session that felt fine turns out to contain five where the visitor asked again immediately.

There is a second benefit that only appears later. A written card from six weeks ago is checkable in a way that a remembered impression is not. When somebody asks whether the assistant has got better, you can answer with what you wrote at the time rather than with what you now believe you thought then.

A fixed day, a fixed size, a stopping rule

Pick a day and keep it. The specific day does not matter and the fixedness does, because a habit attached to a slot survives and a habit attached to intention does not. Weekly is right for most sites. Fortnightly is fine for a quiet one. Monthly is usually too slow to catch a problem while it is still cheap to fix.

Fix the size too, at whatever you can genuinely read properly in the time you are willing to spend, and hold it constant even in weeks when there is more to look at. A varying sample size makes the counts incomparable and reintroduces exactly the selection effect the method exists to avoid.

Then have a stopping rule for the session, which is simply that you stop when the sample is done rather than when you are satisfied. Reading on because the last few were interesting is how a sample turns into a browse. If the interesting ones need more attention, note them and give them a separate half hour under the triage heading, where selecting for interest is the correct method.

Why the sample beats the dashboard for the first three months

Early on, the numbers are small, and small numbers move for reasons that have nothing to do with your assistant. A week with forty conversations and a week with sixty produce rates that look like a trend and are mostly noise. A sample read by a person does not pretend to be a trend, so it does not mislead you into acting on one.

More importantly, the categories are wrong at the start and you do not yet know how. Every counted metric depends on a definition of what is being counted, and those definitions were written before anybody had seen your traffic. What counts as a resolved conversation on your site, with your visitors and your material, is something you learn by reading, and until you have learned it, a number aggregating that definition is precise about the wrong thing.

And a dashboard cannot see the failure that matters most, which is a fluent, confident, well formed answer to a question the visitor did not ask. Nothing about that conversation looks unusual from the outside. No refusal, no repeat, no handover, an ordinary length. It is only visible to somebody who read it and understood what the person wanted, which is precisely the thing a sample is for.

When to promote the card into a number

The sample is not permanent. Its job is partly to tell you what is worth counting, and you will know it has done that job when the outcome column stops surprising you: the same handful of categories keep recurring, the boundaries between them stop being arguable, and you find yourself filling the card in without hesitating.

That is the point at which those categories can become counted metrics, because now they mean something specific on your site rather than something generic. The definitions are yours, you can explain them, and you know what a movement in them would imply. A dashboard built at that point is genuinely useful in a way that a dashboard configured on day one is not.

Keep reading a small sample afterwards anyway, at a reduced size. The counting will drift away from reality slowly and invisibly as the site changes and the visitors change, and a few conversations a week read by a person is the cheapest available check that the number still describes the thing it was built to describe.

If you take one thing away

The one thing
Fix a selection rule, a sample size and a day, record the same five fields on every conversation you draw, and keep the cards so this month can be compared with last month.

Everything above is the reasoning. This is the part that changes what you do on Monday.

Questions

Is fifteen the right number?
It is a number that fits in a sitting and is small enough to survive a busy week, which matters more than any statistical property at this scale. Pick whatever you can genuinely finish, then hold it constant. The consistency is doing the work, not the size.
Who should fill in the card?
Ideally the same person each time, because the judgements in the middle column are subjective and consistency between weeks is worth more than any individual judgement being right. If two people share it, spend twenty minutes once agreeing what partly means, and write that definition down next to the card.
What if the sample is nearly all one topic?
That is a finding, not a problem with the sample. Uneven traffic is normal, and a sample that reflects it is telling you where your visitors actually are. If you also want coverage of the rarer topics, read a second small deliberate set alongside it and keep the two sets separate, so the counted one stays a fair draw.

Keep reading

Try it on your own material

Upload a document or point it at your site, paste one line of HTML, then ask it something only your business could answer.