Written, 14 July 2026

Checking answers when nobody has time to check them all

Every other metric on a support assistant assumes the answers are correct. None of them measure it. Correctness is the hardest thing to evaluate and the only thing that ultimately matters, and the reason it goes unmeasured is that the obvious method, reading everything, is impossible. There is a workable method that is not that.

Nobody reads them all, so stop pretending

The plan that gets written down is that somebody will review the answers. The plan that happens is that somebody reviews them for a fortnight after launch and then stops. This is not a discipline failure. Reviewing every answer costs roughly as much as answering every question, which defeats the point of the exercise.

So the real choice is not between checking everything and checking nothing. It is between an honest small sample, checked properly and repeatedly, and an imaginary comprehensive review that will not survive contact with a busy month. The small sample wins on every practical dimension, including how much you learn.

It also produces something the comprehensive plan never does: a number that can be compared across months. A rate, however roughly measured, gives you a direction. A one off audit gives you a snapshot and no trend.

A sample you can actually draw

Take a fixed number of answers on a fixed schedule, and pick the number by how long checking takes rather than by statistical ambition. Something you can finish in half an hour, weekly, sustained. A sample you keep drawing beats a larger one you draw twice.

Draw it in two parts. Half at random, which is the only part that tells you anything about the overall rate. Half deliberately from the risky end: the topics where a wrong answer costs the most, the newest material, and the answers where the match was weakest. The random half measures, the deliberate half finds.

Stratify by topic if your volume is uneven, which it always is. A pure random sample from a business where most questions are about opening hours will mostly check opening hours, which is the least valuable thing to check. Taking a couple from each significant topic keeps the exercise pointed at the material where being wrong matters.

What right means, precisely

Correctness is not one property and grading it as a single yes or no loses most of the information. Split it into three, each of which fails differently and is fixed differently.

Is it factually true, against the source. This is the check that catches drift between the answer and the material. Is it supported by the material it cited, which is a different question: an answer can be true and still not be supported by the passage it pointed at, and that is a failure even though the outcome was fine, because it means the mechanism is not working. And is it complete: does it leave out a condition, an exception or a caveat that changes what the visitor will do next. Incomplete answers are the most common failure and the least often recorded, because they read perfectly well.

Grade each answer on those three and record which failed. After a few weeks the pattern in the failures tells you what to fix. Mostly incompleteness points at material that buries its exceptions. Mostly unsupported points at a retrieval or threshold problem. Mostly false points at material that is out of date.

Why citations change the cost of checking

Checking an answer with no citation means finding the relevant source yourself. You read the answer, work out which of your documents should govern it, find the passage, and compare. That is several minutes per answer and it is why answer review dies.

Checking an answer that cites its source means opening the cited passage and reading two paragraphs. That is under a minute, and it is a different activity: verification rather than investigation. The difference in cost is the difference between a review that survives and one that does not.

This is the practical argument for grounded, cited answers that has nothing to do with trust or marketing. Every answer in Askably carries numbered citations back to the indexed passage it came from, which exists partly so the visitor can check it, but the larger operational benefit is that it makes routine internal checking cheap enough to actually do. A system whose answers cannot be traced to a source has quietly made its own quality unmeasurable.

The checker has to know the subject

This is the part most often got wrong, and it invalidates the whole exercise when it is. A plausible, well written, confidently phrased wrong answer is indistinguishable from a right one to somebody who does not already know the answer. Handing answer review to whoever has capacity produces a review that approves everything.

The checker needs to be able to say, without looking anything up, that a claim is wrong. That usually means a support agent, a specialist, or whoever wrote the source material. It does not mean a manager, a marketer, or the person who set up the tool, unless they happen to also be one of those.

There is a second reason to use somebody who knows the subject: they catch the incompleteness failures, which are invisible otherwise. Only somebody who knows there is an exception notices that the answer did not mention it. Somebody checking against the document alone will mark it correct, because it is correct as far as it goes.

Recording it so it accumulates

Keep a running record with one row per checked answer: the date, the question, which of the three checks failed, and the fix if there was one. Not a report. A list that grows.

The value is entirely in the accumulation. One session tells you about one session. Three months of rows tells you which topics fail repeatedly, whether the rate is moving, and whether the fixes you made actually held. That is the thing you can put in front of somebody who asks whether the assistant is working, and it is a far better answer than any dashboard number.

Do the checking after each significant change to the material as well as on the schedule. New content is where errors concentrate, both because it is untested and because it competes with existing pages for retrieval in ways nobody predicts. A short check the week after publishing catches the ones that would otherwise run for months.

If you take one thing away

The one thing
Check a small fixed sample every week, half random and half from the risky end, grade each answer as true, supported and complete, and make sure the person checking already knows the right answer.

Everything above is the reasoning. This is the part that changes what you do on Monday.

Questions

How big should the sample be?
Whatever a person can check properly in the time you have committed, drawn on the same schedule every time. The comparability between weeks matters more than the size, because you are looking for a direction rather than estimating a population.
Can the checking be automated?
Partly. A second system can flag answers that are not supported by the passages they cited, which catches one of the three failure types cheaply. It cannot reliably tell you an answer is incomplete or out of date, because both require knowing something that is not in the document, which is precisely why the human check has to know the subject.
What do we do with the errors we find?
Fix the material rather than the answer, almost always. An individual wrong answer is gone. The passage that produced it will produce it again tomorrow, and it is usually one edit away from being unambiguous.

Keep reading

Try it on your own material

Upload a document or point it at your site, paste one line of HTML, then ask it something only your business could answer.