Metrics and terms

The failure that looks exactly like success

In most settings a fabricated answer is embarrassing. In support it is a promise. A visitor told the wrong returns window will hold you to it, and reasonably so, because the assistant was on your site wearing your name. Nothing in this product detects these, and nothing in any product does. Catching them is a reading job.

not measured here

This product does not compute a hallucination rate and does not flag individual answers as fabricated. Nothing in the interface scores accuracy. What it does provide is the material for a manual check: numbered citations bounded to the passages the model was shown, and a threshold that stops weakly matched questions before a model is called.

What it means

A hallucination is a fluent, confident statement that is not supported by the material and is not true. The word is unhelpfully mystical for what happens in a support context, which is more specific and more expensive: a model asked about a policy that is not in its material produces the most plausible policy, because plausible is what it is built to produce, and plausible policies look exactly like real ones. The distinguishing feature is that the failure carries no signal. A retrieval failure announces itself as a refusal; a hallucination arrives with the same tone and the same confidence as every correct answer you have read. In a support setting it also converts into a liability, because a statement made by your assistant to your customer about your terms is something you will be held to.

How it is actually calculated

Why no automatic detection exists

Detecting one requires knowing the correct answer, which is the same problem as producing the correct answer. A checker able to reliably flag a wrong statement about your refund policy would have to know your refund policy better than the thing being checked.

Products that claim detection are usually checking something narrower and worth having: whether the answer is supported by the retrieved passages. That catches the invented claim and misses the answer that faithfully quotes a document which is itself out of date.

What reduces the rate here

The refusal threshold, which stops questions with no strong match from reaching a model at all. This is the single largest control and it is the one most often loosened for a better looking deflection figure.

Citations bounded to the passages shown, which make a claim traceable, and material treated as data rather than instructions, which stops a document from talking the assistant into a different behaviour.

How to actually find them

Sample. Take a fixed number of answered conversations a week, follow the citation, and check that the claim is in the passage. It is dull, it is bounded, and it is the only method that works.

Weight the sample towards consequences rather than volume. Ten checked answers about prices, returns and eligibility are worth several hundred checked answers about opening hours, and a sample drawn at random from traffic will be almost entirely the latter.

How the number gets moved without anything improving

How the apparent rate falls while the real one does not

Stop looking. There is no automatic count, so the observed rate is exactly the rate you sample for. A team that samples ten conversations a month finds fewer problems than a team that samples fifty, and reports a better result from a worse process.

Sample the easy topics. Drawing at random from traffic fills the sample with opening hours and delivery times, which are almost always right, and misses the small number of consequential questions where the failures live.

Loosen the threshold and the invisible failures increase while every visible number improves. Deflection rises, refusals fall, the unanswered list empties, and the answers that replaced those refusals are the thin ones most likely to be wrong. Every dial that flatters the dashboard pushes in this direction, which is worth remembering whenever a figure moves in a pleasing way for a reason nobody chose.

What to look at instead, or alongside

  • A weekly sample of answers on your highest consequence topics, checked against the cited passages by somebody who knows the subject.
  • The refusal threshold set to match the cost of being wrong on that material, rather than to produce a comfortable deflection figure.
  • A list of questions the assistant must never attempt, enforced by material rather than by hoping.
  • Negative ratings read individually, since a visitor who noticed a wrong answer sometimes says so.

Questions

Does a citation prove the answer is right?
It proves the claim came from somewhere and tells you where. The claim can still misread the passage, and the passage itself can be out of date. A citation converts checking from an afternoon into ten seconds, which is its whole value, and somebody still has to spend the ten seconds.
Is a stricter threshold the answer?
It is the main lever and it costs you refusals. Cautious answers only on a strong match, which is usually right for regulated or high consequence material. The judgement is about the cost of being wrong on that particular subject, not about a number you are trying to hit.
How many answers should we check?
A fixed number every week that you will actually do, weighted towards the topics where an error is expensive. A sustainable ten a week beats a heroic hundred once, because the value is in noticing a change rather than in the coverage.

Keep reading

Try it on your own material

Upload a document or point it at your site, paste one line of HTML, then ask it something only your business could answer.