Metrics and terms

The share of questions that got the refusal message

A fallback is the assistant saying it does not know, in wording the owner wrote, without calling a model at all. It is the most informative single event the system produces, because unlike an answer it is unambiguous: something was asked and nothing covered it. The rate is worth watching in both directions, which is the part people miss.

shown in the product

There is no tile called fallback rate. The insights page shows its complement, labelled "Deflection" and hinted "answered, not refused". The raw count of refusals is available through the interface, along with a day by day series of messages and refusals, but neither of those appears on that page.

What it means

The fallback rate is the count of assistant messages that returned the refusal message divided by the count of all assistant messages, over a window. Mechanically it is the complement of the deflection figure, so the two carry identical information and differ only in which direction feels like good news. It is worth naming separately anyway, because the two framings prompt different questions. Deflection invites you to ask how high it can go. Fallback invites you to ask what was in them, which is the question that leads somewhere. A fallback is not a failure of the model; the model was never called. It is a statement that nothing in the indexed material scored above the threshold, which is a fact about your documents rather than about the assistant.

How it is actually calculated

The formula and the flag behind it

Refused assistant messages divided by all assistant messages in the window. The refusal is recorded as a flag written at the moment the answer was produced, not guessed later by matching the text, so the count does not drift if the owner edits their fallback wording.

Refused messages carry no model and no cost, which is what makes the arithmetic downstream work: a refusal is distinguishable from a cached answer even though neither called a model.

What decides it, question by question

The retriever finds passages, the best score is compared against the threshold, and below it the fallback is returned before any model is involved. That ordering is the reason a refusal costs nothing and returns instantly.

So every fallback is one of three things: material that does not exist, material that exists but uses vocabulary the visitor did not, or a threshold set higher than the material can clear. Those need three different fixes, and the rate alone tells you which you have exactly never.

Window and denominator choices

The insights page uses the last thirty days and counts messages. A fallback rate computed per conversation would be a different number and usually a larger looking one, since a conversation only needs one refusal to be tainted.

With no assistant messages in the window the figure is zero, which reads as a perfect record and means no data. Check the message count before believing a flattering rate.

Why a very low rate deserves suspicion

Most people treat this as a number to minimise, and past a point that instinct is wrong. An assistant that never refuses is an assistant that answers questions your documents do not cover, which is the failure mode that produces confident nonsense.

A rate near zero on a modest set of documents means either your visitors ask an unusually narrow set of questions, or the threshold is loose enough that thin matches are being answered from. The second is far more common, and the way to tell is to read a sample of the weakest answers rather than to look at the rate again.

How the number gets moved without anything improving

How the refusal count falls without the material improving

Lower the threshold. Every question that scored between the old bar and the new one now reaches the model, and every one of those becomes an answer rather than a refusal. The rate falls the same afternoon and not one document changed.

Rewrite the fallback message into something that sounds like an answer and the count does not change at all, but the people reading transcripts will stop noticing them. This is the accidental version and it is worse than the deliberate one, because it hides the signal rather than moving it.

Indexing a very large general document has the same effect as lowering the threshold without anybody deciding to lower it. There is always something to match, so there is rarely a refusal, and the answers that replaced those refusals are built on the vaguest passages you own.

What to look at instead, or alongside

  • The unanswered questions list, which is the same events with their text attached and is the only version you can act on.
  • The refused questions grouped by meaning rather than by wording, since one gap arrives in a dozen phrasings.
  • The direction of the rate within one topic over two months, which catches material going stale after a price or policy change.
  • A read of the thinnest answered conversations, to check whether refusals fell because coverage improved or because the bar dropped.

Questions

Is a fallback the same as an escalation?
No. A fallback is the assistant declining to answer. What happens next depends on how the fallback message is worded and whether the visitor asks for a person. When a visitor does ask for one, the assistant offers a form taking a name, an email and a message, and those enquiries are stored and listed per assistant. The two events are separate and only one of them is counted.
Do refusals cost anything?
No. The threshold is checked before the model is called, so a refusal spends nothing and returns immediately. That is also why a high refusal rate never shows up as a cost problem, only as a coverage problem.
Should we raise the threshold if we are worried about accuracy?
It is the correct lever, and it has a visible price. Cautious at 0.5 answers only on a strong match and refuses more often. That trade is usually right on regulated or high consequence material and usually wrong on a marketing site, and the deciding evidence is a sample of your own thin answers rather than a rule of thumb.

Keep reading

Try it on your own material

Upload a document or point it at your site, paste one line of HTML, then ask it something only your business could answer.