Metrics and terms

The single most consequential setting you own

Almost every complaint about an assistant being too cautious or too talkative is a complaint about this one number. It sits before the model, it is chosen by the owner, and it silently sets the level of half the figures on the insights page. It deserves five minutes of your attention roughly once a quarter.

shown in the product

The threshold is a setting rather than a reported figure, and its three options are described in the interface as answering only on a strong match, answering when the material plausibly covers the question, and answering from weaker matches with fewer refusals and more thin answers. The values behind those words are exposed too: 0.5, 0.35 as the default, and 0.22. The score a given question achieved is displayed only for refused questions, as a "best match" percentage.

What it means

A confidence threshold is the minimum match score a question must reach before an answer is attempted. Calling it confidence is a slight misnomer worth keeping in mind: it is not the model's confidence in an answer, because no model has been consulted yet. It is a measure of how close the nearest piece of your material sits to the question. The threshold turns that continuous score into a binary decision, and every binary decision drawn on a continuum makes two kinds of mistake. Set it high and you refuse questions you could have answered. Set it low and you answer questions you should have refused. There is no setting that avoids both, and choosing well means deciding which mistake is more expensive on your material rather than looking for a correct number.

How it is actually calculated

Where the check happens

Retrieval runs, the best passage's score is compared against the threshold, and below it the assistant returns the fallback wording the owner wrote without calling a model at all. Above it the model is called with the retrieved passages.

That ordering has two consequences people find surprising. A refusal costs nothing and returns instantly, and no amount of prompt wording can rescue a question that failed the check, because nothing was ever asked.

The three settings, in the product's own words

Cautious, at 0.5: answers only on a strong match, says it does not know more often, and is wrong less often. Balanced, at 0.35, is the default: it answers when the material plausibly covers the question. Willing, at 0.22: it answers from weaker matches too, with fewer refusals and more thin answers.

The values are absolute cut offs rather than percentiles, so the same setting behaves differently on different material. Dense specific documents clear a high bar for genuine matches; sprawling general documents produce middling scores for everything, and a middling score for everything means a high bar catches nothing useful.

What it moves elsewhere

Nearly everything on the insights page. Deflection, the refusal share, the unanswered list, the latency figure and the average cost per answered message all shift when this dial moves, in the flattering direction as it loosens.

That is the trap in one sentence: the setting that makes the dashboard look better is the setting that makes thin answers more likely. If somebody reports an improvement across every tile at once, ask when the threshold last changed before you believe it.

How the number gets moved without anything improving

The dial that moves every other number at once

Loosening it improves deflection, empties the unanswered list, lowers the latency figure and reduces the average cost per answered message, all on the same afternoon and without a single document changing. Every one of those improvements is the same event counted five ways.

The accidental version is the one to watch for: somebody nudges it during a busy week to reduce refusals, the tiles improve, and nobody connects the two. Three months later the answers are thinner and the dashboard has never looked better.

The honest use of the same lever runs the other way. Tightening it makes several numbers look worse and makes the remaining answers better, which is a difficult thing to present and is often the right call on material with real consequences.

What to look at instead, or alongside

  • The best match percentages on refused questions, which tell you whether you are refusing things that nearly matched.
  • A read of the thinnest answers at your current setting, which is the only direct evidence of what a looser bar costs.
  • The cost of being wrong on your particular material, which is the actual input to this decision.
  • One step at a time, with a week in between, since changing it alongside a content update makes both effects unreadable.

Questions

Which setting should we use?
Balanced unless you have a reason, and the reason is usually about consequences. Material where a wrong answer costs money, safety or a legal position argues for Cautious. A marketing site where the worst case is a mildly unhelpful reply argues the other way. Nothing about the number itself decides this.
Can I set a value between the presets?
The three presets are the supported choices in the interface and the underlying values are shown rather than hidden, precisely because somebody will eventually want the number instead of the adjective. A value that is not one of the three is displayed as custom rather than snapped to the nearest preset, so what is saved is what you see.
Why does the same setting behave differently on two assistants?
Because the score depends on your material. Short focused passages score high for the questions they answer and low for everything else, which is what makes a threshold discriminating. Long general documents score in the middle for everything, and no threshold separates that well.

Keep reading

Try it on your own material

Upload a document or point it at your site, paste one line of HTML, then ask it something only your business could answer.