Metrics and terms

What each answered message actually cost you

There is no tile that does this division for you, and it is the most straightforward useful arithmetic on the page. Both terms are there over the same window. The subtlety is entirely in what the denominator contains, because two of the things it counts cost nothing at all.

shown in the product

The page shows "Inference spend", hinted "last 30 days", and "Messages", hinted "answered", over the same window. Both are computed and displayed. The ratio between them is not shown as its own tile, so the division is yours to do, and the two free categories on that page tell you how to interpret it.

What it means

Cost per answer is the money spent producing answers divided by the number of answers produced. In a support context it is usually the sanest cost figure available, because it is stable against traffic: a busy month and a quiet month should produce similar figures unless something about the assistant changed. The complication specific to this system is that the denominator is not homogeneous. An answered message might have been produced by a model, served from the cache for nothing, or been a refusal returned for nothing, and the three have wildly different costs. Which denominator you pick decides whether the figure means the average cost of a visitor interaction or the average cost of a model call, and those are different questions with different uses.

How it is actually calculated

The arithmetic and where the terms come from

Take the spend figure and divide it by the message count, both over the same thirty day window. That gives you the average cost per answered message, including the free ones.

The spend figure covers model usage over the window. It accumulates only on messages that actually called a model, which is why the two free categories drag the average down without contributing to the numerator.

Choosing a denominator on purpose

All answered messages gives you the cost of serving a visitor turn on average, which is the number to use when forecasting a bill against expected traffic.

Only the messages that called a model gives you the cost of a genuine model call, which is the number that changes when you alter reply length or switch the material being retrieved. You can approximate it by removing the cached share and the refused share from the message count, using the figures on the same page.

The two can differ substantially on an assistant with a high cache rate. Quoting one while thinking about the other is the usual way a cost estimate ends up wrong by a factor.

What is not in the figure

This is inference spend, not the price of the product, and not the cost of a support person answering the conversations the assistant handed over. Do not present it as a cost per resolved contact, because nothing here knows whether anything was resolved.

There is no per topic or per document breakdown either. If you want to know which topics are expensive, the practical proxy is reply length: longer answers cost more, and reply length is a setting rather than an accident.

How the number gets moved without anything improving

How the average falls while nothing gets cheaper

Raise the refusal threshold. Refusals cost nothing and stay in the message count, so the average per answered message falls while your assistant answers fewer questions. The cost line improves because the product does less.

The same happens accidentally after a spike of repeated questions: the cache absorbs them, spend is flat, the message count rises, and the average drops for reasons that have nothing to do with efficiency.

Shortening replies is the legitimate lever and it has a real cost of its own. The reply length setting is a cap rather than a target, so shortening it does not truncate short answers, but it does cut off the long ones that needed the room. Read a few of those before deciding the saving was free.

What to look at instead, or alongside

  • The spend figure on its own, tracked monthly, which is the number that actually appears on a bill.
  • Spend divided by conversations rather than messages, if you are trying to price the assistant per visitor.
  • The cache share, which explains most of the month to month movement in the average.
  • Reply length, since it is the setting with the most direct effect on what a single model call costs.

Questions

Why is there no cost per conversation tile?
Both figures are on the page and the division is one keystroke, but a tile would imply an official denominator, and the honest position is that the right denominator depends on the question you are asking. Naming one would hide that choice rather than make it.
Do refusals and cached answers really cost nothing?
Neither calls a model, so neither adds to the spend figure. A refusal is decided before the model is reached, and a cached answer is returned from a previous one. They are free in the sense that matters here, which is that they do not appear in the numerator.
Can I compare this figure against another product's?
Not usefully. Every product in this category draws its denominator somewhere different and most do not say where. Compare your own figure to your own figure last month, and investigate the direction rather than the level.

Keep reading

Try it on your own material

Upload a document or point it at your site, paste one line of HTML, then ask it something only your business could answer.