Metrics and terms

Reading the meter, and the shape of a month

Every plan below the top one comes with a number of replies, and there is a counter on screen showing how much of it is gone. What that counter counts is not obvious, the ceiling behaves in a way worth knowing before you reach it, and the shape of a month tells you more about whether the plan fits than the total ever will.

shown in the product

The count against the allowance is on screen throughout. The header carries the used figure over the limit, with a bar that turns to a warning colour as it approaches, titled with the count and the period. The usage screen shows the same thing as a meter labelled "Replies", hinted "this period", with the allowance beside it and an infinity sign on the unmetered tier. What is not shown anywhere is the split: the meter is one figure, and it does not say how many of those replies were answered from your material and how many were pleasantries.

What it means

A reply allowance is a count of assistant replies included in a plan for a period, and burn is the rate at which that count is consumed. Metering by replies rather than by conversations or by visitors is the choice that makes a bill legible: a reply is a discrete event, everybody agrees when one happened, and nobody has to argue about where a conversation ended. What it hides is that not every assistant reply is the same kind of thing. Some are produced by consulting your material and writing an answer. Some are returned from the cache without producing anything. Some are refusals decided before any of that begins. Some are a sentence of politeness answered from a list. The meter has one rule about which of those count, that rule is not the same as the rule the insights page uses, and almost every surprise about this number comes from the gap between the two.

How it is actually calculated

What spends one and what does not

A refusal does not spend a reply. The material was checked, nothing cleared the threshold, the fallback wording was returned and no model was involved, so the allowance is untouched. An assistant refusing all day costs nothing and consumes nothing.

An answer served from the cache does not spend one either. Somebody asked a question close enough in meaning to one already answered for that assistant, the previous answer came back, and nothing was produced.

Everything else does. That includes the ordinary case of a question answered from your material, and it also includes the short conversational replies: a greeting, a thank you, an answer to somebody asking what the assistant can do, and the reply to somebody asking for a person. Those are answered without touching your material and without producing anything, they add nothing to the spend figure, and each of them still takes one off the allowance. It is the single most common surprise in this whole area, and it is worth knowing before you compare the meter to anything.

Why the meter and the message count disagree

The messages figure on the insights page counts every assistant message over the last thirty rolling days, for one assistant, including the refusals and the cached answers. The meter counts only the replies that spend one, across the whole workspace, since the current period began.

Three differences at once, then: a different rule about what counts, a different scope, and a different window. Two figures that both look like a count of replies will not match, and neither is wrong. If you are reconciling them, start by noting that the meter is per workspace, so several assistants share one allowance and no screen splits it between them.

The period is the workspace's own rather than a calendar month. It runs from the date the current period began, and that date is set again when a plan change takes effect, which is why an upgrade does not inherit the usage that prompted it.

What the shape of a month tells you

A total tells you whether you fitted. A shape tells you whether the meter is the right size. Even consumption across the period means the allowance matches the traffic, and a plan chosen on the total will keep matching.

Consumption concentrated into a few days means the total is the wrong thing to size against. A launch, a newsletter, an outage on another channel or a seasonal peak can spend a fortnight's worth in an afternoon, and the plan that comfortably fits the month runs out on a Tuesday. The daily series of messages and refusals for an assistant is what shows this, and it is worth looking at before deciding a tier is sufficient.

A rising floor is the third pattern and the one that creeps. Each week is a little heavier than the last, nothing spikes, and the ceiling arrives one month without anything having happened. That is normally growth, which is the good version, and it still needs noticing in advance rather than at the moment the widget stops.

The ceiling, and what a visitor sees

When the count reaches the allowance the assistant stops answering. What appears in the panel is the offline message set on that assistant, whose wording by default says it is having trouble and to try again shortly. It does not mention a plan and it does not mention a limit.

That is the right behaviour for a whitelabel product and it puts an obligation on you. A visitor on a customer's site is going to read that sentence, so it should be a sentence you would be happy to have read. There is a second ceiling behind it, a spend limit, which stops answers in exactly the same way and shows exactly the same message, so the wording covers both.

The failure is total rather than degraded: the assistant does not answer more briefly or refuse more often, it stops. That is why the counter is in the header on every screen rather than tucked into a billing page.

How the number gets moved without anything improving

How the allowance stretches without anybody being better served

Raise the threshold. Refusals do not spend a reply, so an assistant that refuses more lasts longer on the same plan. The meter reads as thrift and the product is doing less, and there is nothing on the usage screen that would show you which of the two happened.

A repetitive week does the same thing without a decision. The cache absorbs the duplicates, none of them spends a reply, and the allowance stretches for reasons entirely to do with the traffic mix. The following week of varied questions burns it faster at identical volume.

The accidental version is sizing a plan on a quiet month. The total fitted, the shape was never looked at, and the first busy week takes the assistant offline on somebody's live site. Read the daily series before believing a comfortable total.

What to look at instead, or alongside

  • The daily series of messages and refusals, which shows the shape of consumption rather than the total that hides it.
  • The refusal share, since refusals are traffic that spent nothing from the allowance and helped nobody either.
  • The spend figure, which is the second ceiling and moves for different reasons than the reply count does.
  • The wording of the offline message, because it is what a visitor reads at either ceiling and it is the one part of this you control in advance.

Questions

Do refusals count against my allowance?
No, and neither do answers served from the cache. A greeting does, and so does somebody typing thanks, even though nothing was looked up and nothing was spent producing them. If your meter looks higher than your real question volume, that is usually where the difference is.
Is the allowance per assistant or per workspace?
Per workspace. Several assistants draw on one pool, and nothing splits the figure between them, so an agency running client sites should watch the total rather than assume each site has its own. The messages figure on an assistant's insights page is a different count over a different window and will not add up to it.
What do visitors see when it runs out?
The offline message set on that assistant, which by default says it is having trouble right now and to try again shortly. It gives no hint that a limit was reached, which is correct for a product sold under somebody else's name, and it means the wording deserves five minutes of thought while nothing is wrong.

Keep reading

Try it on your own material

Upload a document or point it at your site, paste one line of HTML, then ask it something only your business could answer.