Written, 12 May 2026

Drawing the boundary between your own sites

Almost no organisation has one website. There is the main site, and then there is the shop on another domain, the brand that came with an acquisition, the regional site in another language, the documentation, the careers pages, and something a department built four years ago that still ranks. Whoever sets up an assistant runs into a decision nobody framed for them: one of these, or several, and where the line goes.

The tidy version does not exist

The way this is usually described assumes a company with a website. The reality in most organisations above a certain age is a small estate: a primary site, a store on a separate platform because the store had to be on a separate platform, a second brand kept separate because its customers do not overlap, one or two regional sites, and a documentation area that is technically part of the product rather than part of the marketing.

These are not equally maintained and they are not equally owned. Different people update them, at different rates, with different approval processes, and in several cases nobody has read the older ones in a year. That variation is the real subject here, because the question of whether one assistant can serve all of them is mostly a question about whether their contents can safely be mixed.

The decision is made once, usually quickly, and it is expensive to reverse later, because by then there are conversations, settings and habits attached to whatever was chosen. It is worth twenty minutes before the first install rather than after the first complaint.

The honest case for sharing

There is a real argument for one assistant across everything and it is not only about cost. A visitor does not care about your domain structure. Somebody on the documentation site asking about billing is asking a reasonable question that the documentation cannot answer, and an assistant scoped to that site alone has to refuse something it could easily have handled.

Sharing also concentrates the maintenance. One corpus, one review cycle, one owner, one refusal message to keep current. Given how often the failure mode in this whole category is neglect, an arrangement that halves the number of things to neglect deserves a hearing.

So the default should be sharing, with a specific set of conditions that force a split. Those conditions are not about branding or tidiness. They are about whether two sites can make different statements about the same subject, and about who answers when a visitor asks for a person.

What a leak looks like from the visitor's side

A leak is not an error message. It is a fluent, well cited answer built from a document that belongs to a different part of your organisation, and the visitor has no way at all to know that.

The examples are dull and costly. The shop's fourteen day return window quoted to somebody buying a service under a contract with a different cancellation term. A regional price quoted in the wrong currency, or a price that is correct in one country and excludes tax that is included elsewhere. The acquired brand's guarantee attributed to the parent, or the parent's guarantee attributed to the acquisition, which is worse. A staff facing process page from a department site explaining the internal steps for something a customer just asked about.

In every case the visitor will hold you to what they were told, and they are not being unreasonable. Your site said it. The fact that a different one of your sites said it is an internal detail that does not survive first contact with a complaint, and in a dispute about a contract term it may not be much of a defence either.

The one hard rule

Where two sites make different commitments about the same subject, they must not share a corpus. Not because it is untidy, but because of how retrieval behaves in exactly that case.

Retrieval selects passages by similarity to the question. Two versions of the same policy for two parts of your business are maximally similar to each other, which means both are strong matches for the same question, and they differ precisely in the detail that decides the answer. This is the worst possible configuration: near identical documents that disagree on one number. The system cannot prefer the right one, because nothing in the question says which part of the business the visitor is in.

So the split is drawn along commitments rather than along brands. Two sites with the same terms, the same prices and the same team can usually share. Two sites with the same brand and different notice periods cannot, however much they look like one business from the outside.

Region is the hardest case and it is not close

Regional sites are the maximum collision configuration. The pages are often translations or near copies of each other, which makes them similar in wording. They differ on price, on tax treatment, on delivery times, on which products are available, on statutory rights and cancellation periods, and sometimes on which legal entity the customer is actually contracting with.

It is also the case where the consequence of a mistake is highest, because several of those differences are not commercial preferences but legal ones. A statutory cancellation right quoted from the wrong country is not a customer service problem.

Split by region, and duplicate the shared material into each rather than sharing it. The instinct to keep one copy of the pages that are genuinely identical is understandable and it reintroduces exactly the collision you split to avoid, because the identical documents pull questions towards a corpus that also contains the non identical ones. Duplication is the cost of the boundary, and it is a smaller cost than a wrong statutory answer.

How the boundary is actually drawn

Not with a clever filter. Filters fail in the direction of leaking, because the tag that says which site a document belongs to is metadata rather than meaning, and the moment somebody adds a document without the tag the boundary has a hole in it that nobody will find.

The boundary is drawn by running separate deployments. Each site gets its own material, its own threshold, its own refusal wording naming that site by name, its own handover destination, and its own owner. Askably is one script tag per site, so this is a matter of pasting a different tag on each rather than of building anything, and the same is true of most tools in this category.

The handover destination is the test that settles most arguments. If a message left on the shop goes to the person who packs orders and a message left on the services site goes to an account manager, those are two different operations that happen to share a logo, and they should be two assistants. If both go to the same inbox and the same person answers both, sharing is probably fine.

Keeping the cost of a split low

Splitting has real costs and pretending otherwise leads to a decision that gets reversed in six months. You now have several corpora to review, several refusal messages that will drift apart, several sets of settings, and duplicated documents that have to be updated in more than one place.

Three habits keep it manageable. Keep one master copy of any shared document in the place you already maintain it, and index from there into each deployment, so there is still only one thing to edit. Keep a single change checklist naming every deployment, so a change to a shared policy is a list to work through rather than a thing you remember about the site you were looking at. And put the same person's name on all of them, at least at first, because several assistants with several owners is how you end up with several conflicting refusal messages.

Review them together, in one sitting, rather than separately. The point of the sitting is not each individual corpus, it is the comparison: whether the answer to the same question is still what you intended on each site.

A test to run before you decide

Take your ten most asked questions, the real ones from your inbox rather than the ones from a meeting. For each, write the true answer for every site in your estate, in a row.

Any question where the row is identical everywhere is an argument for sharing. Any question where one cell differs decides the boundary on its own, and it usually differs on the questions asked most, because the questions asked most are about money, delivery and cancellation, which are the three things most likely to vary between parts of a business.

There is a second question worth asking beside it: would a visitor to one of these sites be surprised to be told something about another. Sometimes the answer is no, and cross referencing is a service. Somebody on the documentation asking about billing wants the billing answer. Sometimes it is yes, and telling a customer of one brand about another brand's products is a disclosure the brands were kept separate to prevent. That one is a commercial decision rather than a technical one, and it should be made by somebody with the authority to make it.

If you take one thing away

The one thing
Write your ten most asked questions down the page and every site across the top, fill in the true answer for each, and split wherever a row is not identical.

Everything above is the reasoning. This is the part that changes what you do on Monday.

Questions

Can we run one assistant and just tell it which site the visitor is on?
You can pass that information, and it helps with tone and routing. It does not reliably fix the retrieval problem, because the wrong document is still in the corpus and still a strong match. Separation of material is what removes the collision; context only reduces it.
What about a careers site or an investor page?
Keep them out entirely unless somebody has agreed to answer questions from them. Job applicants and investors ask questions with consequences a support corpus cannot handle, and an assistant that answers a question about hiring or results in your company's name is a risk nobody signed off.
Our subsidiary was acquired and still has its own terms. Same brand, one site design.
Different terms means different corpora, regardless of how unified the design is. The visual merge is the reason to be careful rather than a reason to relax: visitors cannot tell which entity they are dealing with, so the answer has to be right without their help.

Keep reading

Try it on your own material

Upload a document or point it at your site, paste one line of HTML, then ask it something only your business could answer.