Read all of them
Not a sample. In the first week the volume is manageable for almost every business, and the value is in the outliers rather than in the average, so sampling removes exactly the conversations worth seeing.
Block out time for it rather than doing it in gaps. An hour a day for five days is enough for most small sites, and the reading gets faster as patterns emerge. What you are building is a mental model of what visitors actually want, and that only forms if you see the whole distribution.
Read the transcripts in full, including the parts where somebody was clearly messing about. Testing behaviour tells you what people expect the thing to be able to do, which is useful even when the question is not serious.
Start with the refusals
Sort so that everything the assistant declined to answer comes first. This is your content gap list, generated by real demand rather than by guessing, and it is the single most valuable output of the launch.
Sort those into three piles. Things you should have documented and did not, which become the writing queue in priority order by frequency. Things that are documented but were not found, which are a shape problem: the page exists, but it is written in a way that does not match how the question was asked. Things that should never be answered by an assistant at all, which are confirmation your boundary is working.
The second pile is the one people misread. A refusal on a question you know you have a page for does not mean the retrieval is broken. It usually means the page buries the answer, or titles it in internal vocabulary, or splits it across sections. The fix is on the page.
Then read the confident answers
This is the harder pass and the one that gets skipped, because an answer that looks fine reads as a success and you move on. Go slower on these.
For each answer, check it against the material yourself. Is it true. Is it complete, or has it stated a rule while dropping the condition attached to it. Does it apply to the person who asked, or has it quoted a policy that covers a different customer type.
Dropped conditions are the most common defect and the hardest to spot, because the answer is not false. Free delivery over a threshold, without the exclusion. A fourteen day return window, without the sale item carve out. Both quote your material accurately and both will produce an argument at some point.
Look for the questions you did not expect
Every launch produces a category nobody predicted, and it is usually the most interesting finding of the week.
Common examples: people asking about a competitor by name, people asking whether you serve their area when you thought that was obvious, people asking questions that reveal they are on the wrong site entirely, people asking about a product you discontinued years ago, and people asking about employment rather than about buying anything.
Each of these is information about a mismatch between who you think visits and who does. Some deserve new material. Some deserve a change to the site rather than to the assistant. And some, like the discontinued product, are telling you that something out there is still sending people to a page you forgot about.
Change the material, not the prompt
The strong temptation in week one is to fix bad answers by editing instructions: telling it to be more careful, to mention the exclusion, to be shorter. Resist it, at least until you have made a full pass at the material.
The reason is that instruction changes are global and unverifiable. You cannot test them against the specific case that prompted the change without waiting for that question to come back. They also interact: three careful instructions added in one week produce behaviour nobody predicted and nobody can attribute to a single change.
Material changes are local and checkable. If an answer dropped a condition, put the condition in the same sentence as the rule on the page, and the next answer will carry it because the passage now contains both. That fix survives, applies to every phrasing of the question, and improves the page for human readers at the same time.
What to do at the end of the week
You should finish with three artefacts. A prioritised list of pages to write, ordered by how many people asked. A list of existing pages to restructure, with a note on what was wrong with each. And a short list of things you learned about your visitors that have nothing to do with the assistant at all.
The third is often the one worth taking to a wider meeting. The first week of unfiltered questions is the closest thing most small businesses ever get to a customer research exercise, and it cost nothing beyond the reading.
Then set the cadence. Weekly full reads for the first month, weekly refusal reviews after that, and a monthly pass on answered questions. Once the reading stops entirely, the drift starts, and nobody notices until a customer points at something.