Why this has to happen first
Human readers do a lot of quiet repair work. Faced with two pages giving different return windows, a person looks at which page seems more official, or more recent, or matches what they were told in the shop. They resolve the conflict using context the document does not contain.
Retrieval has none of that context. Given two contradictory passages, it will surface whichever one matches the question more closely, and that is essentially arbitrary with respect to which one is correct. Sometimes it will surface both, and then you get an answer that hedges between two numbers, which is worse than either.
So the contradiction that has been sitting harmlessly in your help centre for two years becomes an active liability the day you index it. Fixing first is not tidiness. It is the difference between a system that quotes you and a system that misquotes you.
Build the inventory
Start with a list of every document that would be indexed. Help centre pages, policy pages, product pages, any files you were planning to upload, the terms and conditions, the pages nobody has opened in three years but which are still published.
Put them in a spreadsheet with four columns: the URL or filename, the primary question it answers, the date it was last genuinely reviewed, and the person who owns it. The last two columns will be mostly empty, and the emptiness is the finding. A page with no owner and no review date is a page that will be wrong eventually and nobody will notice.
Expect the list to be longer than anyone predicted. Sites accumulate pages the way drawers accumulate cables. The inventory is often the first time anybody has seen the whole set.
Find the duplicates
Sort the inventory by the primary question column. Anywhere two rows carry the same question, you have a duplicate, and duplicates are where contradictions come from: two pages start identical, one gets updated, and the divergence is invisible because nobody reads both.
The most common duplicate pairs are a help centre article and a section of a longer policy page, a current page and an older page that was never unpublished, and a page on the main site duplicated in a separate support subdomain.
For each pair, pick one to survive. Redirect or delete the other. Do not merge them into a longer page, which is the instinct, because a longer page answers more questions and you have just spent the previous step splitting those apart.
Find the contradictions
Duplicates are easy because they are obvious. Contradictions between pages that are not duplicates are harder, and they cluster around a small set of facts that appear in many places.
Search your whole site for each of these, one at a time: your prices, your delivery timeframes, your return window, your opening hours, your minimum order, your notice period for cancellations, your response time commitment, and any figure that appears in marketing copy as well as in policy. Collect every occurrence with its page. Compare.
This exercise usually finds something uncomfortable, most often a marketing page carrying an older, more generous promise than the current policy page. Decide which is true, fix the other, and note that whichever one you did not fix was being read by customers up until this morning.
Find the stale numbers
Go through the inventory looking specifically for figures with a shelf life: prices, percentages you actually publish, dates, staff names, product availability, regulatory thresholds, and anything phrased as new or coming soon.
Check each one against the current truth. The ones you cannot verify in five minutes are the dangerous ones, because if you cannot establish whether a number is current, nobody maintaining the page can either.
For anything you cannot verify, the correct action is to remove the number rather than leave it. A page that says contact us for current pricing is worse than a page with the right price and much better than a page with the wrong one.
Find the questions with no page
Take the list of real top questions from your inbox, and check each one against the inventory. Anything on the question list with no matching row is a gap, and gaps are where the refusals will come from.
Write the missing pages before you index rather than after. It is the same work either way and doing it first means the launch week is spent reading real conversations rather than filling in obvious holes you already knew about.
Two categories deserve special attention. Questions where the answer is no, we do not do that, which are almost never documented because they feel negative and are enormously useful. And questions about how a process works rather than what the policy is, which live in people's heads and have often never been written down anywhere.
Then, and only then, index
By the end of this you should have a smaller set of documents than you started with, each answering a distinct question, each with an owner and a review date, and no unresolved contradictions among the facts that appear more than once.
That is also the point at which the audit pays off a second time. A support team that has just been through this exercise knows exactly what its material says, which means it can evaluate whether an answer is correct rather than merely plausible. Without it, everybody is guessing about their own documentation.