Written, 14 April 2026

The content audit to run before you automate anything

Indexing does not improve your documentation. It exposes it. Every contradiction in your help centre that a human reader silently resolved by picking the newer looking page becomes an answer that could go either way, and you will not know which way it went until a customer tells you. The audit below takes a day for most small sites and it is the highest value day in the whole project.

Why this has to happen first

Human readers do a lot of quiet repair work. Faced with two pages giving different return windows, a person looks at which page seems more official, or more recent, or matches what they were told in the shop. They resolve the conflict using context the document does not contain.

Retrieval has none of that context. Given two contradictory passages, it will surface whichever one matches the question more closely, and that is essentially arbitrary with respect to which one is correct. Sometimes it will surface both, and then you get an answer that hedges between two numbers, which is worse than either.

So the contradiction that has been sitting harmlessly in your help centre for two years becomes an active liability the day you index it. Fixing first is not tidiness. It is the difference between a system that quotes you and a system that misquotes you.

Build the inventory

Start with a list of every document that would be indexed. Help centre pages, policy pages, product pages, any files you were planning to upload, the terms and conditions, the pages nobody has opened in three years but which are still published.

Put them in a spreadsheet with four columns: the URL or filename, the primary question it answers, the date it was last genuinely reviewed, and the person who owns it. The last two columns will be mostly empty, and the emptiness is the finding. A page with no owner and no review date is a page that will be wrong eventually and nobody will notice.

Expect the list to be longer than anyone predicted. Sites accumulate pages the way drawers accumulate cables. The inventory is often the first time anybody has seen the whole set.

Find the duplicates

Sort the inventory by the primary question column. Anywhere two rows carry the same question, you have a duplicate, and duplicates are where contradictions come from: two pages start identical, one gets updated, and the divergence is invisible because nobody reads both.

The most common duplicate pairs are a help centre article and a section of a longer policy page, a current page and an older page that was never unpublished, and a page on the main site duplicated in a separate support subdomain.

For each pair, pick one to survive. Redirect or delete the other. Do not merge them into a longer page, which is the instinct, because a longer page answers more questions and you have just spent the previous step splitting those apart.

Find the contradictions

Duplicates are easy because they are obvious. Contradictions between pages that are not duplicates are harder, and they cluster around a small set of facts that appear in many places.

Search your whole site for each of these, one at a time: your prices, your delivery timeframes, your return window, your opening hours, your minimum order, your notice period for cancellations, your response time commitment, and any figure that appears in marketing copy as well as in policy. Collect every occurrence with its page. Compare.

This exercise usually finds something uncomfortable, most often a marketing page carrying an older, more generous promise than the current policy page. Decide which is true, fix the other, and note that whichever one you did not fix was being read by customers up until this morning.

Find the stale numbers

Go through the inventory looking specifically for figures with a shelf life: prices, percentages you actually publish, dates, staff names, product availability, regulatory thresholds, and anything phrased as new or coming soon.

Check each one against the current truth. The ones you cannot verify in five minutes are the dangerous ones, because if you cannot establish whether a number is current, nobody maintaining the page can either.

For anything you cannot verify, the correct action is to remove the number rather than leave it. A page that says contact us for current pricing is worse than a page with the right price and much better than a page with the wrong one.

Find the questions with no page

Take the list of real top questions from your inbox, and check each one against the inventory. Anything on the question list with no matching row is a gap, and gaps are where the refusals will come from.

Write the missing pages before you index rather than after. It is the same work either way and doing it first means the launch week is spent reading real conversations rather than filling in obvious holes you already knew about.

Two categories deserve special attention. Questions where the answer is no, we do not do that, which are almost never documented because they feel negative and are enormously useful. And questions about how a process works rather than what the policy is, which live in people's heads and have often never been written down anywhere.

Then, and only then, index

By the end of this you should have a smaller set of documents than you started with, each answering a distinct question, each with an owner and a review date, and no unresolved contradictions among the facts that appear more than once.

That is also the point at which the audit pays off a second time. A support team that has just been through this exercise knows exactly what its material says, which means it can evaluate whether an answer is correct rather than merely plausible. Without it, everybody is guessing about their own documentation.

If you take one thing away

The one thing
Before indexing anything, search your entire site for each fact that appears in more than one place, and resolve every contradiction you find.

Everything above is the reasoning. This is the part that changes what you do on Monday.

Questions

Can we index everything and clean up afterwards?
You can, and you will spend the cleanup period fielding wrong answers you could have prevented. The specific risk is contradictions, because they produce answers that are confidently wrong rather than obviously missing, which means you find out from customers rather than from your own review.
What should we do with old pages we do not want to delete?
If they are historically useful but no longer true, keep them out of the indexed set and mark them clearly on the page itself. If nothing links to them and they answer nothing, delete them. An unmaintained page that stays published is a wrong answer waiting for an audience.
Who should own each page?
A named person, not a team. Team ownership means nobody reviews it. The owner does not have to write the page, only to confirm each review cycle that it is still true or say who should check.

Keep reading

Try it on your own material

Upload a document or point it at your site, paste one line of HTML, then ask it something only your business could answer.