Troubleshooting
Between a finished upload and a citation, four things have to go right
An accepted upload is not an indexed document, an indexed document is not a matched passage, and a matched passage is not a source under a reply. Each of those steps can end quietly, and each leaves something behind that you can look at. Working through them in order takes a couple of minutes and beats re-uploading the file, which is the usual first move and never the answer.
The symptom
The file is sitting in the list on the Knowledge tab and no answer has ever quoted it or listed it as a source.
What it usually is
In rough order of how often each one turns out to be the answer. Work down rather than across: each carries a way to tell whether it is yours before you change anything.
- 1
It was read and yielded nothing worth indexing
- Why
- Text is pulled out of the file first and split into passages second. When the extraction finds nothing, or the split produces nothing, the document is recorded as failed rather than stored, and there is nothing in the knowledge base to match against afterwards.
- How to confirm it is this one
- The badge on that source reads Failed, and the line under its name carries the reason. A file that extracted but would not split names itself followed by the words produced no chunks. A scanned document is more explicit and says no text was found, and that scanned documents need to be made searchable before upload.
- Fix
- Upload a version that contains real text rather than a picture of text. If the reason names something else, that sentence is the whole diagnosis and is worth reading literally.
- 2
It belongs to a different assistant
- Why
- A file is uploaded under one assistant and nothing is shared between them. Two assistants in one workspace have separate material, separate settings and separate keys, and neither can see the other's documents.
- How to confirm it is this one
- Open the Knowledge tab of each assistant in turn and look for the file by name. Then view source on the page you are testing from and read the data-key on the tag against the key on that assistant's Install tab.
- Fix
- Upload it under the assistant whose key is actually installed on the site you are testing. Moving it is not possible, so upload it again in the right place and delete the stray source.
- 3
It is indexed and nothing anybody asked came close to it
- Why
- Matching compares the question against the passages, and a gate refuses before the assistant is involved when the closest passage is below the threshold on the Behaviour tab. A file about a subject nobody has asked about in matching words is present, searchable and never reached, which looks identical to a file that was never stored.
- How to confirm it is this one
- Take a phrase that appears in that file and nowhere else in your material, and ask a question built around it. An answer that cites the file proves it is indexed and reachable and that the wording was the obstacle. A refusal instead puts the question on the Insights tab under Unanswered questions with a best match percentage you can read against your threshold.
- Fix
- Put the way visitors phrase it into the document, as a heading with the answer directly under it. That moves the match further than adding more prose around it does.
- 4
It is being used and never credited
- Why
- The chips under a reply are built only from the numbers the assistant wrote into its own sentences. An answer drawn from the file where no number was written carries no chips, and on screen that is indistinguishable from an answer that ignored the file entirely.
- How to confirm it is this one
- Read the reply against the file. Wording, figures or names that exist only in that document, with nothing listed underneath the reply, is this one. It happens more on short answers than long ones.
- Fix
- Nothing to configure. It is worth knowing before you conclude a document is not being used, because the evidence people reach for first is exactly the evidence this cause removes.
If none of those fit
The number to trust is the document count on the source card, which is the count of documents actually stored for that source. A file that indexed shows one. A source card showing Ready next to zero documents has finished and stored nothing, which sends you back to the first cause above with the error line as your evidence. There is no per document list in the dashboard, so that count is the whole picture you get from this screen.
Questions
- Does uploading the same file again help?
- No, and it costs you. Each upload becomes its own source with its own copy of the document, so you end up with two copies of the same material competing to be matched, and a knowledge base that is harder to reason about than the problem you started with.
- Is the original file kept after it is indexed?
- No. Once the text has been extracted and indexed the original is discarded, deliberately, so that we can tell you your uploads are not kept. Re-indexing works from the extracted text, so nothing depends on the original surviving.
- Which formats can be uploaded?
- The upload card on the Knowledge tab lists the extensions it accepts and the size cap for your plan, and both are read from the server rather than written on the page, so what it says there is what will actually be accepted.
Keep reading
- The crawl only indexed one page of my websiteA finished crawl with a single document means discovery found nothing to follow. Six causes, each with the file, setting or address that proves it.
- The crawler is skipping my docs subdomainOnly a leading www counts as the same site, so any other subdomain is a different host and gets dropped. The fix is a second source, not a setting.
- This page builds its content with JavaScriptFetching a page returns the document the server sent, not what a browser assembles afterwards. How to tell, and the three ways round it.
- What to feed itWhy a document produces wrong answers as written, one kind at a time.
- Everything that goes wrongSymptom, cause, how to confirm which one, and the fix.
Try it on your own material
Upload a document or point it at your site, paste one line of HTML, then ask it something only your business could answer.