Before an AI can read your documents, someone has to survive your documents
Most demos of “AI that reads your documents” start at the fun part. You type a question, and an answer comes back, sourced from a contract or an invoice. It looks like magic, and it sells well.
We spend most of our engineering time about three steps earlier than that — on the part nobody demos. Because before a model can answer anything, something has to get every one of a business’s documents into a place where they can be read, indexed and trusted. And a real business’s documents are a mess.
The documents don’t arrive politely
A typical mid-sized company we work with has years of files scattered across shared drives, email inboxes, accounting exports and a scanner in the corner that produces PDFs of wildly varying quality. Bank statements. Supplier invoices, many handwritten or photographed at an angle. Correspondence. The same document, saved four times under three names.
You cannot point a language model at that directly. First you have to ingest it: pull each file out of wherever it lives, work out what it is, convert it into something readable, and store it so nothing is lost and nothing is counted twice. This sounds like plumbing. It is plumbing. It is also where these projects quietly succeed or fail.
The requirements are boring, which is exactly why they matter
When you are moving tens of thousands of a client’s documents, “mostly works” is not a specification. A few things turn out to be non-negotiable:
It has to be resumable. Ingesting a large drive takes hours. Networks drop, permissions get revoked mid-run, a source rate-limits you. If a failure means starting from zero, you will never finish. Every importer we build can stop and pick up exactly where it left off — and it treats a revoked authorization as “pause and report,” not “wipe progress and panic.”
It cannot create duplicates. Pull the same inbox twice and a naive system ingests everything twice, and now your search results and your document counts are wrong. We use durable, shared record-keeping so that two importers running against overlapping sources still land each document exactly once.
Broken files cannot be silently dropped. Real scanners produce malformed PDFs and corrupt images that crash naive tools. The choice at that point is to skip the file — and quietly lose data — or to repair it. We render problem files into clean, lossless pages rather than dropping them, and we log every repair so the failure is visible, not swallowed.
Everything is audited. For every document, you should be able to answer: where did this come from, when did we import it, was it reviewed, and what did we do to it along the way. That provenance trail is what lets a client trust the system with their records in the first place — and it is what lets them stay in control, which is how we think all of this should work.
Why we care about the unglamorous half
None of this is the part that photographs well. There is no chatbot in it. But it is the difference between a document-AI project that a business can actually rely on and a demo that impresses in a meeting and falls over on real data in week two.
The clever retrieval and the language model on top are, increasingly, the commodity. Anyone can wire those up. What’s genuinely hard — and what determines whether the thing works on your documents, not a tidy sample set — is the ingestion layer underneath: resumable, deduplicated, self-repairing, and auditable end to end.
So when someone shows you an AI that reads your documents, the useful question isn’t “how good is the model.” It’s “what happens when you point it at ten years of our actual files.” That’s the question we build for.