A draft reply for the customer
A ticket lands, and next to it appears a suggested reply along with the passage of the procedure it came from. The agent fixes one sentence and sends it. On a harder case they open the passage and read it in full.
Maciej Nuzia · RAG
RAG, short for Retrieval Augmented Generation, searches a selected document set before producing an answer. The response comes with source passages that a reader can open and check. I therefore start with search and organising the source material.
Not every set of documents suits RAG. Further down I describe the three places where it comes apart.
The problem
The policy applies in its amended version, and how that works in practice is known to whoever has been here longest. When a question comes in, somebody puts those pieces together and writes the answer. Then the answer goes out by email to one customer and stays there.
Who it's for
The questions have to return often enough for somebody to notice the system has taken them over. When a sample of the documents tells me that is not the case, I advise against the whole thing and say why. Both come together where documents are in daily use.
Use cases
A ticket lands, and next to it appears a suggested reply along with the passage of the procedure it came from. The agent fixes one sentence and sends it. On a harder case they open the passage and read it in full.
You ask how the notice period works in this particular contract, and what returns is the paragraph where it is written. Across contracts that differ in small details, naming the right place can be the whole value.
Somebody asks about returning goods after the deadline, and the procedure calls that situation an out-of-warranty complaint. Matching on meaning stands a chance of putting those two names together.
A new employee asks the system the things you only get to ask a colleague once.
A body of text that shifts from one year to the next. The system answers from the version marked as current, which means somebody in the company has to mark it. A single decision that has to belong to a named person.
The same machinery, turned outwards. Here I am more careful with the limits: what the model may discuss, and which subjects it has to hand straight to a person.
RAG is sometimes one feature of a larger application. I write about that on the page about AI web applications.
How it is built
The outcome is decided by the material I put into the model and the vector store. Hence the order below.
01
I sit down with what the company keeps: internal policies, contracts, customer correspondence. Some of it drops out straight away, because it duplicates another file or has stopped applying. For each set I need to know who decides which version holds. Without that person, contradictions between documents stay in the store and only surface in the answers.
02
I cut documents into passages that make sense without the rest of the file, and turn them into vectors so they can be searched by meaning. The size of a passage, and what I carry into it from the headings, both come out of trials on your own files. A contract divides differently from the transcript of a call with a customer.
03
Given a question, the system picks passages first and only then builds a sentence out of them. I set how many it takes, what it adds to the query, and what should happen when nothing fits. That last case needs an instruction of its own: what the system should say when the search turns up nothing.
04
I collect the questions people actually ask, and go through them with you one by one: what the search returned, what the model made of it, and whether the source backs it up. Fixes go back onto that list. The documents get a refresh routine of their own.
Technologies
I choose the vector store and the model to fit where your documents are allowed to be kept.
The hard parts
I attach the passages an answer leaned on, and on anything that matters they have to be opened. Retrieval misses sometimes, and the model will build a sentence out of whatever it is handed. A citation is no proof that the answer is right. It only proves what the answer was assembled from.
I design the refresh separately: what feeds the store, how often, and who triggers it. Until the refresh happens the answers come from an older version of the document.
Part of what a company knows was never written down. RAG will not pull it out, because there is nothing to pull it from. Once the system is live you can see what people ask and what the set is missing. Usually that is when a few documents that never existed get written.
Of the three, the last one surprises people most. The list of questions the company has no answer to writes itself, and it is often more interesting than the system.
FAQ
Short for Retrieval Augmented Generation. The model gets a question, but before it answers it searches a document set you point it at and pulls passages out of it. The answer is assembled from those passages and arrives alongside them, so you can see what it stood on.
Fine-tuning changes the behavior of the model itself, and every change in knowledge means training it again. RAG leaves the model alone and gives it a search layer over your files, so fresh knowledge arrives by swapping a document. Where the knowledge shifts month to month, that difference settles it. Where the style and the format of an answer stay fixed, fine-tuning is often the better choice.
The system can be set up to say it found nothing and hand over a contact for a person. That instruction is part of the build, and I test it on questions the set is known to have no answer to.
The data goes where we settle it. The model can be called through a provider that does not train on what you send, or placed closer to your own systems if the documents cannot leave the company. Permissions are a separate matter: the system has to know who may ask about what, or the search will show an employee a clause from a contract they should not see. I raise this early, because the answer changes the architecture.
Documents
A few sentences will do. Tell me what the documents are, what they are stored in, and who answers questions about them today. If you would rather begin with numbers, run the estimation tool first and bring its report with you.