Maciej Nuzia · RAG

RAG answers from company documents and shows its sources

RAG, short for Retrieval Augmented Generation, searches a selected document set before producing an answer. The response comes with source passages that a reader can open and check. I therefore start with search and organising the source material.

Not every set of documents suits RAG. Further down I describe the three places where it comes apart.

The problem

The knowledge is documented, but answers keep being written from scratch

The policy applies in its amended version, and how that works in practice is known to whoever has been here longest. When a question comes in, somebody puts those pieces together and writes the answer. Then the answer goes out by email to one customer and stays there.

  • the same answer has been written before, and nobody knows who wrote it
  • the intranet search returns a list of files to go through, and that is where its job ends
  • documents pile up faster than anyone can put them in order
  • in a dispute with a customer, what counts is a passage you can cite

Who it's for

RAG needs two things: documents and repeated questions

The questions have to return often enough for somebody to notice the system has taken them over. When a sample of the documents tells me that is not the case, I advise against the whole thing and say why. Both come together where documents are in daily use.

  • Customer service and support
  • Legal and compliance
  • HR and internal policies
  • Field service and maintenance
  • Product documentation
  • Sales working from price lists and specifications

Use cases

Where RAG helps in day-to-day work

A draft reply for the customer

A ticket lands, and next to it appears a suggested reply along with the passage of the procedure it came from. The agent fixes one sentence and sends it. On a harder case they open the passage and read it in full.

A question about a clause in a contract

You ask how the notice period works in this particular contract, and what returns is the paragraph where it is written. Across contracts that differ in small details, naming the right place can be the whole value.

Search that understands the question

Somebody asks about returning goods after the deadline, and the procedure calls that situation an out-of-warranty complaint. Matching on meaning stands a chance of putting those two names together.

The first weeks of a new hire

A new employee asks the system the things you only get to ask a colleague once.

Standards and internal rules

A body of text that shifts from one year to the next. The system answers from the version marked as current, which means somebody in the company has to mark it. A single decision that has to belong to a named person.

A chatbot facing customers

The same machinery, turned outwards. Here I am more careful with the limits: what the model may discuss, and which subjects it has to hand straight to a person.

RAG is sometimes one feature of a larger application. I write about that on the page about AI web applications.

How it is built

I start by organising the documents

The outcome is decided by the material I put into the model and the vector store. Hence the order below.

01

What goes into the store

I sit down with what the company keeps: internal policies, contracts, customer correspondence. Some of it drops out straight away, because it duplicates another file or has stopped applying. For each set I need to know who decides which version holds. Without that person, contradictions between documents stay in the store and only surface in the answers.

02

Splitting into passages, and the index

I cut documents into passages that make sense without the rest of the file, and turn them into vectors so they can be searched by meaning. The size of a passage, and what I carry into it from the headings, both come out of trials on your own files. A contract divides differently from the transcript of a call with a customer.

03

Search first, then the answer

Given a question, the system picks passages first and only then builds a sentence out of them. I set how many it takes, what it adds to the query, and what should happen when nothing fits. That last case needs an instruction of its own: what the system should say when the search turns up nothing.

04

A pass through your own questions

I collect the questions people actually ask, and go through them with you one by one: what the search returned, what the model made of it, and whether the source backs it up. Fixes go back onto that list. The documents get a refresh routine of their own.

Technologies

Technology for search and answer generation

I choose the vector store and the model to fit where your documents are allowed to be kept.

  • OpenAI API
  • Claude / Anthropic
  • pgvector
  • Embeddings
  • Python
  • PostgreSQL
  • RAG pipelines
  • LangChain / custom
  • AWS

The hard parts

Where an answer can go wrong

An answer with nothing underneath it

I attach the passages an answer leaned on, and on anything that matters they have to be opened. Retrieval misses sometimes, and the model will build a sentence out of whatever it is handed. A citation is no proof that the answer is right. It only proves what the answer was assembled from.

A document changed yesterday

I design the refresh separately: what feeds the store, how often, and who triggers it. Until the refresh happens the answers come from an older version of the document.

Questions the documents have no answer to

Part of what a company knows was never written down. RAG will not pull it out, because there is nothing to pull it from. Once the system is live you can see what people ask and what the set is missing. Usually that is when a few documents that never existed get written.

Of the three, the last one surprises people most. The list of questions the company has no answer to writes itself, and it is often more interesting than the system.

FAQ

Documents, sources and incorrect answers

What is RAG?

Short for Retrieval Augmented Generation. The model gets a question, but before it answers it searches a document set you point it at and pulls passages out of it. The answer is assembled from those passages and arrives alongside them, so you can see what it stood on.

How is RAG different from fine-tuning?

Fine-tuning changes the behavior of the model itself, and every change in knowledge means training it again. RAG leaves the model alone and gives it a search layer over your files, so fresh knowledge arrives by swapping a document. Where the knowledge shifts month to month, that difference settles it. Where the style and the format of an answer stay fixed, fine-tuning is often the better choice.

What happens when the documents hold no answer?

The system can be set up to say it found nothing and hand over a contact for a person. That instruction is part of the build, and I test it on questions the set is known to have no answer to.

Is company data secure?

The data goes where we settle it. The model can be called through a provider that does not train on what you send, or placed closer to your own systems if the documents cannot leave the company. Permissions are a separate matter: the system has to know who may ask about what, or the search will show an employee a clause from a contract they should not see. I raise this early, because the answer changes the architecture.

Documents

Show me the documents and questions that keep returning

A few sentences will do. Tell me what the documents are, what they are stored in, and who answers questions about them today. If you would rather begin with numbers, run the estimation tool first and bring its report with you.