Language models are trained on public data up to a certain date. They know nothing about your internal policies, product manuals or customer history. Retrieval-augmented generation — RAG — closes that gap by fetching relevant information at the moment a question is asked.
How RAG works
- Ingest: documents are split into smaller chunks.
- Embed: each chunk is converted into a vector that captures its meaning.
- Store: the vectors are saved in a vector index or database.
- Retrieve: a user question is embedded and the most similar chunks are found.
- Generate: the model answers using only those retrieved chunks as context.
Why product teams like it
- Answers can cite the exact source document.
- Updating knowledge means re-indexing documents, not retraining a model.
- Access rules can be applied at retrieval time, so users only see what they are allowed to see.
Where RAG goes wrong
Poor chunking
Chunks that are too small lose context; chunks that are too large dilute relevance. Splitting along the natural structure of documents — headings, sections, table rows — usually works better than fixed character counts.
Retrieval misses
If the right chunk is never retrieved, the model cannot use it. Combining semantic search with keyword search (often called hybrid search) and adding a re-ranking step helps a lot.
Confident wrong answers
Instruct the model to say it does not know when the context does not contain the answer, and show sources so users can verify.
Evaluating a RAG system
Build a small set of real questions with known correct answers and sources. Test retrieval (did we fetch the right chunk?) separately from generation (did the model use it correctly?). This makes problems much easier to diagnose.
Most RAG quality problems are retrieval problems wearing a generation costume.
- RAG
- LLM
- Search


