SHORT ANSWER
Retrieval-augmented generation (RAG) is a pattern in which an AI application first searches your own documents or data for relevant passages, then hands them to a large language model as context. Answers rest on current, company-specific sources and can cite them, without retraining the model. Most quality problems come from the retrieval step, so test it on 100 to 300 real questions.
Retrieval-augmented generation (RAG) pairs a search step with a large language model (LLM). When a user asks a question, the system finds the most relevant content in company sources such as policies, manuals, tickets or contracts. It passes that content to the LLM with the question, and the model writes an answer based on it. The term comes from a 2020 research paper by Lewis and colleagues at Facebook AI Research.
How does a RAG pipeline work?
Ingest: collect and clean documents, then split them into chunks of a few hundred words.
Index: turn each chunk into an embedding and store it in a vector database, often next to a keyword index for hybrid search.
Retrieve: embed the user’s question, find the closest chunks and filter them by the user’s permissions.
Generate: the LLM gets the question plus the retrieved chunks and writes an answer with citations.
Evaluate: test answers for accuracy and grounding, as described in our LLM evaluation guide.
Why do enterprises use RAG?
General-purpose models don’t know your products, contracts or procedures, and their knowledge stops at a cut-off date. RAG lets an assistant use current internal information without sending your data into model training, and it can respect existing access rights. It sits behind most enterprise chatbots, support assistants and knowledge search tools, and it is a building block for AI agents.
RAG vs. fine-tuning vs. plain LLM
| Plain LLM | RAG | Fine-tuning |
|---|---|---|---|
Uses company knowledge | No | Yes, at query time | Partly, baked into weights |
Updating knowledge | Not possible | Re-index documents | Retrain the model |
Source citations | No | Yes | No |
Access control per user | No | Yes | No |
Typical cost | Lowest | Medium | Higher |
Where does RAG go wrong?
Most quality problems start in retrieval, not in the model. Badly split documents, missing metadata, stale content and ignored permissions produce wrong or unsafe answers. As a typical range, a first production assistant for one knowledge domain, with evaluation and access control, takes two to four months. Track progress by scoring answers on a fixed set of 100 to 300 real questions.
Plain vector search isn’t always enough. On an enterprise chatbot project, we first mapped the questions leadership actually asked, such as who owns the at-risk deliverables this sprint. Similarity search couldn’t follow the links between people, projects and timelines. We moved retrieval onto a knowledge graph in Neo4j. Managers now get cited answers in seconds instead of spending hours preparing for team check-ins.
How RUBICON helps with RAG
We design RAG and GraphRAG assistants with evaluation and access control built in from the start, not added after launch. See our AI agents services, or our architects can review your use case and data with you.
Related terms
Frequently asked questions
What is the difference between RAG and fine-tuning?
RAG gives the model relevant information at question time, so you update its knowledge by changing the documents. Fine-tuning changes the model's weights through extra training, which suits style, format or narrow specialised tasks. For company knowledge that changes often and must be cited, RAG is usually cheaper, faster to update and easier to govern.
Does RAG stop hallucinations?
It reduces them but doesn't remove them. If retrieval returns the wrong passages, or the model ignores them, answers can still be wrong. Good RAG systems pair strong retrieval with instructions to answer only from sources, citations, a refusal when nothing relevant turns up and regular evaluation against a test set of real questions.
What do I need to build a RAG system?
You need the source content, a pipeline to split and index it, a retrieval index such as a vector database or hybrid search engine, a large language model and an application layer with access control. For production, add evaluation, monitoring, logging and a process that keeps the index in step with your documents.
What is GraphRAG?
GraphRAG is a variant of RAG that retrieves from a knowledge graph as well as, or instead of, text chunks. It helps with questions that span many documents or follow relationships, such as which suppliers a plant closure affects. It takes more modelling effort, so teams usually choose it for specific, relationship-heavy use cases.
Related case study

Enterprise GraphRAG Chatbot with Neo4j | Case Study
How RUBICON's Two Layer Fixed Entity Architecture eliminated data bottlenecks for a multi team enterprise, delivering a conversational AI system that gives leadership instant project clarity, without hallucinations.
More resources
