What is RAG?An AI that checksbefore it answers
Retrieval augmented generation makes a language model search your own documents before it replies. How the idea works, where it earns its keep and what it still gets wrong.
Get in touch- What is RAG in AI, and what does your business get out of it?
- Retrieval augmented generation: why researchers wanted answers you can trace
- How does RAG work, so the AI finds the right passage?
- RAG vs fine-tuning: which route gets company knowledge into an AI?
- When does RAG pay off for company knowledge?
- RAG hallucinations: where you still need to check the answer
- What to watch for with RAG and data privacy when documents are confidential
- Frequently asked questions
- How to check whether RAG fits your company knowledge
- Where the information on this page comes from

Let's talk about your project.
First we check whether the project fits your business model. Then you get a proposal with phases and effort.
Talk about a RAG project or call: +49 151 1576 5566RAG stands for retrieval augmented generation, and the short version is an AI that looks things up before it answers. Before the language model writes a single sentence, a search step digs through a collection you control (service manuals, price lists, contracts, the help desk wiki) and hands the model the handful of passages that fit the question. The term comes from a 2020 paper by Patrick Lewis and colleagues, accepted at NeurIPS that year. For a business the appeal is twofold: the AI can work with knowledge it never saw in training, without anyone retraining the model, and every claim it makes can be checked against the passage it came from. That does not make it infallible, and we will get to why.
- With RAG, a search step pulls matching passages out of your documents first, and only then does the language model write its reply.
- The term was coined in 2020 by researchers at Facebook AI Research, University College London and New York University.
- Updated documents take effect once they are indexed, no retraining required, and the answer can point to the paragraph it relied on.
- It is not foolproof: in a Stanford study, specialised RAG tools for lawyers got 17 percent or more of the test queries wrong.
On this page
- What is RAG in AI, and what does your business get out of it?
- Retrieval augmented generation: why researchers wanted answers you can trace
- How does RAG work, so the AI finds the right passage?
- RAG vs fine-tuning: which route gets company knowledge into an AI?
- When does RAG pay off for company knowledge?
- RAG hallucinations: where you still need to check the answer
- What to watch for with RAG and data privacy when documents are confidential
- Frequently asked questions
- How to check whether RAG fits your company knowledge
- Where the information on this page comes from
What is RAG in AI, and what does your business get out of it?
What is RAG in AI, minus the jargon? A large language model (LLM) only knows what was in its training data. Your parts list from last quarter was certainly not in there, so when someone asks about it the model can either admit it does not know or, more often than anyone would like, assemble an answer that sounds right and is not.
RAG adds a detour. Say a field technician types on a tablet: “What torque applies to the bearing bolts on series B?” Before the model replies, a second program searches the service manuals, finds two or three relevant paragraphs and passes them along with the question. The model then writes from those paragraphs, a bit like an open-book exam where you are allowed to check the textbook before you start writing.
Two things change in practice. The AI can use material it was never trained on. And it can say which passage it leaned on, so the technician can read it themselves if in doubt (which, with a torque value, they probably should anyway).
Retrieval augmented generation: why researchers wanted answers you can trace
Retrieval augmented generation is a research term, and it is easy to date. On 22 May 2020 Patrick Lewis and eleven co-authors posted “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks” on arXiv; it was later accepted at NeurIPS 2020, one of the major machine learning conferences. The group came from Facebook AI Research (now part of Meta), University College London and New York University.
What bothered them is spelled out in the paper's opening lines: language models store a remarkable amount of factual knowledge, yet they are poor at retrieving it precisely, cannot show where a statement came from and are hard to update. Their answer paired a pre-trained text generator with a searchable index of Wikipedia, queried by a separate neural retriever. By their measurements the result was more specific, more diverse and more factual than a comparable model without retrieval.
When Meta wrote about the work in September 2020, it stressed the point that probably still matters most to companies. You change what such a system knows by swapping or extending the document collection, while the model itself stays exactly as it is.
is the year Patrick Lewis and colleagues published the paper that gave retrieval augmented generation its name.
How does RAG work, so the AI finds the right passage?
How does RAG work under the hood? Read the explanations from AWS, IBM, Google Cloud or Wikipedia and you get much the same picture: a preparation step and a query step. Preparation means cutting documents into chunks of a few paragraphs each and having an embedding model turn every chunk into a long string of numbers that roughly captures its meaning. Those numbers go into a vector database.
When a question arrives, it gets the same treatment, and the database returns the chunks whose numbers sit closest to it. Plenty of systems add an old-fashioned keyword search on top, which vendors call hybrid search, because a part number like 4711-B is often found more reliably by its exact characters than by its meaning. Whatever comes back is attached to the question and sent to the language model.
That sounds manageable, and in principle it is. The catch is that most of the effort sits in the search, not in the model: does it surface the paragraph with the current torque value, or the one from the 2019 manual nobody deleted? If the right passage never arrives, the model cannot use it, however clever it is.
| Step | What happens | Where it tends to go wrong |
|---|---|---|
| Chunking | Documents are split into passages | Chunks too small lose context, chunks too large blur the search |
| Embedding | Each passage becomes a vector and is stored | Changed documents have to be indexed again |
| Retrieval | The question is embedded and matched against stored passages | Part numbers and jargon get lost in pure meaning-based search |
| Augmentation | Retrieved passages are attached to the question | Off-topic hits steer the model the wrong way |
| Generation | The language model writes from the passages | The model can still add claims found nowhere in the sources |
RAG vs fine-tuning: which route gets company knowledge into an AI?
RAG vs fine-tuning is really a choice between two routes for getting your own knowledge into a model, and fine-tuning means training the model further on your data. The model itself changes, whereas with RAG it stays untouched and simply receives relevant reading material with every question. A team at Microsoft in Israel led by Oded Ovadia compared the two for factual knowledge, published at EMNLP 2024. Unsupervised fine-tuning helped a little, but RAG came out ahead throughout, both for knowledge the model had seen before and for entirely new facts.
We would not write fine-tuning off. It is better suited to teaching a model a tone of voice, a fixed output format or a specialist vocabulary, while RAG is the tool when facts change and have to be traceable. For company knowledge we would start with RAG in almost every case and only look at fine-tuning once there is a concrete reason.
| RAG | Fine-tuning | |
|---|---|---|
| What changes | The documents the model can draw on | The model itself |
| New documents | Usable once indexed | Only after another training run |
| Citing sources | Possible, since the passages are known | Not built in |
| Strength | Current, traceable facts | Style, format, vocabulary |
| Factual knowledge per Ovadia et al. | Ahead throughout | Some improvement |
When does RAG pay off for company knowledge?
RAG for company knowledge makes most sense where lots of people keep digging through the same folders and the answer needs to be traceable afterwards: sales staff who need the right dimensions for three hundred products, a support team working from help articles, a quality department with rulebooks that get a new revision every year. Whether an off-the-shelf assistant does the job or a dedicated system is worth it depends on volume and sensitivity, which we unpack in our post on the AI assistant for business.
RAG makes little sense when the whole collection is five documents that fit comfortably into a model's context window, the amount of text it can read in one go. Then you just pass the documents along. One caveat, from Stanford: in “Lost in the Middle”, Nelson Liu and colleagues found that models use information at the beginning and end of long inputs far better than information in the middle. Dumping everything in at once is a bet that the important part does not happen to land in the middle. If you would rather have someone sort out the sources, permissions and hosting with you, our page on RAG pipelines describes how we connect AI to company knowledge for businesses and when a knowledge graph is worth adding.
- 01
Customer service and sales
Often worth itMany recurring questions about products and processes, and answers have to match the current state.
- 02
Internal manuals, quality management, rulebooks
Often worth itLarge, changing document sets where the exact source matters.
- 03
A few short documents
Usually unnecessaryIf everything fits into a single request, the extra setup rarely pays off.
RAG hallucinations: where you still need to check the answer
RAG reduces hallucinations, meaning invented claims, but it does not remove them. The clearest evidence comes from a 2024 study by Stanford RegLab and Stanford HAI, preregistered so the results could not be bent after the fact. It tested legal research tools whose vendors had promoted RAG, some of them promising “hallucination-free” citations. The tools did better than GPT-4 on its own, yet depending on the product they got somewhere between a sixth and a third of the test queries wrong or misleading.
The second limit sits in the documents themselves. A RAG system repeats what it finds, so if the 2024 price list sits next to the 2026 one, or two versions of a work instruction contradict each other, the answer may well rest on the wrong one. Introducing RAG therefore also means deciding who maintains the collection and how changed files get re-indexed, on an ongoing basis.
And third, an answer with a citation is checkable, not automatically correct. Where contracts, health or money are involved, a person should look at the source before acting on it. Some questions also cannot be answered from single passages at all, for example when an amendment changes a contract; for cases like that, Microsoft researchers introduced GraphRAG in 2024, which additionally organises the collection as a network of entities and their relationships.
of test queries or more were answered wrongly or misleadingly by specialised RAG tools for lawyers in a Stanford study, up to a third depending on the tool.
What to watch for with RAG and data privacy when documents are confidential
RAG data privacy tends to mix up two separate questions. The first is where your documents go. If the system runs with a cloud provider, that provider processes the data on your behalf, and the German Data Protection Conference noted in its AI guidance of 6 May 2024 that this often requires a data processing agreement under Article 28(3) GDPR. If you cannot or would rather not do that, RAG can also run on your own servers; what “local” really means technically is explained in our post on local AI for business.
The second question gets asked less often and matters at least as much day to day: who is allowed to see what? A RAG system searches everything you give it. The OWASP project, which catalogues security risks in AI applications, lists weak access controls on vector databases explicitly as LLM08:2025, because otherwise the AI may reveal content the person asking was never meant to see. Salary lists do not belong in the same search index as the product manual, unless the search checks each document's permissions as it goes.
RAG does have one advantage here over a model with data trained into it, and Germany's Data Protection Conference points it out too. In its October 2025 guidance on AI systems using RAG, it notes that a rights and roles concept can be applied to the vector database and the stored documents, whereas inside the language model itself you cannot control who gets which information. It adds that RAG often makes it possible to run a smaller model on your own servers, so personal data never reaches the providers of large language models in the first place.
One technical point that IBM also raises in its explainer: the stored vectors can, under some circumstances, be traced back to the original text. The vector database deserves the same protection as the documents it was built from, not a bit less.
Frequently asked questions
What is RAG in simple terms?
An AI that looks things up before it answers. It searches documents you have chosen and bases its reply on the passages it finds there; the acronym stands for retrieval augmented generation.
What is the difference between an LLM and RAG?
The LLM is the language model that does the writing. RAG is the setup around it that puts a few relevant passages from a search in front of the model before every answer, so it does not have to rely on memory alone.
Is ChatGPT a RAG system?
Not by itself; ChatGPT is first of all a language model service. Once it searches uploaded files or connected sources and works the hits into its reply, it follows the RAG pattern (OpenAI documents a file search tool for developers, built on vector stores, that works exactly this way).
Does RAG train my model?
No. The model stays as it is and only gets reading material with each question, which is why a changed file has to be re-indexed, otherwise the search keeps finding the old version.
Does RAG stop hallucinations?
Unfortunately not, it only makes them rarer. In the 2024 Stanford study, specialised RAG tools for lawyers were wrong on 17 percent of test queries or more.
What types of RAG are there?
Roughly three: the basic form that searches by similarity of meaning, hybrid search that also matches keywords, and variants such as GraphRAG that work with entities and their relationships when a question spans many documents.
Is RAG compatible with the GDPR?
It depends on the setup. With a cloud provider you will often need a data processing agreement, and the search should only return what the individual user is allowed to see. If you want to be on the safe side, run it on your own servers.
How to check whether RAG fits your company knowledge
- 01
Log the questions
For two weeks, keep a shared list where the team notes every question that sent someone hunting through folders or old emails, along with where the answer finally turned up. After a fortnight you will see fairly clearly where RAG would help.
- 02
Clean up the sources
Look at where those answers lived. Three versions of the same form, a price list with no date, a manual nobody has touched in years? Whatever you leave in there, the AI will faithfully repeat later.
- 03
Flag what is confidential
Decide which documents only certain people may see, and whether the data may go to a cloud or has to stay in house. That shapes the setup more than any choice of model.
- 04
Start with one department
A contained area, such as the support team's help articles, is plenty to begin with. Let the subject experts check answers and sources for a few weeks before you connect the next area.
RAG does not turn a language model into an expert on your business, but it does turn it into one that can look things up. How well that works depends less on the model than on your documents and on who is allowed to see them, and that is exactly the part worth getting right before the first answer goes out.
AI agentsAgentic AI Explained: What It Means for Your Business
Google indexingRequest Indexing on Google: What the Button Really Does
AI Act loggingAI Act Logging Requirements: What Applies and When
SEO visibility indexSEO Visibility Index: What It Measures and What It Misses
Gemini 4 ArgonGemini 4 Argon: What the Model Is Actually Built For
Computer useComputer Use Agent: What It Operates and What It May Do
AI voice agentAI Voice Agent Pricing: Cost per Call and Disclosure
Google spam updateGoogle Spam Update: What Can Your Numbers Really Tell You?
LLM costsLLM Cost Optimization: When a Model Switch Pays Off
Detect AI-written textDetect AI-written text: what AI detectors get wrong
Google AI ModeGoogle AI Mode: What It Means for Your Website
AI agentsWhat Is an AI Agent, and How Is It Different From a Chatbot?
Structured DataStructured Data: What It Really Does for AI Search
AI AssistantAI Assistant for Business: Types, Uses and Data Protection
GEOE-E-A-T: Trust Signals for Google and AI Search
Local AILocal AI for Business: What “Local” Really Means
GEOAI Crawlers in robots.txt: Managing GPTBot and Co.
AI for SMEsAI for SMEs: How to Introduce AI Step by Step
Perplexity SEOPerplexity SEO: How to Get Your Site Cited as a Source
ChatGPT SEOChatGPT SEO: How to Get Your Business Found in ChatGPT
llms.txtllms.txt: What the File Does and When It Pays Off
AI SEOAI SEO: What Changes Compared to Traditional SEO
AI OverviewsGoogle AI Overviews: How Google Picks Its Sources
GEO vs AEO vs LLMOGEO vs AEO vs LLMO: The AI Search Terms Explained
Measure AI VisibilityMeasure AI Visibility: Method, Metrics and Limits
ChatGPT AdsChatGPT Ads: How to Advertise on ChatGPT in Germany
AI Literacy ObligationAI Literacy Obligation: What Article 4 Requires Since 2026
AI DisclosureAI Chatbot Disclosure: Article 50 in Practice
AI Text WatermarkAI Text Watermark: What Claude Marks and What It Doesn't
Where the information on this page comes from
- Lewis et al.: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (arXiv)retrieved 9 Oct 2026
- NeurIPS 2020 Proceedings: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasksretrieved 9 Oct 2026
- Meta AI: Retrieval Augmented Generation, Streamlining the creation of intelligent natural language processing modelsretrieved 9 Oct 2026
- Wikipedia (German): Retrieval-Augmented Generationretrieved 9 Oct 2026
- AWS: What is RAG?retrieved 9 Oct 2026
- IBM: What is retrieval-augmented generation?retrieved 9 Oct 2026
- Google Cloud: What is retrieval-augmented generation?retrieved 9 Oct 2026
- Microsoft Learn: Retrieval-augmented generation in Azure AI Searchretrieved 9 Oct 2026
- Ovadia et al.: Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs (arXiv)retrieved 9 Oct 2026
- ACL Anthology: Fine-Tuning or Retrieval? (EMNLP 2024)retrieved 9 Oct 2026
- Liu et al.: Lost in the Middle, How Language Models Use Long Contexts (arXiv)retrieved 9 Oct 2026
- ACL Anthology: Lost in the Middle (TACL 2024)retrieved 9 Oct 2026
- Magesh et al.: Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (arXiv)retrieved 9 Oct 2026
- Stanford HAI: AI on Trial, Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queriesretrieved 9 Oct 2026
- Edge et al.: From Local to Global, A Graph RAG Approach to Query-Focused Summarization (arXiv)retrieved 9 Oct 2026
- Microsoft on GitHub: GraphRAGretrieved 9 Oct 2026
- OWASP Gen AI Security Project: LLM08:2025 Vector and Embedding Weaknessesretrieved 9 Oct 2026
- German Data Protection Conference (DSK): Guidance on AI and data protection, 6 May 2024retrieved 9 Oct 2026
- German Data Protection Conference (DSK): Data protection specifics of generative AI systems using RAG, October 2025retrieved 9 Oct 2026
- LDA Brandenburg: DSK press release on the RAG guidance, 17 Oct 2025retrieved 9 Oct 2026
- EUR-Lex: Regulation (EU) 2016/679 (GDPR), Article 28retrieved 9 Oct 2026
- OpenAI: File search (API documentation)retrieved 9 Oct 2026


