Blog · RAG · 9 Oct 2026
Blog
RAG

What is RAG?An AI that checksbefore it answers

Retrieval augmented generation makes a language model search your own documents before it replies. How the idea works, where it earns its keep and what it still gets wrong.

Get in touch
Enquiry
Nikolai Schöbel und Jeremias Burger, Co-Founder Scalableloops

Let's talk about your project.

First we check whether the project fits your business model. Then you get a proposal with phases and effort.

Talk about a RAG project or call: +49 151 1576 5566
Blog · RAG

RAG stands for retrieval augmented generation, and the short version is an AI that looks things up before it answers. Before the language model writes a single sentence, a search step digs through a collection you control (service manuals, price lists, contracts, the help desk wiki) and hands the model the handful of passages that fit the question. The term comes from a 2020 paper by Patrick Lewis and colleagues, accepted at NeurIPS that year. For a business the appeal is twofold: the AI can work with knowledge it never saw in training, without anyone retraining the model, and every claim it makes can be checked against the passage it came from. That does not make it infallible, and we will get to why.

In brief
  • With RAG, a search step pulls matching passages out of your documents first, and only then does the language model write its reply.
  • The term was coined in 2020 by researchers at Facebook AI Research, University College London and New York University.
  • Updated documents take effect once they are indexed, no retraining required, and the answer can point to the paragraph it relied on.
  • It is not foolproof: in a Stanford study, specialised RAG tools for lawyers got 17 percent or more of the test queries wrong.
Published 9 Oct 2026Nikolai Schöbel and Jeremias Burger9 min read
Nikolai SchöbelJeremias Burger

Nikolai Schöbel and Jeremias Burger

Co-founders of Scalableloops GmbH. Nikolai Schöbel leads online marketing and AI strategy, Jeremias Burger the AI architecture. Both build AI systems and train teams on them in their own agency work.

On this page
  1. What is RAG in AI, and what does your business get out of it?
  2. Retrieval augmented generation: why researchers wanted answers you can trace
  3. How does RAG work, so the AI finds the right passage?
  4. RAG vs fine-tuning: which route gets company knowledge into an AI?
  5. When does RAG pay off for company knowledge?
  6. RAG hallucinations: where you still need to check the answer
  7. What to watch for with RAG and data privacy when documents are confidential
  8. Frequently asked questions
  9. How to check whether RAG fits your company knowledge
  10. Where the information on this page comes from
Basics

What is RAG in AI, and what does your business get out of it?

What is RAG in AI, minus the jargon? A large language model (LLM) only knows what was in its training data. Your parts list from last quarter was certainly not in there, so when someone asks about it the model can either admit it does not know or, more often than anyone would like, assemble an answer that sounds right and is not.

RAG adds a detour. Say a field technician types on a tablet: “What torque applies to the bearing bolts on series B?” Before the model replies, a second program searches the service manuals, finds two or three relevant paragraphs and passes them along with the question. The model then writes from those paragraphs, a bit like an open-book exam where you are allowed to check the textbook before you start writing.

Two things change in practice. The AI can use material it was never trained on. And it can say which passage it leaned on, so the technician can read it themselves if in doubt (which, with a torque value, they probably should anyway).

Origins

Retrieval augmented generation: why researchers wanted answers you can trace

Retrieval augmented generation is a research term, and it is easy to date. On 22 May 2020 Patrick Lewis and eleven co-authors posted “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks” on arXiv; it was later accepted at NeurIPS 2020, one of the major machine learning conferences. The group came from Facebook AI Research (now part of Meta), University College London and New York University.

What bothered them is spelled out in the paper's opening lines: language models store a remarkable amount of factual knowledge, yet they are poor at retrieving it precisely, cannot show where a statement came from and are hard to update. Their answer paired a pre-trained text generator with a searchable index of Wikipedia, queried by a separate neural retriever. By their measurements the result was more specific, more diverse and more factual than a comparable model without retrieval.

When Meta wrote about the work in September 2020, it stressed the point that probably still matters most to companies. You change what such a system knows by swapping or extending the document collection, while the model itself stays exactly as it is.

2020

is the year Patrick Lewis and colleagues published the paper that gave retrieval augmented generation its name.

arXiv 2005.11401; NeurIPS 2020
Mechanics

How does RAG work, so the AI finds the right passage?

How does RAG work under the hood? Read the explanations from AWS, IBM, Google Cloud or Wikipedia and you get much the same picture: a preparation step and a query step. Preparation means cutting documents into chunks of a few paragraphs each and having an embedding model turn every chunk into a long string of numbers that roughly captures its meaning. Those numbers go into a vector database.

When a question arrives, it gets the same treatment, and the database returns the chunks whose numbers sit closest to it. Plenty of systems add an old-fashioned keyword search on top, which vendors call hybrid search, because a part number like 4711-B is often found more reliably by its exact characters than by its meaning. Whatever comes back is attached to the question and sent to the language model.

That sounds manageable, and in principle it is. The catch is that most of the effort sits in the search, not in the model: does it surface the paragraph with the current torque value, or the one from the 2019 manual nobody deleted? If the right passage never arrives, the model cannot use it, however clever it is.

StepWhat happensWhere it tends to go wrong
ChunkingDocuments are split into passagesChunks too small lose context, chunks too large blur the search
EmbeddingEach passage becomes a vector and is storedChanged documents have to be indexed again
RetrievalThe question is embedded and matched against stored passagesPart numbers and jargon get lost in pure meaning-based search
AugmentationRetrieved passages are attached to the questionOff-topic hits steer the model the wrong way
GenerationThe language model writes from the passagesThe model can still add claims found nowhere in the sources
Comparison

RAG vs fine-tuning: which route gets company knowledge into an AI?

RAG vs fine-tuning is really a choice between two routes for getting your own knowledge into a model, and fine-tuning means training the model further on your data. The model itself changes, whereas with RAG it stays untouched and simply receives relevant reading material with every question. A team at Microsoft in Israel led by Oded Ovadia compared the two for factual knowledge, published at EMNLP 2024. Unsupervised fine-tuning helped a little, but RAG came out ahead throughout, both for knowledge the model had seen before and for entirely new facts.

We would not write fine-tuning off. It is better suited to teaching a model a tone of voice, a fixed output format or a specialist vocabulary, while RAG is the tool when facts change and have to be traceable. For company knowledge we would start with RAG in almost every case and only look at fine-tuning once there is a concrete reason.

RAGFine-tuning
What changesThe documents the model can draw onThe model itself
New documentsUsable once indexedOnly after another training run
Citing sourcesPossible, since the passages are knownNot built in
StrengthCurrent, traceable factsStyle, format, vocabulary
Factual knowledge per Ovadia et al.Ahead throughoutSome improvement
Decision

When does RAG pay off for company knowledge?

RAG for company knowledge makes most sense where lots of people keep digging through the same folders and the answer needs to be traceable afterwards: sales staff who need the right dimensions for three hundred products, a support team working from help articles, a quality department with rulebooks that get a new revision every year. Whether an off-the-shelf assistant does the job or a dedicated system is worth it depends on volume and sensitivity, which we unpack in our post on the AI assistant for business.

RAG makes little sense when the whole collection is five documents that fit comfortably into a model's context window, the amount of text it can read in one go. Then you just pass the documents along. One caveat, from Stanford: in “Lost in the Middle”, Nelson Liu and colleagues found that models use information at the beginning and end of long inputs far better than information in the middle. Dumping everything in at once is a bet that the important part does not happen to land in the middle. If you would rather have someone sort out the sources, permissions and hosting with you, our page on RAG pipelines describes how we connect AI to company knowledge for businesses and when a knowledge graph is worth adding.

  1. 01

    Customer service and sales

    Often worth it

    Many recurring questions about products and processes, and answers have to match the current state.

  2. 02

    Internal manuals, quality management, rulebooks

    Often worth it

    Large, changing document sets where the exact source matters.

  3. 03

    A few short documents

    Usually unnecessary

    If everything fits into a single request, the extra setup rarely pays off.

Limits

RAG hallucinations: where you still need to check the answer

RAG reduces hallucinations, meaning invented claims, but it does not remove them. The clearest evidence comes from a 2024 study by Stanford RegLab and Stanford HAI, preregistered so the results could not be bent after the fact. It tested legal research tools whose vendors had promoted RAG, some of them promising “hallucination-free” citations. The tools did better than GPT-4 on its own, yet depending on the product they got somewhere between a sixth and a third of the test queries wrong or misleading.

The second limit sits in the documents themselves. A RAG system repeats what it finds, so if the 2024 price list sits next to the 2026 one, or two versions of a work instruction contradict each other, the answer may well rest on the wrong one. Introducing RAG therefore also means deciding who maintains the collection and how changed files get re-indexed, on an ongoing basis.

And third, an answer with a citation is checkable, not automatically correct. Where contracts, health or money are involved, a person should look at the source before acting on it. Some questions also cannot be answered from single passages at all, for example when an amendment changes a contract; for cases like that, Microsoft researchers introduced GraphRAG in 2024, which additionally organises the collection as a network of entities and their relationships.

17 %

of test queries or more were answered wrongly or misleadingly by specialised RAG tools for lawyers in a Stanford study, up to a third depending on the tool.

Magesh et al., Stanford RegLab and HAI, 2024
Privacy

What to watch for with RAG and data privacy when documents are confidential

RAG data privacy tends to mix up two separate questions. The first is where your documents go. If the system runs with a cloud provider, that provider processes the data on your behalf, and the German Data Protection Conference noted in its AI guidance of 6 May 2024 that this often requires a data processing agreement under Article 28(3) GDPR. If you cannot or would rather not do that, RAG can also run on your own servers; what “local” really means technically is explained in our post on local AI for business.

The second question gets asked less often and matters at least as much day to day: who is allowed to see what? A RAG system searches everything you give it. The OWASP project, which catalogues security risks in AI applications, lists weak access controls on vector databases explicitly as LLM08:2025, because otherwise the AI may reveal content the person asking was never meant to see. Salary lists do not belong in the same search index as the product manual, unless the search checks each document's permissions as it goes.

RAG does have one advantage here over a model with data trained into it, and Germany's Data Protection Conference points it out too. In its October 2025 guidance on AI systems using RAG, it notes that a rights and roles concept can be applied to the vector database and the stored documents, whereas inside the language model itself you cannot control who gets which information. It adds that RAG often makes it possible to run a smaller model on your own servers, so personal data never reaches the providers of large language models in the first place.

One technical point that IBM also raises in its explainer: the stored vectors can, under some circumstances, be traced back to the original text. The vector database deserves the same protection as the documents it was built from, not a bit less.

Frequently asked questions

Frequently asked questions

What is RAG in simple terms?

An AI that looks things up before it answers. It searches documents you have chosen and bases its reply on the passages it finds there; the acronym stands for retrieval augmented generation.

What is the difference between an LLM and RAG?

The LLM is the language model that does the writing. RAG is the setup around it that puts a few relevant passages from a search in front of the model before every answer, so it does not have to rely on memory alone.

Is ChatGPT a RAG system?

Not by itself; ChatGPT is first of all a language model service. Once it searches uploaded files or connected sources and works the hits into its reply, it follows the RAG pattern (OpenAI documents a file search tool for developers, built on vector stores, that works exactly this way).

Does RAG train my model?

No. The model stays as it is and only gets reading material with each question, which is why a changed file has to be re-indexed, otherwise the search keeps finding the old version.

Does RAG stop hallucinations?

Unfortunately not, it only makes them rarer. In the 2024 Stanford study, specialised RAG tools for lawyers were wrong on 17 percent of test queries or more.

What types of RAG are there?

Roughly three: the basic form that searches by similarity of meaning, hybrid search that also matches keywords, and variants such as GraphRAG that work with entities and their relationships when a question spans many documents.

Is RAG compatible with the GDPR?

It depends on the setup. With a cloud provider you will often need a data processing agreement, and the search should only return what the individual user is allowed to see. If you want to be on the safe side, run it on your own servers.

What next

How to check whether RAG fits your company knowledge

  1. 01

    Log the questions

    For two weeks, keep a shared list where the team notes every question that sent someone hunting through folders or old emails, along with where the answer finally turned up. After a fortnight you will see fairly clearly where RAG would help.

  2. 02

    Clean up the sources

    Look at where those answers lived. Three versions of the same form, a price list with no date, a manual nobody has touched in years? Whatever you leave in there, the AI will faithfully repeat later.

  3. 03

    Flag what is confidential

    Decide which documents only certain people may see, and whether the data may go to a cloud or has to stay in house. That shapes the setup more than any choice of model.

  4. 04

    Start with one department

    A contained area, such as the support team's help articles, is plenty to begin with. Let the subject experts check answers and sources for a few weeks before you connect the next area.

RAG does not turn a language model into an expert on your business, but it does turn it into one that can look things up. How well that works depends less on the model than on your documents and on who is allowed to see them, and that is exactly the part worth getting right before the first answer goes out.

or call: +49 151 1576 5566

Further reading
Sources

Where the information on this page comes from

Projekt-Detail

    Got a project in mind?

    We reply personally. First a use-case check, then an architecture proposal.

    Start your inquiry