Learn
GuideInstructions & contextCompanion to Context and memory
Retrieval and RAG: helping a bot answer from your sources
Retrieval means finding relevant information in a collection. Retrieval-augmented generation, usually shortened to RAG, combines finding information with generating an answer that uses it. [1]
A bot might search project notes, load a few useful passages, and then answer your question. The stored documents and the model’s learned knowledge are separate sources of information.
A question with two possible answers
Imagine a plant-swap organizer asks, “Which entrance should visitors use?” The project folder contains:
| Document | What it says |
|---|---|
| Early planning note | Entrance undecided; possibly the west door |
| Approved event note | Use the Oak Street entrance |
| Volunteer instructions | Volunteers enter through the loading area |
A search for “entrance” could return all three. The answer needs more than a word match. It needs the intended audience, the document’s status, and enough surrounding text to interpret it.
A useful reply would be: “Visitors should use the Oak Street entrance, according to the approved event note.” Linking that note lets you check the answer.
If the bot finds only the early planning note, it should explain that the entrance is unconfirmed in the source it found. It should not quietly turn a possibility into a final decision.
Swipe sideways to see the whole diagram, or open it full size.
Figure explanation: An early planning note suggests a possible west door. An approved event note says visitors use the Oak Street entrance. Volunteer instructions name a loading area for volunteers only. Checking the audience and approval status selects the approved event note for the visitor question. The resulting answer names Oak Street and identifies its source. These documents are invented teaching examples.
How information gets found
Some systems search for matching words. Others use embeddings, numerical representations that help compare meaning. A vector search compares these representations to find potentially relevant material. Systems can combine methods. [2]
Long documents are often split into smaller passages called chunks. That makes individual sections easier to retrieve, but a passage can lose meaning when separated from its heading or surrounding explanation. [2]
For example, the instruction “Use the loading area” is misleading if the missing heading says “Volunteers only.” The right excerpt needs enough context to preserve that distinction.
Why a sourced answer can still be wrong
The search might miss the best document. The collection might contain an outdated version. The generated answer might misread the passage or add something it does not support.
Check three things: Did it find the right source? Is that source current and relevant? Does the answer follow from it? A citation is a useful route to evidence; it does not make the claim correct by itself.
For your own files, use clear titles, label drafts, and record significant updates. Keep access restrictions attached to the source collection. A search result should not give someone access they did not already have permission to receive.
A short history
The 2020 paper Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks presented an influential combination of a language generator and retrieved material. It is a landmark for the term RAG, not the beginning of search or the first use of retrieved information in AI. [1]
Try a source-checking question
Ask: “Using the approved event note, tell me the visitor entrance. Identify the source. If it is missing or conflicts with another approved note, say so.”
That request gives you a result you can inspect. For the wider picture, return to context and memory. Product support for searching files, citations, and access controls must be checked separately.
Sources
- Patrick Lewis and colleagues, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, first submitted May 22, 2020; checked September 25, 2026. Supports the historical landmark and retrieval/generation combination.
- Anthropic, Introducing Contextual Retrieval, published September 19, 2024; checked September 25, 2026. Supports embeddings, keyword retrieval, and the loss of context when documents are divided into chunks. The entrance example is original.
