What is RAG, in one sentence? A way to make an LLM answer questions using documents it was never trained on. It searches those documents for relevant passages, then hands the model the results as context before generating a response. No retraining, no fine-tuning — just a search step bolted onto the front of a normal prompt.
The term comes from a 2020 Facebook AI Research paper. It described the approach as combining “pre-trained parametric and non-parametric memory for language generation” — the model’s trained knowledge, plus a live lookup into an external index. Chatbots that answer from your internal docs, coding assistants that search your codebase, support bots that cite your help center — all of it is a variation on that same two-step loop.
What Is RAG?
Retrieval-augmented generation is a pattern, not a specific tool. Retrieve relevant information for a query, then generate an answer using that information as context. The original RAG paper, authored by Patrick Lewis and colleagues at Facebook AI Research, paired a pre-trained sequence-to-sequence model with “a dense vector index of Wikipedia, accessed with a pre-trained neural retriever.”
That’s the core idea, unchanged six years later: a retriever finds the relevant text, a generator writes the answer. What’s changed is the tooling — vector databases, embedding models, and chunking libraries have turned what used to be a research paper’s custom pipeline into something you can build in an afternoon.
RAG exists because LLMs have two hard limits. First, their knowledge is frozen at training time. Ask about something that happened last week, and a model with no retrieval has nothing to work with. Second, they don’t know your private data at all — your internal wiki, your ticket history, your codebase were never part of any public training set. Retrieval solves both problems by pulling the relevant text in at request time instead of baking it into the weights.
What Is RAG Actually Doing When It Retrieves Data?
Under the hood, RAG runs a fixed sequence for every question:
- Chunk your documents. Long documents get split into smaller passages — a paragraph or a few hundred words each — because you retrieve and inject chunks, not entire files.
- Embed each chunk. An embedding model converts each chunk into a vector: a list of numbers that captures its meaning, positioned so that similar text ends up close together in that vector space.
- Store the vectors. Those vectors go into a vector database alongside the original text, indexed for fast similarity search.
- Embed the question. When a user asks something, the same embedding model converts the question into a vector using the exact same process.
- Search for the closest matches. The database returns the chunks whose vectors are closest to the question’s vector — typically the top 3 to 10, depending on how the pipeline is tuned.
- Generate the answer. Those chunks get inserted into the LLM’s prompt as context, and the model writes an answer grounded in that text instead of only its training data.
None of these steps require a specialized model. Any LLM can do the generation step. The retrieval half is what makes RAG RAG.
What Is a Vector Database, and Why Does RAG Need One?
A vector database stores embeddings and answers one specific question fast: “which of these millions of vectors are closest to this new one?” Regular databases index by exact values or ranges. Vector databases index by geometric closeness instead. A search for “how do I reset a password” can match a chunk that says “forgot your login credentials” even though the words barely overlap — the embeddings land near each other because the meaning is similar.
You can run this locally without a hosted service. Ollama serves embedding models through the same API it uses for chat — the Ollama API guide covers the request format for /api/generate and /api/chat. Embeddings work the same way: send {"model": "<embedding model>", "input": "<text>"} to /api/embed. You get back {"embeddings": [[...]], "model": "...", "total_duration": ...} — one array of floats per input string.
nomic-embed-text is a solid default for a local pipeline: pull it with ollama pull nomic-embed-text. Ollama’s own listing states it “surpasses OpenAI text-embedding-ada-002 and text-embedding-3-small performance on short and long context tasks,” at a 274 MB download with a 2K-token context window per chunk. Running an embedding model alongside a chat model adds to your RAM and VRAM budget. The Ollama hardware requirements guide covers the real numbers for sizing that stack.
Why More Retrieved Chunks Doesn’t Mean Better Answers
⚠️ Note: the instinct to fix a bad RAG answer by retrieving more chunks often makes things worse, not better. A 2023 Stanford, Berkeley, and Samaya AI study found that language model performance on long-context tasks “is often highest when relevant information occurs at the beginning or end of the input context” and “significantly degrades when models must access relevant information in the middle of long contexts” — a pattern the paper’s title calls “lost in the middle.”
For a RAG pipeline, that means dumping your top 20 retrieved chunks into the prompt and hoping the model finds the right one is a real failure mode, not a theoretical one. A tighter retrieval step — top 3 to 5 genuinely relevant chunks, ordered so the strongest match sits near the start or end of the context — tends to outperform a looser one with more results and worse ranking. Retrieval quality matters more than retrieval quantity.
RAG vs Fine-Tuning: Which One Do You Actually Need?
| RAG | Fine-tuning | |
|---|---|---|
| Adds new knowledge | Yes, immediately | Yes, but baked into weights |
| Data changes daily | Handles it — just re-index | Needs retraining to update |
| Setup cost | Lower — no training run | Higher — needs a training pipeline and GPU time |
| Answers cite sources | Easy — you know which chunk was retrieved | Hard — knowledge is diffused through the weights |
| Changes model behavior/tone | No | Yes, this is what fine-tuning is actually good at |
| Best for | Answering from a document set that changes | Teaching a consistent style, format, or narrow skill |
The two aren’t competitors. RAG is for giving a model access to facts it doesn’t have. Fine-tuning is for changing how a model behaves, regardless of what facts it’s given. A support bot that needs to answer from this week’s changelog wants RAG. A model that needs to always respond in a specific JSON schema or a specific tone wants fine-tuning. Plenty of production systems use both.
What Is RAG Bad At? Common Mistakes Engineers Make
Chunking without testing retrieval quality. A chunk size that’s too large dilutes the embedding with irrelevant text; too small, and you lose the surrounding context a chunk needs to make sense on its own. Test retrieval on real questions before assuming your chunking strategy works — don’t just eyeball the chunks.
Assuming the retriever always finds something relevant. If nothing in the index actually answers the question, RAG doesn’t stop the model from confidently making something up. It just gives the model bad or irrelevant context to hallucinate on top of. A retrieval step that returns a similarity score below a sane threshold should tell the model there’s no good answer, not force-feed it the closest chunk anyway.
Skipping re-indexing after data changes. RAG’s advantage over fine-tuning is that it stays current. That advantage disappears if nobody re-embeds new or edited documents. A RAG pipeline pointed at a stale index is just an expensive way to answer with outdated information.
Treating every task as a RAG problem. If a task needs the model to reason over information that’s already in its training data, or to follow a consistent output format, adding retrieval doesn’t help and adds latency for nothing. RAG earns its complexity when the answer genuinely lives outside the model’s training data — a distinction worth making before wiring a tool-calling agent into a vector database it doesn’t actually need.
Frequently Asked Questions
Q: What is RAG in simple terms?
A: RAG is a way to make an AI model answer using your own documents. It searches your data for the most relevant passages, then feeds those to the model as context before it writes an answer.
Q: Is RAG the same as fine-tuning?
A: No. RAG adds external knowledge at request time without changing the model’s weights. Fine-tuning retrains the weights themselves, and is better suited to changing behavior, tone, or format than to injecting facts.
Q: What is a vector database used for in RAG?
A: It stores embeddings (numeric representations of text) and finds the ones most similar to a given question, so the pipeline knows which document chunks to hand the model as context.
Q: Can I run RAG locally without a cloud service?
A: Yes. Ollama can serve both the embedding model (like nomic-embed-text) and the generation model locally through its API, and several open-source vector databases run entirely on your own hardware.
Q: Does retrieving more documents always improve RAG answers?
A: No. Research on long-context models shows performance drops when relevant information is buried in the middle of a long context. A smaller set of well-ranked, genuinely relevant chunks usually beats a larger, noisier one.
Quick Summary — what is RAG, in practice:
– RAG pairs a retrieval step (vector search over your documents) with a normal LLM generation step, so the model answers using data it was never trained on
– The pipeline is: chunk documents, embed them, store the vectors, embed the question, retrieve the closest chunks, generate an answer from them
– Ollama’s /api/embed endpoint and models like nomic-embed-text let you run the entire embedding step locally
– More retrieved chunks isn’t automatically better — “lost in the middle” research shows models struggle with relevant info buried mid-context
– RAG and fine-tuning solve different problems: RAG for facts that change, fine-tuning for consistent behavior or format
You can see the retrieval mechanism working before touching a real vector database. Pull nomic-embed-text, then embed a document and a question with the same API call:
curl http://localhost:11434/api/embed -d '{"model": "nomic-embed-text", "input": "your paragraph of text here"}'
curl http://localhost:11434/api/embed -d '{"model": "nomic-embed-text", "input": "your test question here"}'
Do that for a handful of your own paragraphs, save the returned vectors, then compare the question’s vector against each one (cosine similarity is the usual metric) — the closest match is exactly what a vector database automates at scale. That’s the whole retrieval step in miniature, running entirely on your own machine, before you ever reach for a hosted service.