What is RAG?
Learn how Retrieval-Augmented Generation connects an AI model to documents, databases, and current information so it can produce more relevant and grounded answers.
What you will learn
30-second explanation
RAG allows an AI application to search external knowledge before generating an answer.
Instead of asking an LLM to answer only from what it learned during training, the application retrieves relevant information and gives it to the model as additional context.
Understand
Why do we need RAG?
Large Language Models are powerful, but they do not automatically know your private documents, recently updated policies, internal systems, or information created after their training.
Without RAG
- ✕The model may provide generic answers.
- ✕It cannot access private company knowledge.
- ✕Its knowledge may be outdated.
- ✕It may confidently invent missing information.
With RAG
- ✓Answers can use organisation-specific information.
- ✓Knowledge can be updated without retraining the model.
- ✓Responses can include references and citations.
- ✓Retrieved context can make answers more relevant.
Visualize
How a RAG system works
A RAG application usually follows five steps for every user question.
Ask
A user asks a question in natural language.
Search
The system searches the connected knowledge base.
Retrieve
The most relevant pieces of information are selected.
Augment
The retrieved information is added to the LLM prompt.
Generate
The LLM produces an answer grounded in that context.
Real-work example
Employee travel policy assistant
An employee asks:
“Can I claim the cost of a hotel when travelling for a client meeting?”
1. Search
The system searches the latest company travel-policy documents.
2. Retrieve
It finds the section covering hotel eligibility and limits.
3. Generate
The LLM explains the policy in natural language.
4. Cite
The answer links back to the relevant policy section.
Architecture
Core components of a RAG system
Select each component to understand the role it plays in the pipeline.
1Source documents+
The original knowledge, such as PDFs, policies, web pages, product manuals, support tickets, or database records.
2Chunking+
Large documents are divided into smaller sections that can be searched and passed to the model efficiently.
3Embeddings+
Text is converted into numerical representations that capture semantic meaning.
4Vector database+
Embeddings and document metadata are stored so similar content can be found quickly.
5Retriever+
The retriever identifies the document chunks most relevant to the user's question.
6Large Language Model+
The LLM uses the retrieved context and the user's question to generate the final response.
Use
Where RAG is used
Company knowledge assistant
Answer employee questions using internal policies, procedures, and documentation.
Customer support copilot
Retrieve product manuals and previous resolutions before suggesting an answer.
Legal document search
Find relevant clauses and generate answers grounded in contracts or case documents.
Technical documentation assistant
Help developers search APIs, runbooks, architecture notes, and troubleshooting guides.
Compare
RAG vs fine-tuning
These techniques solve different problems and are sometimes used together.
| Area | RAG | Fine-tuning |
|---|---|---|
| Main purpose | Provide external knowledge | Change model behaviour |
| Updating information | Update the connected knowledge base | Train the model again with new examples |
| Private documents | Well suited | Usually not the best way to store facts |
| Citations | Can reference retrieved sources | Source attribution is more difficult |
| Best suited for | Current, private, or frequently changing knowledge | Tone, formatting, specialised tasks, and behaviour |
Consider RAG when
- ✓ Information changes frequently.
- ✓ Answers must use private documents.
- ✓ Users need citations or supporting sources.
- ✓ Knowledge is spread across many documents.
- ✓ Retraining a model would be too slow or expensive.
RAG may not be enough when
- • The task requires exact mathematical computation.
- • The source documents contain incorrect information.
- • The main goal is changing the model's writing style.
- • The system needs complex multi-step reasoning.
- • Reliable retrieval cannot be achieved.
Avoid
Common RAG mistakes
Poor document chunking
Chunks that are too large introduce irrelevant information. Chunks that are too small may lose the context needed to answer correctly.
Retrieving too much information
Sending many loosely related chunks to the LLM can reduce answer quality instead of improving it.
Ignoring metadata
Metadata such as document type, department, version, language, and publication date can greatly improve retrieval.
No evaluation process
A successful demo does not prove that the system is reliable. Test retrieval quality, answer accuracy, citations, latency, and cost.
Treating RAG as a hallucination cure
RAG can reduce hallucinations, but the model can still misunderstand, ignore, or incorrectly combine retrieved information.
Key takeaway
RAG gives an LLM access to the right knowledge at the right time.
A production-quality RAG system is not simply an LLM connected to a vector database. Its success depends on document quality, chunking, retrieval, metadata, prompting, evaluation, security, and source citations working together.
Remember the basic pattern:
Retrieve relevant information → augment the prompt → generate an answer grounded in the retrieved context.
Continue Learning
Related Lessons & Next Steps
Explore more practical AI guides from AIMates.
Stay in the loop
Get practical AI tutorials, frameworks, and real-work insights.