Current Section

Overview

0%

← Back to AI Foundations
AI Foundations · Chapter 7

What is RAG?

Learn how Retrieval-Augmented Generation connects an AI model to documents, databases, and current information so it can produce more relevant and grounded answers.

Beginner8–10 min readPractical AI

What you will learn

✓What Retrieval-Augmented Generation means
✓How a complete RAG request flows through a system
✓The components used in a RAG architecture
✓When RAG is better than fine-tuning
✓Where RAG is used in real organisations
✓Which mistakes reduce production quality

30-second explanation

RAG allows an AI application to search external knowledge before generating an answer.

Instead of asking an LLM to answer only from what it learned during training, the application retrieves relevant information and gives it to the model as additional context.

User question → Retrieve relevant knowledge → Add context to the prompt → Generate a grounded answer

Understand

Why do we need RAG?

Large Language Models are powerful, but they do not automatically know your private documents, recently updated policies, internal systems, or information created after their training.

Without RAG

  • ✕The model may provide generic answers.
  • ✕It cannot access private company knowledge.
  • ✕Its knowledge may be outdated.
  • ✕It may confidently invent missing information.

With RAG

  • ✓Answers can use organisation-specific information.
  • ✓Knowledge can be updated without retraining the model.
  • ✓Responses can include references and citations.
  • ✓Retrieved context can make answers more relevant.

Visualize

How a RAG system works

A RAG application usually follows five steps for every user question.

01

Ask

A user asks a question in natural language.

02

Search

The system searches the connected knowledge base.

03

Retrieve

The most relevant pieces of information are selected.

04

Augment

The retrieved information is added to the LLM prompt.

05

Generate

The LLM produces an answer grounded in that context.

Real-work example

Employee travel policy assistant

An employee asks:

“Can I claim the cost of a hotel when travelling for a client meeting?”

1. Search

The system searches the latest company travel-policy documents.

2. Retrieve

It finds the section covering hotel eligibility and limits.

3. Generate

The LLM explains the policy in natural language.

4. Cite

The answer links back to the relevant policy section.

Architecture

Core components of a RAG system

Select each component to understand the role it plays in the pipeline.

1Source documents+

The original knowledge, such as PDFs, policies, web pages, product manuals, support tickets, or database records.

2Chunking+

Large documents are divided into smaller sections that can be searched and passed to the model efficiently.

3Embeddings+

Text is converted into numerical representations that capture semantic meaning.

4Vector database+

Embeddings and document metadata are stored so similar content can be found quickly.

5Retriever+

The retriever identifies the document chunks most relevant to the user's question.

6Large Language Model+

The LLM uses the retrieved context and the user's question to generate the final response.

Use

Where RAG is used

Company knowledge assistant

Answer employee questions using internal policies, procedures, and documentation.

Customer support copilot

Retrieve product manuals and previous resolutions before suggesting an answer.

Legal document search

Find relevant clauses and generate answers grounded in contracts or case documents.

Technical documentation assistant

Help developers search APIs, runbooks, architecture notes, and troubleshooting guides.

Compare

RAG vs fine-tuning

These techniques solve different problems and are sometimes used together.

AreaRAGFine-tuning
Main purposeProvide external knowledgeChange model behaviour
Updating informationUpdate the connected knowledge baseTrain the model again with new examples
Private documentsWell suitedUsually not the best way to store facts
CitationsCan reference retrieved sourcesSource attribution is more difficult
Best suited forCurrent, private, or frequently changing knowledgeTone, formatting, specialised tasks, and behaviour

Consider RAG when

  • ✓ Information changes frequently.
  • ✓ Answers must use private documents.
  • ✓ Users need citations or supporting sources.
  • ✓ Knowledge is spread across many documents.
  • ✓ Retraining a model would be too slow or expensive.

RAG may not be enough when

  • • The task requires exact mathematical computation.
  • • The source documents contain incorrect information.
  • • The main goal is changing the model's writing style.
  • • The system needs complex multi-step reasoning.
  • • Reliable retrieval cannot be achieved.

Avoid

Common RAG mistakes

Poor document chunking

Chunks that are too large introduce irrelevant information. Chunks that are too small may lose the context needed to answer correctly.

Retrieving too much information

Sending many loosely related chunks to the LLM can reduce answer quality instead of improving it.

Ignoring metadata

Metadata such as document type, department, version, language, and publication date can greatly improve retrieval.

No evaluation process

A successful demo does not prove that the system is reliable. Test retrieval quality, answer accuracy, citations, latency, and cost.

Treating RAG as a hallucination cure

RAG can reduce hallucinations, but the model can still misunderstand, ignore, or incorrectly combine retrieved information.

Key takeaway

RAG gives an LLM access to the right knowledge at the right time.

A production-quality RAG system is not simply an LLM connected to a vector database. Its success depends on document quality, chunking, retrieval, metadata, prompting, evaluation, security, and source citations working together.

Remember the basic pattern:

Retrieve relevant information → augment the prompt → generate an answer grounded in the retrieved context.

Continue Learning

Related Lessons & Next Steps

Explore more practical AI guides from AIMates.

Stay in the loop

Get practical AI tutorials, frameworks, and real-work insights.