Build a Production RAG Application
Build an enterprise retrieval-augmented generation system from scratch. Learn how to chunk documents, generate dense vector embeddings, execute semantic similarity search, and ground LLM completions with verifiable source citations.
Semantic Query
[Doc #42 - Security Handbook v3]: All engineers accessing internal vector databases must connect via AWS PrivateLink...
Grounded Answer
VerifiedWant to Test RAG Retrieval Live?
Launch our pre-configured sandbox with ready-to-run embedding routes and vector search similarity matching.
The 30-Second Recipe
A RAG application grounds LLM completions in private document context via vector similarity search.
Instead of relying on the model's static training memory, your backend vectorizes incoming user queries, searches a vector index for the top-k most semantically relevant document chunks, and injects them into the system prompt with strict citation constraints.
Topology
End-to-End System Architecture
Here is how data flows from source documents to grounded semantic answers:
Document Ingestion
Raw knowledge documents are parsed and split into overlapping text chunks.
Embedding Generation
Each text chunk is converted into high-dimensional vector representations via text-embedding-3-small.
Vector Persistence
Embeddings and source metadata are stored in a vector index (pgvector, Pinecone, Qdrant).
Semantic Retrieval
User query is vectorized and matched against stored chunks using cosine similarity scoring.
Grounded Generation
Retrieved context chunks are injected into the prompt, forcing the LLM to cite exact sources.
Step 1 · Backend Infrastructure
The Vector Retrieval & Grounding API Route
Create the backend route handler at app/api/rag/route.ts. It vectorizes the user query, searches our mock document vector index, and generates a grounded response:
import OpenAI from "openai";
const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
// Mock knowledge base vector store
const KNOWLEDGE_BASE = [
{ id: "doc-1", title: "Security Handbook v3", content: "All engineers accessing internal vector databases must connect via encrypted AWS PrivateLink tunnels." },
{ id: "doc-2", title: "Q3 Roadmap", content: "We are migrating all legacy PostgreSQL instances to pgvector clusters by October." },
{ id: "doc-3", title: "Remote Work Policy", content: "Employees working remotely must use company-approved VPNs and hardware security keys." },
];
export async function POST(req: Request) {
try {
const { query } = await req.json();
if (!query || query.trim().length === 0) {
return new Response(JSON.stringify({ error: "Query is required" }), { status: 400 });
}
// 1. Generate embedding for user query
const embeddingRes = await openai.embeddings.create({
model: "text-embedding-3-small",
input: query,
});
const queryVector = embeddingRes.data[0].embedding;
// 2. Perform semantic search (In production, query pgvector or Pinecone)
// Here we simulate semantic retrieval by matching keywords for demonstration
const retrievedChunks = KNOWLEDGE_BASE.filter(doc =>
query.toLowerCase().split(" ").some(word => doc.content.toLowerCase().includes(word))
);
const context = retrievedChunks.length > 0
.map(d => `[Source: ${d.title}]\n${d.content}`)
.join("\n\n")
: "No direct internal documents found.";
// 3. Generate grounded LLM completion
const completion = await openai.chat.completions.create({
model: "gpt-4o",
messages: [
{
role: "system",
content: `You are an enterprise knowledge assistant. Answer the user's question strictly using the provided context. If the answer is not in the context, state that you do not know.\n\nContext:\n${context}`,
},
{ role: "user", content: query },
],
temperature: 0.2,
});
const answer = completion.choices[0].message.content;
return new Response(JSON.stringify({ answer, sources: retrievedChunks }), {
headers: { "Content-Type": "application/json" },
});
} catch (error) {
console.error("RAG pipeline error:", error);
return new Response(JSON.stringify({ error: "RAG query failed" }), { status: 500 });
}
}Step 2 · Frontend Implementation
The Interactive Q&A Workspace UI
Create the client component at components/tools/RagWorkspace.tsx to allow users to query documents and view verified source citations:
"use client";
import { useState } from "react";
interface Source {
id: string;
title: string;
content: string;
}
export default function RagWorkspace() {
const [query, setQuery] = useState("");
const [loading, setLoading] = useState(false);
const [answer, setAnswer] = useState("");
const [sources, setSources] = useState<Source[]>([]);
const handleSearch = async (e: React.FormEvent) => {
e.preventDefault();
if (!query.trim() || loading) return;
setLoading(true);
try {
const res = await fetch("/api/rag", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ query }),
});
const data = await res.json();
setAnswer(data.answer);
setSources(data.sources || []);
} catch (err) {
console.error(err);
} finally {
setLoading(false);
}
};
return (
<div className="space-y-8">
<form onSubmit={handleSearch} className="space-y-4 p-6 border border-[var(--app-border)] rounded-2xl bg-[var(--app-card)] shadow-sm">
<h3 className="text-sm font-black text-[var(--app-text)]">Internal Knowledge Search</h3>
<div className="flex gap-2">
<input
value={query}
onChange={(e) => setQuery(e.target.value)}
placeholder="Ask a question about internal policies or roadmap..."
className="flex-1 bg-[var(--app-bg)] border border-[var(--app-border)] rounded-xl p-3 text-xs text-[var(--app-text)] focus:ring-1 focus:ring-amber-500"
/>
<button
type="submit"
disabled={loading || !query.trim()}
className="bg-amber-500 hover:bg-amber-400 disabled:opacity-50 text-slate-950 font-bold px-6 py-3 rounded-xl text-xs transition shadow-sm"
>
{loading ? "Searching..." : "Ask RAG ⚡"}
</button>
</div>
</form>
{answer && (
<div className="space-y-6">
{/* Grounded Answer */}
<div className="p-6 border border-[var(--app-border)] rounded-2xl bg-[var(--app-card)] space-y-3 shadow-sm">
<h4 className="text-xs font-bold text-emerald-600 dark:text-emerald-400 uppercase tracking-wider">Grounded Answer</h4>
<p className="text-xs leading-relaxed text-[var(--app-text)] whitespace-pre-wrap">{answer}</p>
</div>
{/* Sources List */}
<div className="p-6 border border-[var(--app-border)] rounded-2xl bg-[var(--app-card)] space-y-3 shadow-sm">
<h4 className="text-xs font-bold text-sky-600 dark:text-sky-400 uppercase tracking-wider">Retrieved Source Documents</h4>
<div className="grid gap-3 sm:grid-cols-2">
{sources.map((src) => (
<div key={src.id} className="p-3 rounded-xl bg-[var(--app-chip)]/40 border border-[var(--app-border)] space-y-1">
<span className="font-bold text-xs text-[var(--app-text)]">{src.title}</span>
<p className="text-[11px] text-[var(--app-text-secondary)] line-clamp-2">{src.content}</p>
</div>
))}
</div>
</div>
</div>
)}
</div>
);
}Data Engineering
Advanced Chunking & Overlap Strategies
RAG retrieval quality is determined entirely by how cleanly source documents are chunked. Naive paragraph splitting cuts sentences in half, destroying semantic meaning:
Fixed-Size Chunking (Brittle)
Splitting text strictly every 500 characters regardless of punctuation. Often severs crucial context and negates embedding accuracy.
Semantic Overlap Chunking (Production)
Splitting text by semantic boundaries (markdown headers, paragraphs) with a 15% token overlap so context flows seamlessly between adjacent chunks.
Level Up
Hands-On Build Challenges
Ready to take this RAG app to production? Implement these three enhancements:
Replace mock vector storage with a real PostgreSQL database running the pgvector extension.
Pass retrieved top-20 chunks through a Cohere or BGE reranker to prune irrelevant context before LLM inference.
Combine dense vector similarity with sparse BM25 keyword search using Reciprocal Rank Fusion (RRF).
Release Gate
Production RAG App Readiness Checklist
Key Takeaways
RAG bridges static LLMs and dynamic enterprise knowledge.
Building a RAG application is the definitive milestone for practical AI engineers. By combining semantic chunking, dense vector embeddings, cosine similarity search, and grounded prompt constraints, you eliminate hallucinations and connect models directly to proprietary corporate documentation.