Current Section

Overview

0%

← Back to Real Projects
Real Projects · Project 3

Build a Production RAG Application

Build an enterprise retrieval-augmented generation system from scratch. Learn how to chunk documents, generate dense vector embeddings, execute semantic similarity search, and ground LLM completions with verifiable source citations.

Next.js App RouterVector EmbeddingsSemantic SearchProduction Recipe
rag-workspace-preview.tsx
Live UI Preview

Semantic Query

"What is our company policy on remote engineering data security?"
Retrieved Chunks (Top 3)

[Doc #42 - Security Handbook v3]: All engineers accessing internal vector databases must connect via AWS PrivateLink...

Run RAG Query ⚡

Grounded Answer

Verified
According to internal security guidelines, engineers must connect to vector storage exclusively through encrypted AWS PrivateLink tunnels to prevent public data exposure [Doc #42].
Source #1: Security Handbook v3
AIMates Hands-On Lab

Want to Test RAG Retrieval Live?

Launch our pre-configured sandbox with ready-to-run embedding routes and vector search similarity matching.

Launch Sandbox Lab →

The 30-Second Recipe

A RAG application grounds LLM completions in private document context via vector similarity search.

Instead of relying on the model's static training memory, your backend vectorizes incoming user queries, searches a vector index for the top-k most semantically relevant document chunks, and injects them into the system prompt with strict citation constraints.

Query Vectorization → Cosine Similarity Search → Context Assembly → Grounded LLM Response

Topology

End-to-End System Architecture

Here is how data flows from source documents to grounded semantic answers:

01

Document Ingestion

Raw knowledge documents are parsed and split into overlapping text chunks.

02

Embedding Generation

Each text chunk is converted into high-dimensional vector representations via text-embedding-3-small.

03

Vector Persistence

Embeddings and source metadata are stored in a vector index (pgvector, Pinecone, Qdrant).

04

Semantic Retrieval

User query is vectorized and matched against stored chunks using cosine similarity scoring.

05

Grounded Generation

Retrieved context chunks are injected into the prompt, forcing the LLM to cite exact sources.

Step 1 · Backend Infrastructure

The Vector Retrieval & Grounding API Route

Create the backend route handler at app/api/rag/route.ts. It vectorizes the user query, searches our mock document vector index, and generates a grounded response:

app/api/rag/route.ts
import OpenAI from "openai";

const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });

// Mock knowledge base vector store
const KNOWLEDGE_BASE = [
  { id: "doc-1", title: "Security Handbook v3", content: "All engineers accessing internal vector databases must connect via encrypted AWS PrivateLink tunnels." },
  { id: "doc-2", title: "Q3 Roadmap", content: "We are migrating all legacy PostgreSQL instances to pgvector clusters by October." },
  { id: "doc-3", title: "Remote Work Policy", content: "Employees working remotely must use company-approved VPNs and hardware security keys." },
];

export async function POST(req: Request) {
  try {
    const { query } = await req.json();

    if (!query || query.trim().length === 0) {
      return new Response(JSON.stringify({ error: "Query is required" }), { status: 400 });
    }

    // 1. Generate embedding for user query
    const embeddingRes = await openai.embeddings.create({
      model: "text-embedding-3-small",
      input: query,
    });
    const queryVector = embeddingRes.data[0].embedding;

    // 2. Perform semantic search (In production, query pgvector or Pinecone)
    // Here we simulate semantic retrieval by matching keywords for demonstration
    const retrievedChunks = KNOWLEDGE_BASE.filter(doc =>
      query.toLowerCase().split(" ").some(word => doc.content.toLowerCase().includes(word))
    );

    const context = retrievedChunks.length > 0
      .map(d => `[Source: ${d.title}]\n${d.content}`)
      .join("\n\n")
      : "No direct internal documents found.";

    // 3. Generate grounded LLM completion
    const completion = await openai.chat.completions.create({
      model: "gpt-4o",
      messages: [
        {
          role: "system",
          content: `You are an enterprise knowledge assistant. Answer the user's question strictly using the provided context. If the answer is not in the context, state that you do not know.\n\nContext:\n${context}`,
        },
        { role: "user", content: query },
      ],
      temperature: 0.2,
    });

    const answer = completion.choices[0].message.content;

    return new Response(JSON.stringify({ answer, sources: retrievedChunks }), {
      headers: { "Content-Type": "application/json" },
    });
  } catch (error) {
    console.error("RAG pipeline error:", error);
    return new Response(JSON.stringify({ error: "RAG query failed" }), { status: 500 });
  }
}

Step 2 · Frontend Implementation

The Interactive Q&A Workspace UI

Create the client component at components/tools/RagWorkspace.tsx to allow users to query documents and view verified source citations:

components/tools/RagWorkspace.tsx
"use client";

import { useState } from "react";

interface Source {
  id: string;
  title: string;
  content: string;
}

export default function RagWorkspace() {
  const [query, setQuery] = useState("");
  const [loading, setLoading] = useState(false);
  const [answer, setAnswer] = useState("");
  const [sources, setSources] = useState<Source[]>([]);

  const handleSearch = async (e: React.FormEvent) => {
    e.preventDefault();
    if (!query.trim() || loading) return;

    setLoading(true);
    try {
      const res = await fetch("/api/rag", {
        method: "POST",
        headers: { "Content-Type": "application/json" },
        body: JSON.stringify({ query }),
      });
      const data = await res.json();
      setAnswer(data.answer);
      setSources(data.sources || []);
    } catch (err) {
      console.error(err);
    } finally {
      setLoading(false);
    }
  };

  return (
    <div className="space-y-8">
      <form onSubmit={handleSearch} className="space-y-4 p-6 border border-[var(--app-border)] rounded-2xl bg-[var(--app-card)] shadow-sm">
        <h3 className="text-sm font-black text-[var(--app-text)]">Internal Knowledge Search</h3>
        <div className="flex gap-2">
          <input
            value={query}
            onChange={(e) => setQuery(e.target.value)}
            placeholder="Ask a question about internal policies or roadmap..."
            className="flex-1 bg-[var(--app-bg)] border border-[var(--app-border)] rounded-xl p-3 text-xs text-[var(--app-text)] focus:ring-1 focus:ring-amber-500"
          />
          <button
            type="submit"
            disabled={loading || !query.trim()}
            className="bg-amber-500 hover:bg-amber-400 disabled:opacity-50 text-slate-950 font-bold px-6 py-3 rounded-xl text-xs transition shadow-sm"
          >
            {loading ? "Searching..." : "Ask RAG ⚡"}
          </button>
        </div>
      </form>

      {answer && (
        <div className="space-y-6">
          {/* Grounded Answer */}
          <div className="p-6 border border-[var(--app-border)] rounded-2xl bg-[var(--app-card)] space-y-3 shadow-sm">
            <h4 className="text-xs font-bold text-emerald-600 dark:text-emerald-400 uppercase tracking-wider">Grounded Answer</h4>
            <p className="text-xs leading-relaxed text-[var(--app-text)] whitespace-pre-wrap">{answer}</p>
          </div>

          {/* Sources List */}
          <div className="p-6 border border-[var(--app-border)] rounded-2xl bg-[var(--app-card)] space-y-3 shadow-sm">
            <h4 className="text-xs font-bold text-sky-600 dark:text-sky-400 uppercase tracking-wider">Retrieved Source Documents</h4>
            <div className="grid gap-3 sm:grid-cols-2">
              {sources.map((src) => (
                <div key={src.id} className="p-3 rounded-xl bg-[var(--app-chip)]/40 border border-[var(--app-border)] space-y-1">
                  <span className="font-bold text-xs text-[var(--app-text)]">{src.title}</span>
                  <p className="text-[11px] text-[var(--app-text-secondary)] line-clamp-2">{src.content}</p>
                </div>
              ))}
            </div>
          </div>
        </div>
      )}
    </div>
  );
}

Data Engineering

Advanced Chunking & Overlap Strategies

RAG retrieval quality is determined entirely by how cleanly source documents are chunked. Naive paragraph splitting cuts sentences in half, destroying semantic meaning:

Fixed-Size Chunking (Brittle)

Splitting text strictly every 500 characters regardless of punctuation. Often severs crucial context and negates embedding accuracy.

Semantic Overlap Chunking (Production)

Splitting text by semantic boundaries (markdown headers, paragraphs) with a 15% token overlap so context flows seamlessly between adjacent chunks.

Level Up

Hands-On Build Challenges

Ready to take this RAG app to production? Implement these three enhancements:

Challenge 1: pgvector Integration

Replace mock vector storage with a real PostgreSQL database running the pgvector extension.

Challenge 2: Cross-Encoder Reranking

Pass retrieved top-20 chunks through a Cohere or BGE reranker to prune irrelevant context before LLM inference.

Challenge 3: Hybrid Search

Combine dense vector similarity with sparse BM25 keyword search using Reciprocal Rank Fusion (RRF).

Release Gate

Production RAG App Readiness Checklist

✓Document chunks maintain a 10% to 15% token overlap to prevent splitting sentences across boundaries.
✓Vector search queries enforce strict tenant_id or document-level filters at the database index layer.
✓Retrieved context is passed through a cross-encoder reranker to prune irrelevant chunks before LLM generation.
✓UI displays inline source citations linking directly to originating document titles and page numbers.
✓Empty queries and payloads exceeding context window token limits are rejected at the edge.
✓Error boundaries gracefully handle vector database timeouts and embedding provider rate limits.
✓API keys are sequestered safely in server environment variables, never exposed to browser bundles.

Key Takeaways

RAG bridges static LLMs and dynamic enterprise knowledge.

Building a RAG application is the definitive milestone for practical AI engineers. By combining semantic chunking, dense vector embeddings, cosine similarity search, and grounded prompt constraints, you eliminate hallucinations and connect models directly to proprietary corporate documentation.

Semantic Chunking → Vector Embeddings → Cosine Similarity → Grounded Citations.