Current Section

Overview

0%

โ† Back to Real Projects
Real Projects ยท Project 8

Build an AI Voice Assistant

Build a production-grade multimodal voice interface that captures live microphone audio, transcribes speech via Whisper STT, reasons through GPT-4o, and synthesizes natural-sounding spoken audio responses using Text-to-Speech.

Next.js App RouterOpenAI Whisper STTText-to-Speech AudioProduction Recipe
voice-assistant-preview.tsx
Live UI Preview
๐ŸŽ™๏ธ

Listening to user voice input...

"Explain how sliding window memory works."

AIMates Hands-On Lab

Want to Test Voice Pipelines Live?

Launch our pre-configured sandbox with ready-to-run Whisper STT and TTS audio streaming routes.

Launch Sandbox Lab โ†’

The 30-Second Recipe

An AI voice assistant chains Speech-to-Text, LLM reasoning, and Text-to-Speech into a closed conversational loop.

Instead of reading text on a screen, your backend route accepts recorded microphone audio blobs, transcribes them via OpenAI Whisper, processes the query through GPT-4o, synthesizes natural audio via OpenAI TTS, and streams the spoken MP3 response back to the browser for auto-playback.

Microphone Blob โ†’ POST /api/voice โ†’ Whisper STT โ†’ GPT-4o โ†’ OpenAI TTS โ†’ Audio Playback

Topology

End-to-End System Architecture

Here is how data flows from spoken microphone input to synthesized voice output:

01

Audio Capture

User clicks record; browser MediaRecorder API captures microphone PCM audio buffers.

02

Speech-to-Text (Whisper)

Backend receives audio blob and invokes OpenAI Whisper API to transcribe spoken audio into text.

03

LLM Reasoning Engine

Transcribed text is passed into GPT-4o with conversation context to generate a helpful response.

04

Text-to-Speech Synthesis

Model text response is converted back into spoken audio streams (MP3/WAV) using OpenAI TTS models.

05

Audio Playback

Frontend receives the audio stream blob, creates an object URL, and auto-plays the response voice.

Step 1 ยท Backend Infrastructure

The Whisper STT & TTS API Route Handler

Create the backend route handler at app/api/voice/route.ts. It transcribes incoming audio files via Whisper, generates a response via GPT-4o, and synthesizes audio via TTS:

app/api/voice/route.ts
import OpenAI from "openai";

const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });

export async function POST(req: Request) {
  try {
    const formData = await req.formData();
    const audioFile = formData.get("audio") as File;

    if (!audioFile) {
      return new Response(JSON.stringify({ error: "Audio file is required" }), { status: 400 });
    }

    // 1. Transcribe speech using Whisper STT
    const transcription = await openai.audio.transcriptions.create({
      file: audioFile,
      model: "whisper-1",
    });

    const userText = transcription.text;
    if (!userText || userText.trim().length === 0) {
      return new Response(JSON.stringify({ error: "Could not transcribe audio" }), { status: 400 });
    }

    // 2. Generate intelligent response using GPT-4o
    const chatCompletion = await openai.chat.completions.create({
      model: "gpt-4o",
      messages: [
        { role: "system", content: "You are a concise, helpful voice assistant. Keep verbal responses under 3 sentences." },
        { role: "user", content: userText },
      ],
      temperature: 0.5,
    });

    const assistantText = chatCompletion.choices[0].message.content || "I am sorry, I could not process that.";

    // 3. Convert response text to spoken audio via OpenAI TTS
    const ttsResponse = await openai.audio.speech.create({
      model: "tts-1",
      voice: "alloy",
      input: assistantText,
    });

    const audioBuffer = Buffer.from(await ttsResponse.arrayBuffer());

    return new Response(audioBuffer, {
      headers: {
        "Content-Type": "audio/mpeg",
        "X-User-Transcript": encodeURIComponent(userText),
        "X-Assistant-Text": encodeURIComponent(assistantText),
      },
    });
  } catch (error) {
    console.error("Voice assistant error:", error);
    return new Response(JSON.stringify({ error: "Voice processing failed" }), { status: 500 });
  }
}

Step 2 ยท Frontend Implementation

The Microphone Recorder Workspace UI

Create the client component at components/tools/VoiceAssistant.tsx to handle MediaRecorder audio capture and automatic voice playback:

components/tools/VoiceAssistant.tsx
"use client";

import { useState, useRef } from "react";

export default function VoiceAssistant() {
  const [isRecording, setIsRecording] = useState(false);
  const [status, setStatus] = useState("Tap microphone to start speaking");
  const [transcript, setTranscript] = useState("");
  const [reply, setReply] = useState("");
  const mediaRecorderRef = useRef<MediaRecorder | null>(null);
  const audioChunksRef = useRef<Blob[]>([]);

  const startRecording = async () => {
    audioChunksRef.current = [];
    try {
      const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
      const mediaRecorder = new MediaRecorder(stream, { mimeType: "audio/webm" });
      mediaRecorderRef.current = mediaRecorder;

      mediaRecorder.ondataavailable = (e) => {
        if (e.data.size > 0) audioChunksRef.current.push(e.data);
      };

      mediaRecorder.onstop = async () => {
        setStatus("Processing voice pipeline...");
        const audioBlob = new Blob(audioChunksRef.current, { type: "audio/webm" });
        await sendAudioToBackend(audioBlob);
      };

      mediaRecorder.start();
      setIsRecording(true);
      setStatus("Listening... Speak into your microphone.");
    } catch (err) {
      console.error(err);
      setStatus("Microphone permission denied.");
    }
  };

  const stopRecording = () => {
    if (mediaRecorderRef.current && isRecording) {
      mediaRecorderRef.current.stop();
      setIsRecording(false);
      // Stop all audio tracks
      mediaRecorderRef.current.stream.getTracks().forEach((track) => track.stop());
    }
  };

  const sendAudioToBackend = async (blob: Blob) => {
    const formData = new FormData();
    formData.append("audio", blob, "voice-input.webm");

    try {
      const res = await fetch("/api/voice", {
        method: "POST",
        body: formData,
      });

      if (!res.ok) throw new Error("Voice pipeline failed");

      const userQuery = decodeURIComponent(res.headers.get("X-User-Transcript") || "");
      const assistantResponse = decodeURIComponent(res.headers.get("X-Assistant-Text") || "");
      setTranscript(userQuery);
      setReply(assistantResponse);

      // Play audio response
      const audioBlob = await res.blob();
      const audioUrl = URL.createObjectURL(audioBlob);
      const audio = new Audio(audioUrl);
      audio.play();
      setStatus("Tap microphone to speak again");
    } catch (err) {
      console.error(err);
      setStatus("Error processing voice request.");
    }
  };

  return (
    <div className="p-8 border border-[var(--app-border)] rounded-2xl bg-[var(--app-card)] text-center space-y-6 shadow-sm">
      <h3 className="text-sm font-black text-[var(--app-text)]">Live Voice Interaction Workspace</h3>

      <div className="flex justify-center">
        <button
          onClick={isRecording ? stopRecording : startRecording}
          className={`h-24 w-24 rounded-full flex items-center justify-center text-3xl transition shadow-xl ${
            isRecording ? "bg-rose-500 text-white animate-pulse" : "bg-amber-500 text-slate-950 hover:bg-amber-400"
          }`}
        >
          {isRecording ? "โน๏ธ" : "๐ŸŽ™๏ธ"}
        </button>
      </div>

      <p className="text-xs font-semibold text-[var(--app-text-secondary)]">{status}</p>

      {(transcript || reply) && (
        <div className="space-y-3 pt-4 text-left">
          {transcript && (
            <div className="p-3 rounded-xl bg-[var(--app-chip)]/40 border border-[var(--app-border)] text-xs">
              <strong className="block text-[var(--app-muted)] text-[10px] uppercase mb-1">You Spoke</strong>
              {transcript}
            </div>
          )}
          {reply && (
            <div className="p-3 rounded-xl bg-[var(--app-chip)]/40 border border-[var(--app-border)] text-xs">
              <strong className="block text-[var(--app-muted)] text-[10px] uppercase mb-1">Assistant Spoke</strong>
              {reply}
            </div>
          )}
        </div>
      )}
    </div>
  );
}

Performance Engineering

Optimizing Voice Latency for Real-Time Conversation

Traditional HTTP round-trips for STT โ†’ LLM โ†’ TTS can introduce noticeable delays. Production voice assistants employ two key optimizations:

OpenAI Realtime WebSockets API

Bypass separate REST roundtrips by streaming raw PCM audio bi-directionally over WebSockets directly to OpenAI's native Realtime model endpoints for sub-second responses.

Sentence Chunking & Streaming TTS

Begin synthesizing audio for the first sentence of an LLM response before the model finishes generating the final paragraph, eliminating dead air.

Level Up

Hands-On Build Challenges

Ready to take this voice assistant to production? Implement these three enhancements:

Challenge 1: Wake Word Detection

Integrate client-side Porcupine or Web Audio API wake word detection ("Hey AIMates") to trigger recording hands-free.

Challenge 2: Voice Selection

Allow users to toggle between OpenAI TTS voices (Alloy, Echo, Fable, Onyx, Nova, Shimmer).

Challenge 3: Tool Execution

Equip your voice assistant with calendar lookup or calculator tools using OpenAI Function Calling.

Release Gate

Production Voice Assistant Checklist

โœ“Microphone permission states are handled gracefully with clear browser permission prompts and fallback notices.
โœ“Audio blobs are compressed to web-optimized formats (WebM/MP3) before uploading to reduce network latency.
โœ“Server routes integrate OpenAI Whisper STT and TTS models with zero local client dependency.
โœ“UI provides clear visual states for listening, processing, speaking, and idle modes.
โœ“Error boundaries gracefully handle microphone hardware disconnects or network timeout failures.
โœ“API keys remain securely stored in server environment variables, never exposed to browser bundles.
โœ“Audio playback automatically cleans up object URLs to prevent browser memory leaks.

Key Takeaways

Voice interfaces make AI interactions as natural as human conversation.

Building an AI voice assistant demonstrates how to chain Speech-to-Text, LLM reasoning, and Text-to-Speech into a seamless multimodal product. By managing browser media recording and server audio streams efficiently, you unlock the future of hands-free computing.

Microphone Recording โ†’ Whisper STT โ†’ GPT-4o Reasoning โ†’ OpenAI TTS Playback.