Build an AI Voice Assistant
Build a production-grade multimodal voice interface that captures live microphone audio, transcribes speech via Whisper STT, reasons through GPT-4o, and synthesizes natural-sounding spoken audio responses using Text-to-Speech.
Listening to user voice input...
"Explain how sliding window memory works."
Want to Test Voice Pipelines Live?
Launch our pre-configured sandbox with ready-to-run Whisper STT and TTS audio streaming routes.
The 30-Second Recipe
An AI voice assistant chains Speech-to-Text, LLM reasoning, and Text-to-Speech into a closed conversational loop.
Instead of reading text on a screen, your backend route accepts recorded microphone audio blobs, transcribes them via OpenAI Whisper, processes the query through GPT-4o, synthesizes natural audio via OpenAI TTS, and streams the spoken MP3 response back to the browser for auto-playback.
Topology
End-to-End System Architecture
Here is how data flows from spoken microphone input to synthesized voice output:
Audio Capture
User clicks record; browser MediaRecorder API captures microphone PCM audio buffers.
Speech-to-Text (Whisper)
Backend receives audio blob and invokes OpenAI Whisper API to transcribe spoken audio into text.
LLM Reasoning Engine
Transcribed text is passed into GPT-4o with conversation context to generate a helpful response.
Text-to-Speech Synthesis
Model text response is converted back into spoken audio streams (MP3/WAV) using OpenAI TTS models.
Audio Playback
Frontend receives the audio stream blob, creates an object URL, and auto-plays the response voice.
Step 1 ยท Backend Infrastructure
The Whisper STT & TTS API Route Handler
Create the backend route handler at app/api/voice/route.ts. It transcribes incoming audio files via Whisper, generates a response via GPT-4o, and synthesizes audio via TTS:
import OpenAI from "openai";
const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
export async function POST(req: Request) {
try {
const formData = await req.formData();
const audioFile = formData.get("audio") as File;
if (!audioFile) {
return new Response(JSON.stringify({ error: "Audio file is required" }), { status: 400 });
}
// 1. Transcribe speech using Whisper STT
const transcription = await openai.audio.transcriptions.create({
file: audioFile,
model: "whisper-1",
});
const userText = transcription.text;
if (!userText || userText.trim().length === 0) {
return new Response(JSON.stringify({ error: "Could not transcribe audio" }), { status: 400 });
}
// 2. Generate intelligent response using GPT-4o
const chatCompletion = await openai.chat.completions.create({
model: "gpt-4o",
messages: [
{ role: "system", content: "You are a concise, helpful voice assistant. Keep verbal responses under 3 sentences." },
{ role: "user", content: userText },
],
temperature: 0.5,
});
const assistantText = chatCompletion.choices[0].message.content || "I am sorry, I could not process that.";
// 3. Convert response text to spoken audio via OpenAI TTS
const ttsResponse = await openai.audio.speech.create({
model: "tts-1",
voice: "alloy",
input: assistantText,
});
const audioBuffer = Buffer.from(await ttsResponse.arrayBuffer());
return new Response(audioBuffer, {
headers: {
"Content-Type": "audio/mpeg",
"X-User-Transcript": encodeURIComponent(userText),
"X-Assistant-Text": encodeURIComponent(assistantText),
},
});
} catch (error) {
console.error("Voice assistant error:", error);
return new Response(JSON.stringify({ error: "Voice processing failed" }), { status: 500 });
}
}Step 2 ยท Frontend Implementation
The Microphone Recorder Workspace UI
Create the client component at components/tools/VoiceAssistant.tsx to handle MediaRecorder audio capture and automatic voice playback:
"use client";
import { useState, useRef } from "react";
export default function VoiceAssistant() {
const [isRecording, setIsRecording] = useState(false);
const [status, setStatus] = useState("Tap microphone to start speaking");
const [transcript, setTranscript] = useState("");
const [reply, setReply] = useState("");
const mediaRecorderRef = useRef<MediaRecorder | null>(null);
const audioChunksRef = useRef<Blob[]>([]);
const startRecording = async () => {
audioChunksRef.current = [];
try {
const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
const mediaRecorder = new MediaRecorder(stream, { mimeType: "audio/webm" });
mediaRecorderRef.current = mediaRecorder;
mediaRecorder.ondataavailable = (e) => {
if (e.data.size > 0) audioChunksRef.current.push(e.data);
};
mediaRecorder.onstop = async () => {
setStatus("Processing voice pipeline...");
const audioBlob = new Blob(audioChunksRef.current, { type: "audio/webm" });
await sendAudioToBackend(audioBlob);
};
mediaRecorder.start();
setIsRecording(true);
setStatus("Listening... Speak into your microphone.");
} catch (err) {
console.error(err);
setStatus("Microphone permission denied.");
}
};
const stopRecording = () => {
if (mediaRecorderRef.current && isRecording) {
mediaRecorderRef.current.stop();
setIsRecording(false);
// Stop all audio tracks
mediaRecorderRef.current.stream.getTracks().forEach((track) => track.stop());
}
};
const sendAudioToBackend = async (blob: Blob) => {
const formData = new FormData();
formData.append("audio", blob, "voice-input.webm");
try {
const res = await fetch("/api/voice", {
method: "POST",
body: formData,
});
if (!res.ok) throw new Error("Voice pipeline failed");
const userQuery = decodeURIComponent(res.headers.get("X-User-Transcript") || "");
const assistantResponse = decodeURIComponent(res.headers.get("X-Assistant-Text") || "");
setTranscript(userQuery);
setReply(assistantResponse);
// Play audio response
const audioBlob = await res.blob();
const audioUrl = URL.createObjectURL(audioBlob);
const audio = new Audio(audioUrl);
audio.play();
setStatus("Tap microphone to speak again");
} catch (err) {
console.error(err);
setStatus("Error processing voice request.");
}
};
return (
<div className="p-8 border border-[var(--app-border)] rounded-2xl bg-[var(--app-card)] text-center space-y-6 shadow-sm">
<h3 className="text-sm font-black text-[var(--app-text)]">Live Voice Interaction Workspace</h3>
<div className="flex justify-center">
<button
onClick={isRecording ? stopRecording : startRecording}
className={`h-24 w-24 rounded-full flex items-center justify-center text-3xl transition shadow-xl ${
isRecording ? "bg-rose-500 text-white animate-pulse" : "bg-amber-500 text-slate-950 hover:bg-amber-400"
}`}
>
{isRecording ? "โน๏ธ" : "๐๏ธ"}
</button>
</div>
<p className="text-xs font-semibold text-[var(--app-text-secondary)]">{status}</p>
{(transcript || reply) && (
<div className="space-y-3 pt-4 text-left">
{transcript && (
<div className="p-3 rounded-xl bg-[var(--app-chip)]/40 border border-[var(--app-border)] text-xs">
<strong className="block text-[var(--app-muted)] text-[10px] uppercase mb-1">You Spoke</strong>
{transcript}
</div>
)}
{reply && (
<div className="p-3 rounded-xl bg-[var(--app-chip)]/40 border border-[var(--app-border)] text-xs">
<strong className="block text-[var(--app-muted)] text-[10px] uppercase mb-1">Assistant Spoke</strong>
{reply}
</div>
)}
</div>
)}
</div>
);
}Performance Engineering
Optimizing Voice Latency for Real-Time Conversation
Traditional HTTP round-trips for STT โ LLM โ TTS can introduce noticeable delays. Production voice assistants employ two key optimizations:
OpenAI Realtime WebSockets API
Bypass separate REST roundtrips by streaming raw PCM audio bi-directionally over WebSockets directly to OpenAI's native Realtime model endpoints for sub-second responses.
Sentence Chunking & Streaming TTS
Begin synthesizing audio for the first sentence of an LLM response before the model finishes generating the final paragraph, eliminating dead air.
Level Up
Hands-On Build Challenges
Ready to take this voice assistant to production? Implement these three enhancements:
Integrate client-side Porcupine or Web Audio API wake word detection ("Hey AIMates") to trigger recording hands-free.
Allow users to toggle between OpenAI TTS voices (Alloy, Echo, Fable, Onyx, Nova, Shimmer).
Equip your voice assistant with calendar lookup or calculator tools using OpenAI Function Calling.
Release Gate
Production Voice Assistant Checklist
Key Takeaways
Voice interfaces make AI interactions as natural as human conversation.
Building an AI voice assistant demonstrates how to chain Speech-to-Text, LLM reasoning, and Text-to-Speech into a seamless multimodal product. By managing browser media recording and server audio streams efficiently, you unlock the future of hands-free computing.