← INDEX Agent systems · 2 of 2

Building a Production-Ready Voice AI Agent with <500ms Latency

How I built a high-performance voice AI agent using Node.js, Deepgram, and OpenAI. A deep dive into system architecture, latency optimization techniques, and handling real-time audio streams.

Project Snapshot

Goal: Create a human-like voice assistant for restaurant reservations. Result: sub-500ms latency, seamless interruption handling, and robust state management. Stack: Node.js, Deepgram (Nova-2), OpenAI (GPT-4), WebSocket streams.

Building a voice AI that feels “natural” is notoriously difficult. The uncanny valley isn’t just about how the voice sounds—it’s about how it responds. A few hundred milliseconds of delay can turn a fluid conversation into an awkward walkie-talkie exchange.

In this post, I’ll walk you through the architecture of my Voice AI Agent, specifically focusing on how I engineered it for sub-500ms latency and high reliability.

The Latency Challenge

In a typical voice interaction loop, latency accumulates at every step:

  1. VAD (Voice Activity Detection): Waiting for the user to finish speaking.
  2. STT (Speech-to-Text): Transcribing audio to text.
  3. LLM Inference: Generating a response token-by-token.
  4. TTS (Text-to-Speech): Converting text back to audio.
  5. Network Overhead: Round-trips between services.

If you chain these sequentially, you easily hit 2-3 seconds of lag. That’s unacceptable for a real-time agent.

Architecture Guidelines

To solve this, I moved from a request-response model to a full-duplex streaming architecture.

1. WebSocket-First Design

Instead of HTTP REST requests, the entire pipeline operates over persistent WebSockets. This minimizes handshake overhead and allows for bidirectional data flow.

// Simplified conceptual flow
ws.on('message', (audioChunk) => {
  // 1. Immediately stream audio to Deepgram
  deepgramLive.send(audioChunk);
});

deepgramLive.on('transcript', (text) => {
  // 2. Stream transcript text to LLM
  llmChain.stream(text);
});

2. Optimizing Speech-to-Text (STT)

I used Deepgram’s Nova-2 model for its speed and accuracy. The critical configuration here is the endpointing (VAD) setting.

  • Too sensitive: It cuts you off while thinking.
  • Too relaxed: It introduces massive delays.

I tuned the VAD to a dynamic window of ~500ms and implemented a “Barge-in” mechanism. When the user starts speaking, the system emits a SpeechStarted event which immediately clears the audio buffer and stops the AI from talking.

Graceful Interruptions

“Barge-in” is the most important feature for perceived naturalness. If the user interrupts, the AI must shut up immediately. I achieved this by maintaining a server-side state of is_speaking and flushing the TTS audio queue the moment user audio input is detected.

3. Intelligent Filler Words

Even with optimization, LLM inference takes time. To mask this, I built an Intelligent Filler Word System.

When the STT pipeline detects a completed sentence, but the LLM hasn’t generated the first token, the system plays a context-aware filler sound (e.g., “Hmm,” “Let me check,” “One moment”). This buys the system ~800ms of “thinking time” without the user feeling ignored.

The “Brain”: State Management

Managing a conversation isn’t just about text; it’s about state. The agent needs to know:

  • Did I just ask a question?
  • Is the user confirming an order?
  • Did the call drop?

I implemented a finite state machine (FSM) to handle these transitions. This ensures the AI doesn’t hallucinate an order confirmation when the user was just asking about hours.

Performance Metrics

The results of these optimizations were significant:

MetricStandard PipelineOptimized PipelineImprovement
STT Latency600ms+less than 250ms58%
TTFB (Time to First Byte)800ms+~300ms62%
End-to-End Latency2.5s+~500ms5x Faster

Conclusion

Building a production-ready Voice AI Agent requires 50% AI and 50% distributed systems engineering. By decoupling the components and streaming everything, we can achieve human-level response times.

The code for this project is part of my portfolio and demonstrates that with the right architecture, we can bridge the gap between “chatbot” and “digital assistant.”

// REFERENCES
  1. Deepgram API Documentation Speech-to-text API reference
  2. OpenAI Realtime API Guidelines for low-latency AI interactions
  3. Effective Node.js WebSocket Handling Node.js documentation for UDP/Datagram sockets