> Source: https://meetrix.io/blogs/webrtc-for-ai-voice-agents/
> Markdown copy of that page. Cite the URL above, not this file.

Development

# WebRTC for AI Voice Agents

[By Shalomi Umeshika](https://meetrix.io/blogs/authors/shalomi-umeshika/) • September 22, 2026 • 17 min read

When a voice agent feels slow, the model usually gets the blame. Sometimes that's fair. But plenty of "the AI is laggy" complaints actually come from the transport layer: the pipe that carries the caller's voice to the agent and the agent's voice back. It clips the first word, lets the agent talk over people, or freezes for half a second when the mobile signal dips.

In 2026, that pipe is often WebRTC, especially for browser and mobile clients. OpenAI recommends it for browser and mobile clients using its Realtime API. LiveKit is built around it. Pipecat supports it alongside other transports. The phone network is the big exception, and I'll get to why.

When we compared [LiveKit and Jitsi](https://meetrix.io/blogs/livekit-vs-jitsi/) a few weeks ago, the voice agent angle was what readers asked about most. Then we covered [the full self-hosted voice agent stack](https://meetrix.io/blogs/self-host-voice-ai-agent-stack/), layer by layer. This article zooms in on the layer both of those pieces took for granted: what WebRTC actually does for an AI voice agent, how to wire it up, and where it quietly breaks.

The short version

-   Browser or mobile app talking to an agent: use WebRTC. The difference shows up on bad networks, not in your office demo.
-   Phone callers only: SIP or your telephony provider's WebSocket stream is fine. WebRTC can't fix 8 kHz phone audio.
-   Your backend talking to a model API: WebSockets are fine. That hop runs on a clean datacenter network.
-   Humans and agents in the same call: WebRTC with an SFU like LiveKit, so the agent joins as just another participant.

## What WebRTC does for agents

WebRTC is the browser standard for real-time audio and video. It's what powers Google Meet, Jitsi, and most in-browser calling. For a voice agent, it does one job: move small chunks of audio in both directions, continuously, with as little delay as possible, over networks that are often bad.

That last part matters more than it sounds. Your test calls come from a laptop on fibre, three metres from the router. Your users call from a train, from a hospital basement, or from a café where forty people share one connection. A voice agent's transport gets judged in those conditions, not yours.

Under the hood, WebRTC is really a bundle of pieces: ICE, STUN, and TURN for getting through NATs and firewalls, DTLS-SRTP for encryption, RTP for timestamped media over UDP, the Opus codec, and SCTP data channels for everything that isn't audio. If you want the full vocabulary, our [WebRTC architecture explainer](https://meetrix.io/blogs/webrtc-architecture-explained/) goes through each part. For voice agents, you mostly need to know what each piece saves you from building.

## Why not just WebSockets?

This is the question everyone asks first, and it's a fair one. A WebSocket is simpler. Every backend developer already knows how to open one. You can send base64 audio chunks down it in an afternoon, and in a demo it works fine.

The problem is TCP.

WebSockets run on TCP, and TCP promises that every byte arrives, in order. That's exactly what you want for a chat message or a JSON payload. For live audio it's the wrong promise. When one packet goes missing, TCP holds back everything behind it until the lost packet is resent and arrives. LiveKit's engineers [have described the stall](https://livekit.com/blog/why-webrtc-beats-websockets-for-voice-ai-agents) as "possibly for hundreds of milliseconds." This is called head-of-line blocking, and in a conversation it sounds like the agent froze mid-sentence.

WebRTC sends audio over UDP instead. A lost packet is just lost. The receiver papers over the 20 ms gap and keeps playing. You'd rather hear a tiny glitch than silence followed by a burst of stale audio, and so would your caller.

So is WebSocket audio always worse? No. On a clean network the two feel the same, both well under 50 ms of transport delay inside a region. The gap appears when the network gets bad, and that's precisely when a user decides whether your agent is any good. We covered the general trade-offs in [WebRTC vs WebSocket](https://meetrix.io/blogs/webrtc-vs-websocket-real-time-communication/). For voice agents specifically, the verdict is simpler: WebSockets between servers, WebRTC to the user.

| Problem | WebRTC | WebSocket |
| --- | --- | --- |
| Lost packet | Skipped and concealed, playback continues | Stream stalls until TCP resends it |
| Uneven arrival (jitter) | Adaptive jitter buffer built in | You write your own buffer |
| Echo from speakers | Browser echo cancellation on the mic | Only if you capture the mic yourself with the right settings |
| Codec | Opus negotiated automatically | Raw PCM or whatever you encode by hand |
| Congestion | Media-aware congestion control adjusts bitrate | TCP backs off in bursts |
| Firewalls | Needs UDP, or TURN as a fallback | Goes anywhere HTTPS goes |
| Effort to start | Signalling, ICE, SDP to learn | An afternoon |

Look at the last two rows, though. WebSockets win on simplicity and on getting through firewalls. That's a real advantage, and it's why many teams start with WebSockets and move to WebRTC once real users show up.

## What you get for free

The best argument for WebRTC isn't UDP. It's the list of things the browser already does before a single byte leaves the device. Here are the ones that matter most for an agent.

### Echo cancellation

This is the one people underrate, and it's arguably the most important piece for a voice agent. When the agent speaks through a laptop speaker or a phone on speakerphone, the microphone picks that audio back up. Without echo cancellation, your voice activity detector hears the agent's own voice, decides the user is talking, and the agent interrupts itself. Over and over. WebRTC's acoustic echo cancellation removes the agent's playback from the mic signal before it's sent, because the browser knows exactly what it just played.

### The Opus codec

WebRTC endpoints must support Opus, along with G.711 for phone compatibility. Opus was designed for speech over the internet. It adapts its bitrate to the network, has built-in forward error correction that can rebuild a lost packet from data carried in the next one, and supports discontinuous transmission, which sends almost nothing during silence. A speaking user typically averages around 20 kbps. That's a rounding error next to a video stream.

### The jitter buffer

Packets don't arrive at even intervals, especially on Wi-Fi and mobile. WebRTC's adaptive jitter buffer holds a few tens of milliseconds of audio and releases it smoothly, growing when the network gets choppy and shrinking when it calms down. Over a WebSocket you build this yourself, and most people's first version either stutters or adds far too much delay.

### Congestion control

WebRTC watches delay growing between packets and lowers the bitrate before packets start getting dropped. TCP only reacts after loss, then recovers slowly. For audio this matters less than for video because voice is small, but it's the difference between graceful degradation and sudden dropouts when a shared connection fills up.

### Data channels

Not everything in a voice agent session is audio. Live transcripts, tool call results, "the agent is thinking" states, session config changes. WebRTC data channels carry all of that on the same connection as the audio, so you don't need a second WebSocket alongside it. OpenAI's Realtime API does exactly this with a channel named `oai-events`.

## Three ways to wire it

Nearly every production voice agent I've seen described uses one of three shapes. The right one depends on who's on the call and whose model you're using.

### 1\. Browser straight to the model

Browser → WebRTC → Model vendor (e.g. OpenAI Realtime)

The user's browser opens a WebRTC connection directly to the model provider. Your server's only job is to mint a short-lived credential so your real API key never touches the browser. OpenAI supports this with its Realtime API: [the official WebRTC guide](https://developers.openai.com/api/docs/guides/voice-webrtc) recommends it over WebSockets for browser and mobile clients.

It's the fastest way to a working demo, and there's no media server to run. The trade-off: your audio goes straight to a third party, you're tied to one vendor's speech-to-speech model, and you can't easily add a human or a phone caller to the same session. Speech-to-speech models also tend to be the most expensive option per minute.

### 2\. Agent as a participant in an SFU room

Browser / app → SFU (LiveKit) ← Agent worker → STT / LLM / TTS or realtime model

The user joins a room on a Selective Forwarding Unit. Your agent, running on your servers, joins the same room as another participant. The agent then talks to whatever models you like, hosted or self-hosted, over the backend network.

This is the pattern LiveKit Agents is built around, and it's the most flexible of the three. Want a supervisor to listen in? They join the room. A second agent to hand off to? Same room. A phone caller? A SIP bridge puts them in the room too. And because your agent sits between the user and the model, you can swap models without touching the client. It's also the pattern that lets you keep everything on your own infrastructure: our guide to [self-hosting LiveKit on AWS](https://meetrix.io/blogs/self-host-livekit-aws/) covers the server side.

The cost is one more hop and one more service to run. For a strictly one-to-one call, an SFU is honestly more than you need.

### 3\. Peer-to-peer with your own agent server

Browser → WebRTC P2P → Your agent process

Your agent process is itself the WebRTC endpoint. No SFU in between. Pipecat does this with its [SmallWebRTCTransport](https://docs.pipecat.ai/api-reference/server/services/transport/small-webrtc), built on the Python aiortc library. Teams building their own gateways at larger scale often reach for Pion, the Go WebRTC library.

This is the sweet spot for simple one-user-one-agent products. Fewer moving parts, one less network hop, and you still get everything WebRTC offers on the user's side. The limit is that it's strictly two parties. The moment you need a third, you're back to pattern 2.

### Pick an SFU pattern if

More than two parties will ever share a call, you need phone callers alongside browser users, or you expect to run many concurrent calls across several machines.

### Pick peer-to-peer if

Every session is one user and one agent, and you'd rather run one fewer service. You can move to an SFU later without changing your models.

## When WebRTC isn't worth it

Here's the part most "WebRTC for voice AI" articles skip.

If your callers dial a phone number, WebRTC probably won't help you. In March 2026, a developer running inbound Twilio calls through OpenAI's Realtime API [tested WebSocket against WebRTC](https://dev.to/nick_lackman/i-tested-our-websocket-audio-pipeline-with-webrtc-heres-why-i-switched-it-back-3g1j) and switched back. Average response time was about 1,920 ms over WebSocket and 2,060 ms over WebRTC. In other words, essentially the same, because the transport was under 5 percent of the total. That's one field test, not a universal rule, but the reasoning behind it is solid.

The reason is simple once you see it. Phone audio comes out of the phone network as G.711 at 8 kHz. By the time it reaches your servers, the quality is already set. Carrying those bytes over a nicer protocol doesn't add back the frequencies the phone line removed. And the lossy, jittery part of the journey (the caller's mobile network) is handled by the phone carrier, not by you.

For phone-only agents, you have two sensible options: a SIP trunk that delivers calls straight to your agent (OpenAI's Realtime API has [a SIP connector](https://developers.openai.com/api/docs/guides/realtime-sip), and LiveKit has an open-source SIP service), or your telephony provider's WebSocket media stream. If you've ever added dial-in to a meeting server, the idea is the same as [adding a SIP gateway to Jitsi with Jigasi](https://meetrix.io/blogs/add-sip-gateway-jitsi-meet-jigasi/), just pointed at an agent instead of a meeting room.

The same logic applies between servers. Your agent worker talking to a model API runs over a clean datacenter network where packet loss is close to zero. A WebSocket is perfectly fine there. That's why Gemini Live's native API is WebSocket-based, and why frameworks usually put WebRTC on the user side and WebSockets on the backend side.

So WebRTC earns its complexity on one specific hop: the last mile to a browser or mobile app. That happens to be the hop where most new voice products live.

## Interruptions and barge-in

Real people interrupt. They say "no, wait" halfway through the agent's answer, or "yeah, yeah" to hurry it along. How gracefully your agent handles that is where the transport shows up in ways that aren't obvious.

Two things have to work for barge-in to feel natural.

First, the agent has to hear the interruption at all, over the sound of its own voice. That's the echo cancellation point from earlier. Without it, you're stuck choosing between an agent that interrupts itself and one that can't be interrupted, which is how early IVR systems felt.

Second, when the user cuts in, the agent's audio has to stop right away. Not two seconds later, once the queued audio drains. This is a sneaky WebSocket problem. Text-to-speech generates audio faster than real time, so a WebSocket pipeline often pushes several seconds of speech to the client, which queues it for playback. When the user interrupts, the client has to be told to flush that queue, and the server needs to know how much the user actually heard, or the conversation history says the agent finished a sentence the user never heard. OpenAI's Realtime API, for example, asks WebSocket clients to report where playback was cut off. Over WebRTC, the audio is paced in real time and only a jitter buffer's worth sits on the client, so stopping the stream actually stops the voice.

Burst audio is your problem either way

That March 2026 case study found WebRTC initially sounded worse, because dumping TTS audio into a WebRTC track in bursts confused the jitter buffer. Audio has to be fed in at real-time pace, typically 20 ms frames. LiveKit Agents and Pipecat do this for you. If you're writing your own gateway, pacing is the first thing to get right.

## Where transport latency hides

A voice agent's response time is everything that happens between the user stopping and the agent starting: turn detection, speech-to-text, the model's first tokens, the first sentence of synthesised speech, and the network in both directions. The transport is usually the smallest of those. Usually.

Here's where it grows, in the order I'd check.

1.  **Distance.** Physics wins every time. LiveKit measured real round trips of 230 to 280 ms between Singapore and US East, against a theoretical one-way minimum of about 75 ms. If your users are in Asia and your agent runs in Virginia, no protocol choice will save you. Put the media server near the users, and the agent and models near the media server.
2.  **TURN relays in the wrong place.** When a user can only connect through TURN, every packet detours through the relay. A relay on another continent adds that continent's round trip to every turn of the conversation.
3.  **Extra hops between clouds.** Browser to a hosted SFU in one cloud, to your agent in a second cloud, to a model API in a third. Each hop is small. Stacked, they're not.
4.  **Oversized jitter buffers.** On a really bad connection the buffer grows to protect playback. That's the right trade-off, but it's why the same agent can feel slower on mobile data.

Everything else (turn detection, model choice, streaming TTS) we covered in [where the latency hides in a voice agent stack](https://meetrix.io/blogs/self-host-voice-ai-agent-stack/#latency). That's where most of your milliseconds live.

## Connecting a browser, in code

Here's what pattern 1 looks like in practice, a browser talking directly to OpenAI's Realtime API over WebRTC. Even if you end up using LiveKit or Pipecat, it's worth reading once, because it shows how little is really going on: capture the mic, attach it to a peer connection, open a data channel, swap SDP descriptions over HTTPS.

Start with the microphone. These constraints are the defaults in most browsers, but set them explicitly, because one missing flag is the most common cause of an agent that interrupts itself.

```plaintext
const mic = await navigator.mediaDevices.getUserMedia({
  audio: {
    echoCancellation: true,   // stops the agent hearing its own voice
    noiseSuppression: true,
    autoGainControl: true,
    channelCount: 1,          // mono is all speech recognition needs
  },
});
```

Then the connection. The ephemeral key is fetched from your own server first, so your real API key never reaches the browser.

```plaintext
// EPHEMERAL_KEY comes from your server, which calls
// POST /v1/realtime/client_secrets with your real API key.
const pc = new RTCPeerConnection();

// Play whatever the model says
const speaker = document.createElement("audio");
speaker.autoplay = true;
pc.ontrack = (event) => (speaker.srcObject = event.streams[0]);

// Send the microphone (with the constraints above)
pc.addTrack(mic.getAudioTracks()[0], mic);

// Transcripts, tool calls and session updates travel here
const events = pc.createDataChannel("oai-events");
events.onmessage = (msg) => console.log(JSON.parse(msg.data).type);

const offer = await pc.createOffer();
await pc.setLocalDescription(offer);

const answer = await fetch("https://api.openai.com/v1/realtime/calls", {
  method: "POST",
  body: offer.sdp,
  headers: {
    Authorization: "Bearer " + EPHEMERAL_KEY,
    "Content-Type": "application/sdp",
  },
});

await pc.setRemoteDescription({ type: "answer", sdp: await answer.text() });
```

That's the whole transport. No STUN configuration appears here because the model provider's servers have public addresses. The moment you host the other end yourself (pattern 2 or 3), you'll need to think about NAT, which brings us to the part that breaks most often.

Check the model name before you ship

OpenAI renames and versions its realtime models often. The session config (model, voice, instructions) is set when you create the client secret, so check the current model list in the docs rather than copying a name from a tutorial.

## Firewalls, NAT and TURN

UDP is WebRTC's biggest strength and its biggest deployment headache. Plenty of networks block it outright: corporate offices, hospitals, hotel Wi-Fi, some mobile carriers. When UDP can't get through, a WebRTC connection has one fallback: a TURN server that relays media over TCP or TLS, ideally on port 443 where it looks like ordinary HTTPS.

Skip TURN, and the failure mode is nasty. The call doesn't error. It just never connects, or connects with no audio. The user assumes your product is broken and leaves. They almost never file a bug report, so your dashboards look fine while a slice of your users silently can't use the product.

If you self-host the WebRTC side, here's the short list:

-   Open your media server's UDP range. For LiveKit that's typically 50000-60000 UDP plus TCP 7881, or a single UDP mux port such as 7882 in place of the range. Check your version's docs.
-   Run TURN on 443 over TLS. That's the one that gets through the strictest networks.
-   Put TURN in the same region as your media server, for the latency reason above.
-   Test from a phone on mobile data with Wi-Fi switched off, before launch. It takes two minutes and catches most of these problems.

If STUN, TURN and ICE are still a bit of a blur, our [STUN vs TURN vs ICE explainer](https://meetrix.io/blogs/stun-vs-turn-vs-ice-webrtc-nat-traversal/) covers what each one does and when you need your own relay. LiveKit ships a built-in TURN server, but plenty of teams run a dedicated Coturn box so several services can share it. Meetrix has a [pre-configured Coturn image](https://meetrix.io/store/coturn/) for exactly that job.

[

### Meetrix Coturn - Developer Guide

Welcome to the Meetrix Coturn Developer Guide! This guide is designed to assist you in seamlessly integrating Coturn into your AWS environment. Whether you're new to AWS or an experienced developer, you'll discover step-by-step instructions, configuration details, and troubleshooting tips to ensure a smooth experience.

By Dinesh Chathuranga • Meetrix.io

](https://meetrix.io/blogs/coturn-developer-guide/)

## Encryption and privacy

WebRTC media is always encrypted with DTLS-SRTP. There's no plaintext mode to accidentally leave on. That's a genuine plus next to a hand-rolled WebSocket pipeline, where encryption is only as good as your TLS configuration.

But be careful what you claim from it. DTLS-SRTP protects audio on the wire, between the user and whatever terminates the WebRTC connection. That endpoint (an SFU, your agent, or a model vendor) decrypts the audio, because it has to. You can't transcribe encrypted speech. For a voice agent, "encrypted in transit" is table stakes. The real privacy question is where the decrypting servers run and who operates them.

That's why pattern 2 matters so much in regulated industries. With a self-hosted SFU and self-hosted models, caller audio never leaves your own network. With pattern 1, it goes straight to the model vendor. Neither is wrong, but your compliance team will want to know which one you picked. If you're serving users in Europe, our piece on [the EU AI Act and WebRTC](https://meetrix.io/blogs/eu-ai-law-compliance-webrtc/) covers the other obligations, including telling callers they're talking to an AI.

## Picking your transport stack

Here's how the main options compare in September 2026. None of them is the best in the abstract. They sit at different points between "runs itself" and "you control everything."

| Option | What it is | Pattern | Self-hostable | Good fit for |
| --- | --- | --- | --- | --- |
| LiveKit server + Agents | Go SFU plus Python/Node agent framework, Apache 2.0 | 2 | Yes | Multi-party, phone plus browser, many concurrent calls |
| Pipecat SmallWebRTC | Peer-to-peer transport on aiortc, BSD 2-clause | 3 | Yes | One user, one agent, minimal infrastructure |
| OpenAI Realtime over WebRTC | Browser connects straight to OpenAI's model | 1 | No | Fast prototypes, speech-to-speech on OpenAI |
| Gemini Live | WebSocket API; WebRTC via LiveKit, Pipecat or Daily | 2 or 3 | Transport only | Gemini models behind a WebRTC front end |
| Daily | Hosted WebRTC infrastructure, maintainers of Pipecat | 2 | No | Pipecat teams who don't want to run media servers |
| Pion | Go WebRTC library, MIT | 3 (custom) | Yes | Teams building their own WebRTC-to-model gateway |

My honest take: most teams should start with LiveKit Agents or Pipecat, not a raw library. Both handle the pacing, interruption and turn-taking logic that's genuinely hard to get right. Building your own gateway on Pion or aiortc makes sense once you're at a scale where shaving a hop matters more than development speed. Hardly anyone is there on day one.

## Where I'd start

If you're building a voice agent people will reach from a browser or an app, use WebRTC from the first prototype. It's not much harder than WebSockets once a framework handles the signalling, and it's the version that holds up when your first real user calls from a train.

If every call comes from a phone number, don't let anyone talk you into WebRTC for its own sake. A SIP trunk is simpler and just as fast.

And whichever you choose, deploy TURN before launch day, not after the first silent failure.

**The one test worth doing this week:** call your agent from a phone on mobile data, Wi-Fi off, and interrupt it mid-sentence. If it connects, hears you, and stops talking, your transport is in good shape. If it doesn't, you now know which section of this article to reread.

## Frequently Asked Questions

Why do AI voice agents use WebRTC instead of WebSockets?

WebRTC sends audio over UDP, so one lost packet costs a tiny glitch instead of stalling the stream while TCP retransmits. It also brings Opus, echo cancellation, jitter buffering and congestion control built in. Over WebSockets you rebuild those pieces yourself.

Does the OpenAI Realtime API support WebRTC?

Yes. OpenAI recommends WebRTC over WebSockets for browser and mobile clients. Your server mints a short-lived client secret, the browser posts its SDP offer to the Realtime calls endpoint, and events flow over a data channel named oai-events.

Does Gemini Live support WebRTC?

Not natively. The Gemini Live API itself streams over WebSockets. To reach browser users over WebRTC, you put a framework such as LiveKit Agents or Pipecat in between: it speaks WebRTC to the user and keeps a WebSocket open to Gemini.

Do I need WebRTC for a phone-based voice agent?

Usually not. Phone audio arrives as 8 kHz G.711 from the PSTN, so WebRTC cannot improve its quality. A SIP trunk into your agent, or a WebSocket media stream from your telephony provider, is simpler and just as fast for phone-only traffic.

Do I need an SFU or media server for a voice agent?

Not for simple one-to-one calls. A peer-to-peer connection between the browser and your agent works fine. An SFU such as LiveKit starts to pay off once you add humans to the call, like a supervisor listening in, or need phone callers in the same room.

Why does my voice agent work on office Wi-Fi but fail on mobile data?

Almost always NAT or a firewall blocking UDP. Some mobile carriers and most corporate networks need a TURN server relaying media over TCP or TLS on port 443. Deploy TURN before launch, because users who fail to connect never report it.

Is WebRTC audio encrypted?

Always. WebRTC media is encrypted with DTLS-SRTP, and there is no switch to turn it off. That covers the network hop only. The agent server still decrypts the audio to transcribe it, so where that server runs is the real privacy decision.

How much latency does WebRTC itself add to a voice agent?

Very little inside one region. In one March 2026 case study, the transport was under 5 percent of total response time. Distance matters much more: LiveKit measured 230 to 280 ms real round trips between Singapore and US East.

## Add a TURN Server to Your Voice Agent Stack

Launch a pre-configured Coturn TURN/STUN server on AWS so users behind strict firewalls and mobile networks can still reach your WebRTC voice agent.

[Deploy Coturn from AWS Marketplace](https://aws.amazon.com/marketplace/pp/prodview-zrea7eq3c4jbe)

Meetrix Store

Coturn

TURN/STUN for WebRTC, no per-minute relay fees

[Deploy it](https://meetrix.io/store/coturn/)

Meetrix Store New

Deploy what this guide covers, pre-configured.

-    [Coturn TURN/STUN for WebRTC, no per-minute relay fees](https://meetrix.io/store/coturn/)
-    [Jitsi Meet Self-hosted video calls for 50 to 500 users](https://meetrix.io/store/jitsi-meet/)
-    [RustDesk Remote desktop AMI, a TeamViewer alternative](https://meetrix.io/store/rustdesk/)
-    [OpenVPN Encrypted remote access, no per-user fees](https://meetrix.io/store/openvpn/)
-    [Mailcow Business email on your own domain](https://meetrix.io/store/mailcow/)
-    [Listmonk Newsletters with no per-subscriber fees](https://meetrix.io/store/listmonk/)

[Browse all products](https://meetrix.io/store/)
