> Source: https://meetrix.io/blogs/self-host-voice-ai-agent-stack/
> Markdown copy of that page. Cite the URL above, not this file.

LLMs & APIs

# Self-Host a Voice AI Agent Stack

[By Shalomi Umeshika](https://meetrix.io/blogs/authors/shalomi-umeshika/) • September 21, 2026 • 18 min read

You really can run a voice AI agent entirely on your own hardware these days, without handing a per-minute fee to anybody. Every piece of the puzzle has a solid open-source option: something to carry the audio (WebRTC), a model that turns speech into text, a language model for the thinking part, a voice to speak back with, and a framework that ties it all together and figures out when it's the agent's turn to jump in.

The tricky bit isn't tracking down the components. It's choosing the right ones and getting them to respond fast enough that the person on the other end doesn't just hang up.

A couple of weeks back, when we put [LiveKit and Jitsi head to head](https://meetrix.io/blogs/livekit-vs-jitsi/), the AI voice agent angle kept surfacing as the thing that really sets LiveKit apart from a plain meeting server. A lot of the follow-up questions boiled down to some version of: "Okay, but how do I run the whole thing myself?" Well, this is that article. We'll go through each layer, name the open-source project I'd genuinely reach for (with versions as they stand in September 2026), show you a working agent hooked up to self-hosted models, and be upfront about when it's just not worth the hassle.

The quick version

-   Transport: LiveKit server
-   Agent framework: LiveKit Agents (or Pipecat if you want finer control)
-   Speech-to-text: faster-whisper
-   Brain: an open-weights LLM running on vLLM
-   Voice: Kokoro-82M
-   Turn-taking: Silero VAD plus a turn detection model

Stick all of that in a single region, on one private network, with a GPU handling the STT and LLM. That's your self-hosted voice agent.

**By the numbers:**

0 API calls A fully self-hosted pipeline never lets caller audio or transcripts leave your network

1 x 24 GB GPU Enough for a pilot: an 8B-class LLM plus Whisper on a single card, with TTS and VAD running on CPU

~$720/month What a GPU instance at roughly $1/hour costs running around the clock, the number to beat a managed per-minute price

## What Exactly Is a Voice Agent Stack?

A voice AI agent is a program that listens to someone speak, works out what they meant, and talks back, in real time, usually over a phone line or in a browser. Underneath, it's a relay race:

Caller audio in → VAD + turn detection → Speech-to-text → LLM → Text-to-speech → Agent audio out

People often call this the cascaded or STT-LLM-TTS pipeline. The alternative is a speech-to-speech model that takes audio in and spits audio out directly, like OpenAI's realtime models. Those are genuinely impressive, but right now nearly every good one is a hosted API. If self-hosting is the goal, the cascaded pipeline is the realistic path in 2026. And honestly, it's easier to debug, because you can read the transcript at every step and pinpoint exactly which model got it wrong.

Around that pipeline you also need a transport (something that moves audio between caller and agent with minimal delay, which in practice means WebRTC), and optionally a telephony bridge so actual phone numbers can reach it.

## Why Bother Self-Hosting?

Managed voice AI platforms are good. Vapi, Retell, LiveKit Cloud and their ilk will have you talking to an agent in an afternoon. So why take on the servers?

From the conversations I have with customers, it comes down to three reasons, and usually it's the first one.

-   **The audio can't leave.** Clinics, banks, insurers, government departments, anyone taking calls in the EU. A recorded phone call is personal data (often sensitive personal data), and every third-party API it touches is another processor to add to your DPA and another place it can leak. Self-hosting turns "we send caller audio to four US vendors" into "the audio never leaves our VPC." If you're in Europe, our piece on [the EU AI Act and WebRTC](https://meetrix.io/blogs/eu-ai-law-compliance-webrtc/) covers the other half of this, including the rule that callers must be told they're talking to an AI.
-   **Per-minute pricing stops making sense at volume.** A managed pipeline bills you per minute for STT, again for the LLM tokens, again for TTS, and again for the platform. At a few hundred minutes a month that's nothing. At a few hundred thousand, a GPU box you own starts looking cheap.
-   You want to change things the platform won't let you change.

That last one is short on purpose. It's real (a custom fine-tuned model, an unusual telephony setup, a language the vendors handle badly), but it's rarely the reason on its own.

And here's the other side, stated plainly: if you need a working agent for a demo next week and you'll handle a few hundred calls a month, don't self-host. Use a managed platform, prove the idea works, and come back to this article when the invoice or the compliance team forces your hand.

## The Seven Layers, Picked

Here's each layer, what it does, and what I'd reach for. I've checked every version and license below against the project's own GitHub or docs this month, because this space moves fast enough that a six-month-old recommendation can already be wrong.

### 1\. Transport: LiveKit Server

Apache 2.0 v1.13.7 CPU only

[LiveKit](https://github.com/livekit/livekit) is a WebRTC SFU written in Go, and it's the default answer for the transport layer. Your agent joins a LiveKit room as just another participant. That sounds like a small design detail until you need a human supervisor to listen in on a call, or want to add a second agent, or put video in front of the voice. All of that just works, because everyone is a participant in the same room.

It's a single binary. One node handles a pilot comfortably. Going multi-node requires Redis as a shared message bus, and one thing the docs state but most tutorials skip: a single room must fit on a single node. For voice agents that's almost never a problem, since each call is its own small room, but it matters if you're planning big group sessions.

If you're new to SFUs and ICE, our [WebRTC architecture explainer](https://meetrix.io/blogs/webrtc-architecture-explained/) covers the vocabulary you'll run into when you open the firewall.

### 2\. Voice Activity Detection: Silero VAD

MIT v6.2.2 CPU only

[Silero VAD](https://github.com/snakers4/silero-vad) tells the pipeline whether a given slice of audio is speech or not. It's tiny, it runs on CPU, both major frameworks ship a plugin for it, and there's no serious reason to pick anything else. Moving on.

### 3\. Turn Detection: The Layer People Forget

LiveKit turn detector Smart Turn v3.2, BSD 2-clause CPU only

VAD knows when you stopped making sound. It doesn't know whether you're finished. "My account number is..." followed by a pause while you find the card is silence, but it isn't the end of your turn. An agent that only uses VAD will interrupt people constantly, and the usual fix (wait longer before replying) makes every single response feel slow.

Turn detection models fix this by looking at what was said, or how it was said, and predicting whether the speaker is done. Two options worth knowing:

-   **LiveKit's turn detector.** A 0.1B-parameter model that reads the transcript, runs on CPU in under 500 MB of memory, and covers 14 languages. Here's the catch almost nobody mentions: it's open-weights, not open source. The plugin source code is Apache 2.0, but the end-of-turn model itself ships under the LiveKit Model License, which is proprietary. You can use it freely, but only together with the LiveKit Agents framework: you can't extract it and run it standalone. For most companies that's fine. For anyone whose legal team asks for OSI-approved licenses across the stack, it isn't.
-   **[Smart Turn](https://github.com/pipecat-ai/smart-turn)**, from the Pipecat team. It listens to the audio itself rather than the transcript, so it picks up on intonation. Version 3.2 is BSD 2-clause, the CPU build is an 8 MB int8 model, the project quotes under 100 ms on most cloud instances, and it supports 23 languages.

If I had to point to one layer that decides whether your agent feels natural or robotic, it's this one. Not the LLM.

### 4\. Speech-to-Text: faster-whisper

MIT v1.2.1 GPU recommended

[faster-whisper](https://github.com/SYSTRAN/faster-whisper) reimplements OpenAI's Whisper on CTranslate2 and is noticeably quicker and lighter than the original. The easiest way to run it as a service is [Speaches](https://github.com/speaches-ai/speaches) (MIT), which wraps it in an OpenAI-compatible `/v1/audio/transcriptions` API. That compatibility is the trick that makes this whole stack pleasant: your agent framework thinks it's talking to OpenAI, and it's actually talking to a container on your own network.

Model size is the trade-off you'll keep revisiting. The small and distil models are fast enough for real-time on a modest GPU and good enough for clean English. Accents, background noise and domain vocabulary (drug names, product codes) push you towards the larger models, which cost latency.

One honest limitation: Whisper was built for transcribing finished clips, not streaming audio. Frameworks work around that by using VAD to cut speech into utterances and transcribing each one. It works well. If you need true word-by-word streaming, [sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx) (Apache 2.0) is the project to look at, with more assembly required.

### 5\. The LLM: vLLM or Ollama

vLLM 0.29.0, Apache 2.0 Ollama 0.34.2, MIT GPU

The brain. Both servers expose an OpenAI-compatible chat completions API, so your agent code doesn't care which one you pick. The difference is what happens under load.

Ollama is the fastest way to get a model answering on a laptop or a single box. Great for development. vLLM is built for serving many requests at once, with continuous batching and paged attention, and that's exactly what ten simultaneous phone calls look like. For production voice agents I'd use vLLM, and I'd say that even if Meetrix didn't package it. Our [guide to running a self-hosted OpenAI-compatible API](https://meetrix.io/blogs/openai-compatible-api-self-hosted/) goes deeper on the serving side.

On model choice: for voice, smaller is usually better. Time-to-first-token matters far more than benchmark scores, because the caller is sitting in silence while the model thinks. An 8B-class instruction-tuned model that handles tool calling well will beat a 70B model that makes people wait. Our [Llama vs Mistral vs DeepSeek comparison](https://meetrix.io/blogs/best-self-hosted-llm/) is a good shortlist to start from.

### 6\. Text-to-Speech: Kokoro-82M

Apache 2.0 82M parameters CPU capable

[Kokoro](https://huggingface.co/hexgrad/Kokoro-82M) is the open-weights TTS model that made self-hosted voices stop sounding like a GPS. It has 82 million parameters, 54 voices across 8 languages in v1.0, and it's small enough to run on CPU, which frees your GPU for the models that actually need it. Kokoro-FastAPI (Apache 2.0) wraps it in an OpenAI-compatible speech endpoint, same trick as Speaches.

Piper deserves a mention because it's even lighter and very popular in the Home Assistant world. But check the license before you build on it. The original `rhasspy/piper` repository is archived, and active development moved to `OHF-Voice/piper1-gpl`, which, as the name says, is GPL-3.0. That's a very different set of obligations from the MIT license most older tutorials quote.

Voice cloning models exist too. I'd leave them alone for anything customer-facing unless you have written consent from the person whose voice it is.

### 7\. Telephony: LiveKit SIP

Apache 2.0 Optional

If callers dial a phone number, you need something to translate between the phone network and WebRTC. LiveKit's open-source SIP service bridges an inbound or outbound SIP trunk into a room, so a phone caller becomes a participant like anyone else. You still need a SIP trunk provider for the actual number, and you'll open 5060 (and 5061 for TLS) plus an RTP range of 10000-20000 UDP. If you've already added phone dial-in to Jitsi, the concept is the same as [adding a SIP gateway with Jigasi](https://meetrix.io/blogs/add-sip-gateway-jitsi-meet-jigasi/), just pointed at an agent instead of a meeting.

## Stack at a Glance

| Layer | Pick | Alternative | License | Runs on |
| --- | --- | --- | --- | --- |
| Transport | LiveKit server 1.13.7 | Pipecat SmallWebRTC, WebSockets | Apache 2.0 | CPU |
| Agent framework | LiveKit Agents 1.8.2 | Pipecat 1.11.0 | Apache 2.0 / BSD 2-clause | CPU |
| VAD | Silero VAD 6.2.2 | (no real contender) | MIT | CPU |
| Turn detection | LiveKit turn detector | Smart Turn v3.2 | LiveKit Model License / BSD 2-clause | CPU |
| Speech-to-text | faster-whisper via Speaches | sherpa-onnx for streaming | MIT / Apache 2.0 | GPU |
| LLM serving | vLLM 0.29.0 | Ollama 0.34.2 | Apache 2.0 / MIT | GPU |
| Text-to-speech | Kokoro-82M via Kokoro-FastAPI | Piper (piper1-gpl) | Apache 2.0 / GPL-3.0 | CPU or GPU |
| Telephony | LiveKit SIP | Pipecat + Twilio or Telnyx websockets | Apache 2.0 | CPU |

## LiveKit Agents or Pipecat?

These are the two serious open-source frameworks for the orchestration layer, and they're both good. Anyone telling you one is clearly better hasn't used both for the same job.

[LiveKit Agents](https://github.com/livekit/agents) (Apache 2.0, 1.8.2 released September 15, 2026) is the higher-level of the two. You describe an agent, give it an STT, LLM and TTS, and the framework handles interruptions, turn-taking, tool calls and hand-offs between agents. Recent releases added duplex model support (speech models that speak and listen simultaneously), PII redaction for chat history and recordings, OpenTelemetry tracing, and a preforking worker that cuts agent start-up time. Its biggest advantage is that the transport, the framework and the SIP bridge come from one project and are designed to fit. The downside is that you're buying into LiveKit's room model for everything.

[Pipecat](https://github.com/pipecat-ai/pipecat) (BSD 2-clause, 1.11.0 released September 18, 2026), maintained by Daily and the community, models the agent as an explicit pipeline of processors that frames flow through. You see and control every step. It's transport-agnostic: its own SmallWebRTC transport for simple peer-to-peer, Daily, LiveKit, or plain websockets for Twilio, Telnyx, Vonage and Plivo phone streams. It has a huge list of service integrations and ships new releases very often.

Note that the question above is about the orchestration framework, not the transport underneath it. If it's LiveKit's room model itself you're trying to get away from, not just the Agents framework, our [LiveKit alternatives roundup](https://meetrix.io/blogs/livekit-alternatives/) covers Daily, OpenVidu and the rest of the field.

### Pick LiveKit Agents if

You want one coherent stack from WebRTC to phone line, expect to run many calls at once across several machines, or want human supervisors and agents in the same room.

### Pick Pipecat if

Your team wants to own every step of the pipeline, you're mostly bridging phone calls over websockets, or you need a service integration LiveKit doesn't have yet.

For the rest of this article I'll use LiveKit Agents, mainly because the transport and SIP layers are already LiveKit. Everything about the model layers applies to Pipecat just as well.

## Build It: A Working Agent

This is the skeleton, not a production deploy. It assumes four services on one private network: LiveKit server, a Speaches container for STT, vLLM (or Ollama) for the LLM, and Kokoro-FastAPI for TTS.

### Step 1: Run LiveKit Server

Save this as `livekit.yaml`. The single UDP mux port keeps firewall rules short. The default range is 50000-60000 UDP if you'd rather not mux.

```plaintext
port: 7880
rtc:
  tcp_port: 7881
  udp_port: 7882          # single UDP mux port, simpler firewall rules
  use_external_ip: true
keys:
  voiceagent: change-this-to-a-long-random-secret
logging:
  level: info
```

```plaintext
docker run -d --name livekit \
  -p 7880:7880 -p 7881:7881 -p 7882:7882/udp \
  -v $PWD/livekit.yaml:/livekit.yaml \
  livekit/livekit-server --config /livekit.yaml
```

Put port 7880 behind a load balancer or reverse proxy that terminates TLS, so clients connect over `wss://`. LiveKit has a built-in TURN server (5349 for TLS, 3478 for UDP) you can enable in the same config. If you'd rather run a dedicated TURN box that other services can share, our [Coturn developer guide](https://meetrix.io/blogs/coturn-developer-guide/) walks through it.

Open the ports before you debug anything else

TCP 7881 plus UDP 7882 (or 50000-60000) must be reachable from the internet. If they aren't, the agent joins the room, the logs look healthy, and nobody hears a thing. It's the most common "it's broken" report for new LiveKit installs.

### Step 2: Write the Agent

The trick is the OpenAI plugin. Its STT, LLM and TTS classes all accept a `base_url`, so you point them at your own containers. The `api_key` has to be set to something even though your local servers ignore it, because the OpenAI client refuses to start without one.

```plaintext
from livekit import agents
from livekit.agents import Agent, AgentServer, AgentSession, TurnHandlingOptions
from livekit.plugins import openai, silero
from livekit.plugins.turn_detector.multilingual import MultilingualModel

class Receptionist(Agent):
    def __init__(self) -> None:
        super().__init__(
            instructions=(
                "You answer the phone for a dental clinic. "
                "Keep every reply under two sentences. "
                "Plain spoken English only: no lists, no markdown, no emoji."
            ),
        )

server = AgentServer()

@server.rtc_session(agent_name="receptionist")
async def entrypoint(ctx: agents.JobContext):
    session = AgentSession(
        vad=silero.VAD.load(),
        # Speaches (faster-whisper) exposes an OpenAI-compatible STT API
        stt=openai.STT(
            base_url="http://stt.internal:8000/v1",
            api_key="local",
            model="Systran/faster-whisper-small",
            use_realtime=False,
        ),
        # vLLM or Ollama, both speak the OpenAI chat completions API
        llm=openai.LLM(
            base_url="http://llm.internal:8000/v1",
            api_key="local",
            model="your-served-model-name",
        ),
        # Kokoro-FastAPI exposes an OpenAI-compatible speech endpoint
        tts=openai.TTS(
            base_url="http://tts.internal:8880/v1",
            api_key="local",
            model="kokoro",
            voice="af_heart",
        ),
        turn_handling=TurnHandlingOptions(turn_detection=MultilingualModel()),
    )

    await session.start(room=ctx.room, agent=Receptionist())
    await session.generate_reply(instructions="Greet the caller and ask how you can help.")

if __name__ == "__main__":
    agents.cli.run_app(server)
```

Install the pieces with `pip install "livekit-agents[openai,silero,turn-detector]"`. Notice the instructions in the prompt. "No lists, no markdown, no emoji" is not decoration: LLMs love bullet points, and a TTS model will cheerfully read "asterisk asterisk" out loud to your caller.

### Step 3: Connect It and Start the Worker

```plaintext
export LIVEKIT_URL=wss://rtc.example.com
export LIVEKIT_API_KEY=voiceagent
export LIVEKIT_API_SECRET=change-this-to-a-long-random-secret

python agent.py download-files   # pulls Silero and turn detector weights once
python agent.py start
```

The worker registers with LiveKit and gets dispatched into rooms as calls arrive. Run more workers on more machines and LiveKit spreads calls across them. That's your horizontal scaling story for the agent side.

[

### vLLM on AWS - Developer Guide

Deploy vLLM, the high-throughput OpenAI-compatible LLM inference server, on AWS with our step-by-step CloudFormation guide. Learn how to launch a GPU instance, pick a model, secure it with SSL, and call the API.

By Binuka Ranatunga • Meetrix.io

](https://meetrix.io/blogs/vllm-developer-guide/)

## Where the Latency Hides

A voice agent lives or dies on how long the caller waits after they stop talking. In normal conversation people answer each other within a fraction of a second, and a gap much longer than a second starts to feel like the line dropped. You won't hit human speed with a cascaded pipeline, but you can get close enough that nobody comments on it.

The delay stacks up like this: the turn detector decides the caller is done, STT finishes the transcript, the LLM produces its first few tokens, TTS turns the first sentence into audio, and that audio crosses the network. Here's where I'd look, in order.

1.  **Turn detection settings.** Usually the biggest single chunk, and it's a setting, not a model. A fixed "wait 800 ms of silence" rule costs 800 ms on every turn. A turn detection model lets you reply sooner when the sentence is obviously finished and wait longer when it isn't.
2.  **LLM time-to-first-token.** Smaller model, shorter system prompt, prefix caching turned on in vLLM. And tell the model to keep replies short, because a shorter answer is also a faster one to speak.
3.  **Stream everything.** TTS should start on the first sentence while the LLM is still writing the second. Both frameworks do this by default. Don't break it by waiting for the full LLM response in a tool.
4.  **Distance.** Put the agent workers, the GPU box and LiveKit in the same region, ideally the same private subnet. Every hop between clouds adds a round trip to every turn.
5.  **Cold starts.** The first call after a deploy is always the slow one because models are loading. Warm them up with a dummy request at start-up.

LiveKit Agents also has preemptive generation, which starts the LLM on a probable final transcript before the turn is confirmed. It trades a little wasted compute for a noticeably snappier feel. Worth switching on once everything else works.

## Hardware and Real Costs

For a pilot handling a handful of calls at once, you need less than you'd think:

-   One GPU instance with 24 GB of VRAM for STT and the LLM. On AWS that's a g5.xlarge (NVIDIA A10G, 24 GiB GPU memory) or a g6.xlarge (NVIDIA L4, 24 GiB GPU memory). An 8B-class model plus a small Whisper fits on one card.
-   One modest CPU instance for LiveKit server, the agent workers, Kokoro and the VAD and turn detection models.

That's it for a start. As concurrency grows, the LLM is what you'll scale first, then STT. Keeping them on separate GPUs once you're past the pilot stops them fighting over memory at the worst moment. If you only need the same STT and LLM pairing after a call, not live, the [Jitsi notetaker with Whisper](https://meetrix.io/blogs/jitsi-ai-notetaker-whisper/) shows a batch version.

The cost maths is where people fool themselves, so let's do it properly. A GPU instance at about $1 an hour on demand is roughly $720 a month if it runs around the clock. That cost is the same whether it handles ten calls or ten thousand. Now take your all-in managed cost per minute. Suppose it's 5 cents once you add up platform, STT, LLM and TTS. Then $720 buys you 14,400 managed minutes, about 240 hours of talk time. Below that, the managed platform is cheaper, before you even count engineering time. Well above it, and self-hosting starts saving real money every month.

Reserved instances or savings plans lower the GPU line considerably. Your engineer's time doesn't get cheaper. Budget for someone who owns this stack, because it will page them.

If you'd rather not build the GPU and LLM layer by hand, Meetrix packages [vLLM as a ready-to-run image](https://meetrix.io/store/vllm/) for AWS and GCP, which covers the biggest piece of this stack: GPU drivers, the OpenAI-compatible API and SSL already set up.

## What Breaks First

Here's what goes wrong in roughly the order you'll hit it.

**No audio, healthy logs.** Blocked UDP. Covered above, but it's worth repeating because it eats afternoons.

**Works at the office, fails for customers.** Corporate firewalls and some mobile networks block UDP completely. Without TURN on 443 or 5349, those users can't connect at all. Enable TURN before launch, not after the first complaint.

**The agent talks over people.** You're running on VAD alone, or your endpointing delay is too short. Add a turn detector.

**The agent reads out formatting.** Markdown, bullet points, URLs spelled letter by letter. Fix it in the system prompt, and filter the text before TTS for anything the prompt misses.

**Great for one call, awful for five.** STT and the LLM are sharing a GPU and queueing behind each other. Either give them separate cards or cap concurrency per box honestly and scale out.

**Numbers and names get mangled.** Whisper hears "fifteen" when the caller said "fifty", or can't spell your product name. Pass a prompt with expected vocabulary to STT, and have the agent read back anything important (booking references, amounts) before acting on it.

**A licence surprise in legal review.** The LiveKit turn detector and Piper are the two to check. Everything else in the table is permissive.

## Should You Self-Host?

If caller audio can't go to third parties, or you're heading past a few hundred hours of calls a month, yes. The open-source stack in 2026 is good enough that "self-hosted" no longer means "noticeably worse." LiveKit or Pipecat for the plumbing, faster-whisper for ears, an open-weights LLM on vLLM for the brain, Kokoro for the voice, and a real turn detector so it doesn't talk over people.

If you're still finding out whether anyone wants a voice agent at all, rent one first. The pipeline you build on a managed platform maps almost one to one onto the stack above, so nothing you learn is wasted when you move it in-house.

**Where I'd start:** get the LLM layer running first, on a GPU you control, behind an OpenAI-compatible API. It's the heaviest piece, the one the other layers depend on, and the part you'll be tuning the longest.

## Frequently Asked Questions

Can you run a voice AI agent completely offline?

Yes. LiveKit server, LiveKit Agents or Pipecat, Silero VAD, faster-whisper, an LLM on vLLM or Ollama, and Kokoro TTS all run on your own hardware with no external API calls. Only telephony needs an outside SIP trunk if you want real phone numbers.

What GPU do I need to self-host a voice agent?

For a pilot, a single 24 GB card (an NVIDIA A10G on an AWS g5.xlarge, or an L4 on a g6.xlarge) fits an 8B-class LLM plus a Whisper model. Kokoro, Silero VAD and the turn detector run fine on CPU.

Is LiveKit Agents or Pipecat better for self-hosting?

LiveKit Agents if you want the transport, agent framework and SIP bridge from one project, and plan to run more than a handful of concurrent calls. Pipecat if you want fine control over every step of the pipeline and a free choice of transport.

Is the whole LiveKit voice stack open source?

The server, the Agents framework and the SIP service are Apache 2.0. The LiveKit turn detector model is not: it ships under the LiveKit Model License. If you need a fully OSI-licensed stack, Smart Turn (BSD 2-clause) is the swap.

Why can my agent join the room but nobody hears it?

Almost always blocked UDP. LiveKit needs UDP 50000-60000 (or the single mux port 7882) plus TCP 7881 open, and users on strict corporate networks need a TURN server on 443 or 5349 to reach it at all.

Is self-hosting cheaper than a managed voice AI platform?

Only at steady volume. A GPU billed around $1 an hour costs roughly $720 a month running 24/7, whether you take ten calls or ten thousand. Work out your managed per-minute cost first, then find the break-even point.

Can a self-hosted voice agent answer real phone calls?

Yes. LiveKit's open-source SIP service bridges a SIP trunk into a room, so the caller becomes a participant and your agent talks to them the same way it talks to a browser user. You still buy the phone number from a SIP trunk provider.

Which open-source text-to-speech model sounds most natural?

For a self-hosted English agent in 2026, Kokoro-82M is where I'd start: Apache 2.0, 54 voices across 8 languages, small enough for CPU. Piper is lighter still, but its current home is GPL-3.0, so check that works for you.

## Run the LLM Layer of Your Voice Agent on AWS

Launch a pre-configured vLLM server with an OpenAI-compatible API on your own AWS GPU instance, ready for LiveKit Agents or Pipecat to call.

[Get vLLM on AWS Marketplace](https://aws.amazon.com/marketplace/pp/prodview-rvsnp5xf66l4m)

Meetrix Store

vLLM

Serve open models on your own GPU

[Deploy it](https://meetrix.io/store/vllm/)

Meetrix Store New

Deploy what this guide covers, pre-configured.

-    [vLLM Serve open models on your own GPU](https://meetrix.io/store/vllm/)
-    [OpenWebUI A private ChatGPT-style assistant](https://meetrix.io/store/openwebui/)
-    [Llama 4 Scout Mixture-of-Experts, 17B active parameters](https://meetrix.io/store/llama-4-scout/)
-    [Jitsi Meet Self-hosted video calls for 50 to 500 users](https://meetrix.io/store/jitsi-meet/)
-    [RustDesk Remote desktop AMI, a TeamViewer alternative](https://meetrix.io/store/rustdesk/)
-    [Coturn TURN/STUN for WebRTC, no per-minute relay fees](https://meetrix.io/store/coturn/)

[Browse all products](https://meetrix.io/store/)
