WebRTC load testing is about throwing a bunch of fake participants at your media server and keeping an eye on two things at the same time: what the server is doing (CPU, bandwidth) and what each participant would actually feel (packet loss, jitter, round-trip time, frozen video). A server's CPU graph on its own can be misleading. You might see 60% CPU while everyone hears garbled audio, and the opposite can happen too.

This guide isn't tied to any single server. It walks through the method, the Chrome flags that let you fake a camera, a rundown of the tools worth knowing, a Puppeteer script you can tweak for any browser-based app, and the common mistakes that quietly mess up your numbers. If you're running Jitsi and want the step-by-step for that platform, our Jitsi load testing guide with Malleus goes deeper. For background on what you're actually testing, start with WebRTC architecture explained.

The Quick Verdict

Testing LiveKit

Use lk load-test. It simulates publishers and subscribers without a browser, so a single machine can handle a lot.

Testing Jitsi

Use Malleus from Jitsi Meet Torture on a Selenium Grid. We cover that in a separate guide.

Any other server with a real front end

webrtcperf, or a short Puppeteer script like the one below, driving headless Chrome with fake media.

Testing TURN, not the SFU

Force relayed connections in your test clients and watch the TURN server's bandwidth, not the media server's.

3,000 Subscribers LiveKit served from 10 audio publishers at 80% CPU on a 16-core machine, in its own benchmark.
10x Roughly how much cheaper an emulated OpenVidu test participant is than a real Chrome one (about 0.1 vCPU versus about 1). Treat this as a vendor-published figure.
65535 File handles LiveKit tells you to allow with ulimit -n before running its load tester.

Prerequisites

  • A staging copy of the server you want to test, sized like production. Testing against a different instance size only tells you about that size.
  • Separate machines for the load generators. They need their own CPU and bandwidth, and they must not share a host with the server under test.
  • Monitoring on the server before you start: CPU, network in and out, and memory. Without it you'll have client-side numbers and no explanation for them.
  • Node.js 20 or later if you follow the Puppeteer example, and ffmpeg to prepare the test media.

Only test what you own

Aim load tests at infrastructure you run, or have written permission to test. Pointing hundreds of fake participants at someone else's service is an attack, and cloud providers have their own rules about that kind of traffic.

What to Measure

Decide your pass/fail numbers before you run anything. "200 participants for 30 minutes with packet loss under your limit and no frozen video" is a test. "See how it goes" is a demo. The thresholds depend on your product (a webinar tolerates things a doctor's video call does not), so pick yours, write them down, then test.

On the client side, every browser exposes what you need through RTCPeerConnection.getStats(). These are the fields worth collecting:

  • Packet loss: packetsLost against packets received, on inbound streams.
  • Jitter: how uneven packet arrival is. High jitter is what turns into choppy audio.
  • Round-trip time: currentRoundTripTime on the active candidate pair.
  • Frame rate and freezes, plus resolution. A call can have zero packet loss and still drop to 180p because the server is shedding quality.
  • Join time: how long until the first frame renders.

On the server side, watch CPU, network throughput in and out, and memory. An SFU forwards packets instead of decoding them, so the usual bottleneck is bandwidth and packet-forwarding CPU, not video processing. That's also why the load generators, which do decode, end up costing more than the server they're testing.

Fake Cameras With Chrome Flags

Real participants have cameras. Test participants don't, and a headless browser on a server will ask for a camera it doesn't have. Chrome has a set of flags to fix that, documented on webrtc.org's testing page:

  • --use-fake-ui-for-media-stream accepts the camera and microphone permission prompt automatically.
  • --use-fake-device-for-media-stream replaces the camera and microphone with a synthetic test pattern and beep.
  • --use-file-for-fake-video-capture and --use-file-for-fake-audio-capture play a file instead of the test pattern.

The test pattern is fine for a first check. For numbers you can trust, use real footage, because a flat test pattern encodes at a very different bitrate from a person moving around. Chrome accepts .y4m and .mjpeg files as fake video sources, which ffmpeg can make from any clip:

# 10 seconds of 640x360 video at 15 fps, raw y4m (Chrome loops it)
ffmpeg -i source.mp4 -t 10 -vf scale=640:360 -r 15 -pix_fmt yuv420p test.y4m

# 10 seconds of mono 48 kHz audio
ffmpeg -i source.mp4 -t 10 -ac 1 -ar 48000 test.wav

Keep the clip short and the resolution modest. Raw video is enormous (that 10-second clip is already tens of megabytes), and Chrome loops it for the length of the test. Then start a browser with all four flags and open your app:

google-chrome \
  --headless=new \
  --use-fake-ui-for-media-stream \
  --use-fake-device-for-media-stream \
  --use-file-for-fake-video-capture=/opt/media/test.y4m \
  --use-file-for-fake-audio-capture=/opt/media/test.wav \
  https://your-app.example.com/room/load-test

If that one browser joins a room and shows moving video on the other side, the plumbing works. Everything below is about doing it many times at once.

Pick a Load Testing Tool

There are two designs. Some tools drive real browsers, which tests your actual front end but costs a lot of CPU per participant. Others use a lightweight client built on a WebRTC library, which is far cheaper but tests the server and not your web app. OpenVidu's own documentation puts a real Chrome participant at about 1 vCPU and 1 GB of memory, and its emulated participant at about 0.1 vCPU and 0.2 GB. That figure comes from the vendor, so treat it as a rough guide rather than a hard number. A 200-participant browser test, by that math, needs something like 200 vCPU of load generators.

ToolBuilt forHow it makes load
lk load-testLiveKitGo SDK clients, no browser
OpenVidu Load TestOpenVidu 3Emulated clients, or real Chrome and Firefox
webrtcperfAny browser-based WebRTC appPuppeteer-driven headless Chrome
Malleus (Jitsi Meet Torture)Jitsi MeetSelenium Grid browsers
Your own scriptAnything with a web front endHeadless Chrome with fake media

License details matter if you plan to modify a tool or build it into a product. OpenVidu Load Test is Apache 2.0. webrtcperf is AGPL-3.0, which is fine for running tests in-house and something to read carefully before you embed it anywhere. webrtcperf also pushes its metrics to Prometheus through a Pushgateway, so the results land on the same Grafana boards as your server monitoring.

Not sure which server to test in the first place? Our rundown of the best open-source WebRTC media servers covers the options, and LiveKit vs Jitsi compares the two we see most.

Run a LiveKit Test

LiveKit ships its load tester inside its CLI. It uses the Go SDK to act as publishers and subscribers, and its video publishers send simulcast layers at 720p, 360p and 180p, so the server sees realistic fan-out. Raise the file handle limit first, as LiveKit's docs advise, then start small:

ulimit -n 65535

# Large meeting shape: everyone publishes, everyone subscribes
lk load-test \
  --url wss://livekit.example.com \
  --api-key YOUR_KEY --api-secret YOUR_SECRET \
  --room load-test \
  --video-publishers 10 --subscribers 50 \
  --duration 5m

Flags can change between CLI versions, so check lk load-test --help on yours. For scale, LiveKit publishes its own results from a 16-core c2-standard-16 instance on Google Cloud:

ScenarioPublishersSubscribersBandwidth in / outCPU
Large audio room103,0007.3 kBps / 23 MBps80%
Large meeting15015050 MBps / 93 MBps85%
Livestream13,000233 kBps / 531 MBps92%

Source: LiveKit's self-hosting benchmark page. Treat these as one vendor's results on one machine shape, not a capacity promise for yours. A user on LiveKit's GitHub reported hitting only 1,500 subscribers on the same hardware, and a maintainer noted the SFU "may not be as efficient as it used to be as we have added more features." Your mileage will vary.

93MBps out, 150 people on camera

Look at the middle row. A meeting where everyone publishes and everyone subscribes pushed out 93 MBps while the audio room pushed 23 MBps to twenty times as many listeners. The shape of the room moves the load far more than the head count does.

That's why "how many users can it handle?" has no honest one-number answer, and why you should test the room shape you actually run. Our self-hosted LiveKit on AWS guide covers sizing the server side.

A Browser Test for Any App

When no purpose-built tool exists for your stack, a short script gets you most of the way. The idea has three parts: launch headless Chrome with the fake media flags, wrap RTCPeerConnection before the app loads so you can reach its connections, then poll getStats() on a timer and print what you find.

// load-test.mjs   (Node 20+, npm install puppeteer)
// usage: node load-test.mjs https://your-app.example.com/room/load-test 20 180
import puppeteer from 'puppeteer';

const [, , url, count = '10', seconds = '120'] = process.argv;

// Runs inside every page before the app does, so we can reach its connections later.
const hook = () => {
  const Original = window.RTCPeerConnection;
  window.__pcs = [];
  window.RTCPeerConnection = function (...args) {
    const pc = new Original(...args);
    window.__pcs.push(pc);
    return pc;
  };
  window.RTCPeerConnection.prototype = Original.prototype;
};

async function join(i) {
  const browser = await puppeteer.launch({
    headless: true,
    args: [
      '--no-sandbox', // needed in most containers
      '--use-fake-ui-for-media-stream',
      '--use-fake-device-for-media-stream',
      '--use-file-for-fake-video-capture=/opt/media/test.y4m',
      '--use-file-for-fake-audio-capture=/opt/media/test.wav',
    ],
  });
  const page = await browser.newPage();
  await page.evaluateOnNewDocument(hook);
  // How a name is passed in depends on your app. Adjust this line.
  await page.goto(url + '?name=load-' + i, { waitUntil: 'load' });
  return { browser, page };
}

function sample(page) {
  return page.evaluate(async () => {
    const out = { lost: 0, received: 0, jitter: 0, rtt: 0 };
    for (const pc of window.__pcs || []) {
      const stats = await pc.getStats();
      stats.forEach((r) => {
        if (r.type === 'inbound-rtp') {
          out.lost += r.packetsLost || 0;
          out.received += r.packetsReceived || 0;
          out.jitter = Math.max(out.jitter, r.jitter || 0);
        }
        if (r.type === 'candidate-pair' && r.state === 'succeeded' && r.nominated) {
          out.rtt = Math.max(out.rtt, r.currentRoundTripTime || 0);
        }
      });
    }
    return out;
  });
}

const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
const clients = [];

for (let i = 0; i < Number(count); i++) {
  clients.push(await join(i));
  await sleep(2000); // ramp up, never join everyone at once
}

console.log('seconds,clients,loss_pct,max_jitter_ms,max_rtt_ms');
const started = Date.now();
while (Date.now() - started < Number(seconds) * 1000) {
  const rows = await Promise.all(clients.map((c) => sample(c.page)));
  const lost = rows.reduce((n, r) => n + r.lost, 0);
  const received = rows.reduce((n, r) => n + r.received, 0);
  const jitter = Math.max(...rows.map((r) => r.jitter));
  const rtt = Math.max(...rows.map((r) => r.rtt));
  const loss = (100 * lost) / Math.max(1, lost + received);
  console.log([
    Math.round((Date.now() - started) / 1000),
    clients.length,
    loss.toFixed(2),
    (jitter * 1000).toFixed(0),
    (rtt * 1000).toFixed(0),
  ].join(','));
  await sleep(10000);
}

await Promise.all(clients.map((c) => c.browser.close()));

Be honest about what this is: a starting point. I haven't run it against your app. It assumes the app creates its peer connections in the page itself, and the line that passes a display name is a placeholder you'll have to change. Join flows that need a login, a lobby or a click on a "Join" button need a few lines of Puppeteer to get through, using page.click and page.waitForSelector.

It also doesn't scale forever. Each browser is a full Chrome, so at roughly 1 vCPU per participant you'll outgrow one machine quickly. Run the script on several load generators with a different count on each, and add up the results. If you outgrow even that, webrtcperf already does this job with more polish.

Reading the Results

Plot server CPU, server bandwidth and client packet loss on one timeline against the number of participants. You're looking for the knee: the point where loss or jitter starts climbing while participants keep being added. That participant count, not the point where the server falls over, is your real capacity.

One research result is worth knowing before you trust CPU as your only signal. OpenVidu's team compared Kurento, mediasoup and Pion and reported that for Kurento and Pion, CPU alone tracked quality degradation well, while for mediasoup it did not, and additional WebRTC metrics were needed because quality could stay good even at high CPU. Their research page has the papers. The practical lesson: always collect the client metrics, whatever server you run.

Look at the worst participants, not only the average. An average loss of 0.5% can hide a handful of people at 8%, and those are the ones who file support tickets. When something looks wrong in a single session, debugging WebRTC applications shows how to dig into it with chrome://webrtc-internals.

Test the TURN Path

Most load tests run on a clean network where every call connects directly, so they never touch your TURN server at all. Then launch day arrives, a corporate firewall forces thousands of users through relays, and the load is completely different. Relayed media passes through your TURN server in both directions, which means its bandwidth bill and CPU matter as much as the SFU's.

To test it, force relayed connections in the test clients:

// In the page under test, force every call through TURN
const pc = new RTCPeerConnection({
  iceServers: [{ urls: 'turn:turn.example.com:3478', username: 'user', credential: 'pass' }],
  iceTransportPolicy: 'relay',
});

Run the same test again with everyone relayed and watch the TURN server. If you haven't set one up, our Coturn developer guide deploys one on AWS, and STUN vs TURN vs ICE explains when relays kick in.

Mistakes That Skew the Numbers

  • The load generator is the bottleneck. If the machine running the browsers is at full CPU, you're measuring the generator. Watch its CPU too, and add machines before it saturates.
  • Everyone joins at once. Real rooms fill gradually. A thundering herd of joins tests your signalling server and ICE handling, which is a separate test. Ramp up, as the script above does.
  • Generators in the same rack as the server. No latency, no loss, flattering results. Put them in the same region at least, and ideally somewhere with a real network path in between.
  • A test clip that is unlike real video. A static image encodes at a fraction of the bitrate of a person moving. Check the bitrate your fake participants actually send and compare it with a real session.
  • Too short. Memory leaks and slow degradation show up after 30 minutes, not 3. Run a ramp test to find the knee, then a long soak test at a participant count just below it.
  • Reporting only averages. Use percentiles and the worst few participants.
  • Forgetting the cost. A big test moves a lot of bandwidth. Check what your cloud charges for traffic before a 3,000-subscriber run, as our Jitsi self-hosting cost breakdown shows for one platform.

Where I'd Start

Just getting started

One browser with the fake media flags, joining a room against your staging server. If that works, you have the foundation for everything else.

Running LiveKit

Use lk load-test with a room shape that matches your product, then ramp the counts up in steps.

Running Jitsi

Follow our Malleus guide, then compare against load balancing and videobridge monitoring to read the results.

Before a launch

A ramp test to find the knee, a soak test just below it, and a relayed-only run against TURN. Three tests, three different failure modes.

Write the pass or fail numbers down before the first run. If you decide what counts as a failure after you've seen the graphs, you'll talk yourself into whatever the graphs say.

Frequently Asked Questions

What is WebRTC load testing?

It's the practice of simulating many participants against a media server to see how it performs and what quality each participant experiences.

How do I simulate WebRTC users without cameras?

Use Chrome's fake media flags (--use-fake-device-for-media-stream, etc.) to feed a test pattern or a file instead of a real camera. .y4m and .mjpeg files both work as fake video sources.

Do I need real browsers for a WebRTC load test?

Not always. Lightweight clients built on WebRTC libraries are much cheaper, but if you need to test your actual front end, real browsers are required.

How many users can a WebRTC server handle?

There's no single number. It depends on the room shape, the server hardware, and the network. You have to test your specific scenario. LiveKit's own benchmarks show 3,000 audio subscribers or 150 video participants on a 16-core machine, but a user reported only hitting 1,500 on the same setup.

Which tool should I use for WebRTC load testing?

It depends on your server. For LiveKit, use lk load-test. For Jitsi, use Malleus. For anything else with a web front end, try webrtcperf or a Puppeteer script.

How do I load test a TURN server?

Force relayed connections in your test clients and monitor the TURN server's bandwidth and CPU while running the same load test.

Need a Server to Test Against?

Meetrix's pre-configured Jitsi Meet servers for AWS and Google Cloud come in sizes for 50 to 500 users, so you can load test a production-shaped server you control.

Explore Jitsi Meet on the Store