> Source: https://meetrix.io/blogs/webrtc-monitoring-production/
> Markdown copy of that page. Cite the URL above, not this file.

Development

# Monitoring WebRTC in Production: Metrics and Alerts

[By Shalomi Umeshika](https://meetrix.io/blogs/authors/shalomi-umeshika/) • October 2, 2026 • 14 min read

Monitoring WebRTC in production means keeping an eye on three layers at once: what users actually feel in their browsers, what your media server is up to, and what your TURN server is relaying. Most teams start with the middle one because it is the easiest to scrape, then discover the hard way that a healthy-looking server graph tells you nothing about the person whose audio keeps dropping.

This guide stays platform-neutral. It covers what to pull from the browser, which server and TURN metrics are worth collecting, how to wire it all into Prometheus and Grafana, and the short list of alerts that actually deserve to wake someone up. If you want the narrower versions, our [Jitsi videobridge monitoring on CloudWatch](https://meetrix.io/blogs/monitor-jitsi-videobridge-cloudwatch/) guide covers one bridge on AWS, and [debugging WebRTC applications](https://meetrix.io/blogs/debugging-webrtc-applications/) shows how to dig into a single session by hand. And if you have not load tested yet, do that first: [WebRTC load testing](https://meetrix.io/blogs/webrtc-load-testing/) gives you the baseline that makes production numbers mean something.

The Quick Verdict

Starting from nothing

Collect `getStats()` deltas from the browser first. It is the only layer that shows what users feel, and it needs no server access.

Running Jitsi

Scrape the videobridge's `/metrics` and watch `stress_level` alongside participants and bitrate. Be aware that JVB's documentation still lists some conference size bucket metrics that were actually removed from recent builds. The JSON stats endpoint no longer includes them, and histogram-style conference size data now lives in the Prometheus output instead.

Running coturn

Start it with `--prometheus`, and track relay share from the client side to see when users fall back to TURN. Coturn 4.14.0 removed the dependency on the old prometheus-client-c library and rewrote the Prometheus exporter from scratch, but it still relies on libmicrohttpd.

Need per-call forensics

Add rtcstats on top, which stores full stats dumps for the sessions you choose, instead of only aggregates. The original fippo repository was archived in September 2025, but Jitsi maintains an active fork that is still getting commits as of April 2026.

5 s Default period of Jitsi Videobridge's statistics reports, adjustable with videobridge.stats.interval

9641 Default port for coturn's Prometheus metrics, once you start it with the --prometheus flag

0 to 1 Range of the videobridge's stress\_level, where 1 means full capacity (values above 1 are allowed)

Prerequisites

-   A Prometheus server and Grafana, or a hosted equivalent. Our [Grafana developer guide](https://meetrix.io/blogs/grafana-developer-guide/) deploys Grafana on AWS if you need one.
-   Access to the media server's metrics interface. On the Jitsi Videobridge, that means enabling its REST and Colibri APIs (in a Docker setup, `JVB_ENABLE_APIS=rest,colibri`).
-   The ability to ship a small piece of JavaScript with your front end, plus a tiny collector service to receive what it sends.
-   A baseline from a load test. Alert thresholds copied from a blog post are guesses; thresholds from your own test are measurements.

## The Three Layers

Each layer answers a different question, and each can be fine while another is on fire.

-   **Client quality** answers "is the call good for the person on it?" Packet loss, jitter, round-trip time, frozen video, and whether the connection came up at all.
-   **The media server** answers "is the box coping?" CPU, bandwidth, participants, conferences, and for Jitsi a purpose-built load number.
-   **TURN** answers "who is being relayed, and what does that cost?" Relayed calls pass through your TURN server in both directions, so they cost bandwidth and show up differently in quality than direct ones. Our [STUN vs TURN vs ICE](https://meetrix.io/blogs/stun-vs-turn-vs-ice-webrtc-nat-traversal/) article explains when relays kick in.

If you only build one, build the client layer. A server can sit at modest CPU while a regional network problem ruins calls, and only the browsers will tell you.

## Client-Side Telemetry

Every browser exposes quality data through `RTCPeerConnection.getStats()`. The fields worth collecting, per the [W3C WebRTC statistics spec](https://www.w3.org/TR/webrtc-stats/):

-   `packetsLost` and packets received, on inbound streams.
-   `jitter`, measured in seconds.
-   `freezeCount` and `totalFreezesDuration` for video. A freeze is counted when the gap between two rendered frames is at least the larger of three times the average frame duration, or the average plus 150 ms.
-   `qualityLimitationReason` on outbound video, which says whether bandwidth or CPU is holding quality down. The values are "none", "cpu", "bandwidth", or "other".
-   `currentRoundTripTime` on the active candidate pair.
-   The local candidate's `candidateType`, where `relay` means the call goes through TURN.

Sample every ten seconds or so, send the change since the last sample, and keep each message tiny. Here is a starting point you can drop next to your peer connection code:

```javascript
// rtc-telemetry.js
// Call watch(pc, { region, platform }) for every RTCPeerConnection you create.
const ENDPOINT = '/api/rtc-metrics';
const INTERVAL_MS = 10000;

export function watch(pc, labels) {
  let last = { lost: 0, received: 0, freezes: 0 };

  const send = (body) =>
    navigator.sendBeacon(ENDPOINT, JSON.stringify({ ...labels, t: Date.now(), ...body }));

  // Failures to connect never reach getStats, so report state changes too.
  pc.addEventListener('connectionstatechange', () => {
    if (pc.connectionState === 'connected' || pc.connectionState === 'failed') {
      send({ event: pc.connectionState });
    }
  });

  const timer = setInterval(async () => {
    if (pc.connectionState === 'closed') return clearInterval(timer);

    const now = { lost: 0, received: 0, freezes: 0, jitter: 0, rtt: 0, relay: false, limit: 'none' };
    const stats = await pc.getStats();
    const byId = new Map();
    stats.forEach((r) => byId.set(r.id, r));

    stats.forEach((r) => {
      if (r.type === 'inbound-rtp') {
        now.lost += r.packetsLost || 0;
        now.received += r.packetsReceived || 0;
        now.freezes += r.freezeCount || 0;
        now.jitter = Math.max(now.jitter, r.jitter || 0);
      }
      if (r.type === 'outbound-rtp' && r.qualityLimitationReason && r.qualityLimitationReason !== 'none') {
        now.limit = r.qualityLimitationReason;
      }
      if (r.type === 'candidate-pair' && r.state === 'succeeded' && r.nominated) {
        now.rtt = r.currentRoundTripTime || 0;
        const local = byId.get(r.localCandidateId);
        now.relay = !!local && local.candidateType === 'relay';
      }
    });

    // The counters are cumulative, so send the change since the last sample.
    // Clamp at zero: a stream that ends makes the running total drop.
    send({
      event: 'sample',
      lost: Math.max(0, now.lost - last.lost),
      received: Math.max(0, now.received - last.received),
      freezes: Math.max(0, now.freezes - last.freezes),
      jitter_ms: Math.round(now.jitter * 1000),
      rtt_ms: Math.round(now.rtt * 1000),
      relay: now.relay,
      limit: now.limit,
    });
    last = { lost: now.lost, received: now.received, freezes: now.freezes };
  }, INTERVAL_MS);
}
```

Two details in there matter more than they look. The stats are cumulative counters, so the code sends deltas, clamped at zero because a stream that ends makes the running total drop. And it reports `connectionstatechange` separately, for the reason in the next panel. I have not run this against your app, so treat it as a template and test it with a real call before shipping it.

0samples from users who never connected

Client telemetry has a blind spot built in. A browser that fails to connect never gets as far as sending quality stats, so your loss and jitter graphs only describe the people it worked for. The angriest users are missing from your data.

Count failures separately: report the `failed` connection state, and compare attempts on your signalling server with connections that actually came up.

On the receiving end, a small collector turns each sample into Prometheus metrics: a histogram for loss ratio, jitter and round-trip time, and counters for samples, relayed samples and connection events. Label them by coarse, low-cardinality values such as region and platform. Do not label by user, room or session. That is the fastest way to bring a Prometheus server down.

If you want to keep full per-call dumps for the sessions you pick, rather than only aggregates, look at [rtcstats](https://github.com/rtcstats/rtcstats), an open-source client SDK and collector. The original fippo/rtcstats-server repository was archived in September 2025, but the Jitsi fork is active and still receiving commits as of April 2026.

## Media Server and TURN Metrics

**Jitsi Videobridge.** The bridge exports statistics as JSON at `/colibri/stats` and in Prometheus format at `/metrics`, both on its private interface. The report period defaults to 5 seconds and can be changed with `videobridge.stats.interval` in `jvb.conf`. The numbers worth graphing are `stress_level` (0 is idle, 1 is full capacity), `participants`, `conferences`, `bit_rate_download` and `bit_rate_upload`, `packet_rate_download` and `packet_rate_upload`, and `endpoints_sending_audio` and `endpoints_sending_video`. The full list is in the [videobridge statistics documentation](https://github.com/jitsi/jitsi-videobridge/blob/master/doc/statistics.md).

One thing to know: JVB has been migrating toward a shim layer that translates modern Prometheus metrics into legacy formats for backward compatibility. This means metric names in the Prometheus output can differ slightly from the JSON keys, and some metrics that used to appear in JSON are now only available as Prometheus histograms. Check `/metrics` on your version before writing queries. Running more than one bridge? Our [Jitsi load balancing](https://meetrix.io/blogs/jitsi-meet-load-balancing/) guide shows how the bridges are selected, which is what these numbers feed into.

**LiveKit.** Set `prometheus_port` in the server config and LiveKit serves metrics on that port at `/metrics`. The port is not set by default, so you have to enable it explicitly. Recent versions added Prometheus metrics for join latency and peer connection state, which are useful for catching connection issues early. The [deployment docs](https://docs.livekit.io/transport/self-hosting/deployment/) show the setting, and our [self-hosted LiveKit on AWS](https://meetrix.io/blogs/self-host-livekit-aws/) guide covers sizing the server these metrics describe.

**coturn.** Start it with the `--prometheus` flag (or `--prometheus-tls` for HTTPS) and it serves metrics on port 9641 at `/metrics`. Coturn 4.14.0 removed the dependency on the old prometheus-client-c library and rewrote the exporter from scratch, but it still needs libmicrohttpd at build time (see its [Prometheus documentation](https://github.com/coturn/coturn/blob/master/docs/Prometheus.md)). Some distribution packages leave Prometheus out, so confirm your build has it. The documented metrics are mostly about unauthenticated request handling, allocation lifecycles, and traffic volume, labeled by realm and optionally username. For relay bandwidth, add host-level network metrics from the TURN machine, and use the client-side relay share from earlier. If you have not deployed coturn yet, see our [Coturn developer guide](https://meetrix.io/blogs/coturn-developer-guide/).

One coturn warning: the `--prometheus-username-labels` flag labels traffic metrics with client usernames, and it is off by default because it can leak memory with ephemeral usernames such as those from the TURN REST API. Leave it off.

## Wire It Into Prometheus

The scrape config is short. Replace the targets with your own hosts and ports:

```yaml
scrape_configs:
  - job_name: jvb
    metrics_path: /metrics
    static_configs:
      # the private HTTP interface of each videobridge
      - targets: ['jvb1.internal:8080', 'jvb2.internal:8080']

  - job_name: coturn
    static_configs:
      # coturn started with --prometheus listens on port 9641
      - targets: ['turn1.internal:9641']

  - job_name: rtc-collector
    static_configs:
      # your own collector that turns client samples into metrics
      - targets: ['collector.internal:9100']
```

The `rtc-collector` job is the small service from the client section. It receives the browser samples and exposes them as metrics for Prometheus to scrape. Keep it boring: parse, validate, increment, done. Drop anything malformed, because it is a public endpoint that anyone can post to.

If you also keep long-term data in a time-series database, our [TimescaleDB with Grafana and Prometheus](https://meetrix.io/blogs/timescaledb-grafana-prometheus-aws/) article shows one way to do that.

## Dashboards Worth Building

Resist the urge to graph everything. Five panels cover most incidents:

1.  Participants and conferences over time, so every other graph has context.
2.  Client p95 packet loss and jitter, split by region.
3.  Connection success rate, from the `connected` and `failed` events.
4.  Media server load: `stress_level` and bitrate in and out for each bridge.
5.  Relay share, and TURN bandwidth next to it.

Put the user-facing panels at the top and the server panels below. When a page comes in, the first question is "are users hurting?", and the second is "which layer?".

## Alerts That Matter

Alert on user pain and on things that are about to cause it, not on every metric that moves. A short list:

-   **Connection failures rising.** The most direct sign that users cannot join. Page on it.
-   **Client p95 loss or freezes up.** Quality is degrading for real people. Page on it, scoped by region so you can see where.
-   **Relay share jumps compared with last week.** A sudden rise usually means UDP is being blocked somewhere or a direct path broke. Open a ticket.
-   **Bridge `stress_level` near 1.** The bridge is close to full capacity, so add capacity before the next big meeting.
-   **TURN down, or its certificate close to expiry.** If you serve TURN over TLS, an expired certificate quietly strands the users who depend on it most.

Example rules for the first and third, using the metrics the collector exposes. The numbers are placeholders, and the comparison with last week avoids needing an absolute threshold at all:

```yaml
groups:
  - name: webrtc
    rules:
      - alert: ClientLossHigh
        # 0.05 is a placeholder. Start from your own baseline.
        expr: histogram_quantile(0.95, sum by (le, region) (rate(webrtc_client_loss_ratio_bucket[10m]))) > 0.05
        for: 10m
        labels: { severity: page }

      - alert: ConnectionFailuresUp
        expr: |
          sum(rate(webrtc_connection_events_total{event="failed"}[10m]))
            / sum(rate(webrtc_connection_events_total[10m])) > 0.02
        for: 10m
        labels: { severity: page }

      - alert: RelayShareJumped
        expr: |
          sum(rate(webrtc_samples_relay_total[15m])) / sum(rate(webrtc_samples_total[15m]))
            > 2 * (sum(rate(webrtc_samples_relay_total[15m] offset 7d))
                   / sum(rate(webrtc_samples_total[15m] offset 7d)))
        for: 15m
        labels: { severity: ticket }
```

## Privacy and Cost

Stats are less innocent than they look. Candidate objects contain IP addresses, so send only the candidate type, as the script above does. Do not attach user names, emails or room names to metrics, both for privacy and because those labels explode Prometheus's memory. Tell users in your privacy policy that you collect connection quality data, and check what your local rules require before shipping it.

On cost, ten-second samples from every participant add up. If volume is high, sample a share of sessions, or send every thirtieth second instead. Aggregated histograms are cheap to store. Raw per-session samples are not.

## Common Mistakes

-   **Monitoring only the server.** CPU looks fine while a regional ISP problem wrecks calls.
-   **Averages instead of percentiles.** An average loss of 0.5% can hide a group of users at 8%.
-   **Alerting on CPU alone.** It tells you the box is busy, not that anyone is unhappy.
-   **High-cardinality labels.** User or room labels make your monitoring fall over before your product does.
-   **Counting only survivors.** Missing the users who failed to connect makes the dashboards flatteringly green.
-   **No baseline.** Without a load test or a few weeks of history you do not know whether 3% loss is normal for your users.

## Where I'd Start

This week

Ship the client sampler to a small share of sessions and chart connection success rate and p95 loss. That alone beats most setups.

Next

Scrape your media server and TURN, and put server load on the same dashboard as the client numbers.

After a month of data

Set alert thresholds from the baseline you now have, then trim any alert that fired without anyone caring.

When an incident is unexplained

Add per-session dumps with rtcstats for the affected users, and read them with the techniques in our debugging guide.

An alert nobody acts on is worse than no alert, because it teaches the team to ignore the pager. Start with three that you would genuinely get out of bed for, and add more only when an incident proves one was missing.

## Frequently Asked Questions

What should I monitor in a production WebRTC service?

Three layers: client-side quality (packet loss, jitter, RTT, freezes, connection success), media server load (CPU, bandwidth, participants, stress level), and TURN relay share and bandwidth.

How do I collect WebRTC stats from real users?

Use RTCPeerConnection.getStats() in the browser, sample every ten seconds, and send deltas to a small collector that turns them into Prometheus metrics.

Which metrics does Jitsi Videobridge expose?

stress\_level, participants, conferences, bitrate and packet rates, and endpoints sending audio or video, among others. They are available as JSON at /colibri/stats and as Prometheus metrics at /metrics. Note that the JSON endpoint and the Prometheus output do not always list the same metrics, and some conference size data is now only in the Prometheus histograms.

How do I monitor a TURN server?

Start coturn with --prometheus, scrape port 9641, and track relay share from the client side. Add host-level network metrics for relay bandwidth.

What WebRTC alerts are worth setting?

Connection failures rising, client p95 loss or freezes up, relay share jumping, bridge stress near capacity, and TURN down or certificate expiring.

Does collecting WebRTC stats create privacy problems?

Yes. Candidate objects contain IP addresses, so send only the candidate type. Never attach user names or room names to metrics.

## Run Your Own Video and TURN Servers

Meetrix's pre-configured Jitsi Meet and Coturn servers for AWS and Google Cloud give you production-shaped infrastructure to monitor, and the metrics endpoints described here work on both.

[Explore Jitsi Meet on the Store](https://meetrix.io/store/jitsi-meet/)

Meetrix Store

Jitsi Meet

Self-hosted video calls for 50 to 500 users

[Deploy it](https://meetrix.io/store/jitsi-meet/)

Meetrix Store New

Deploy what this guide covers, pre-configured.

-    [Jitsi Meet Self-hosted video calls for 50 to 500 users](https://meetrix.io/store/jitsi-meet/)
-    [Coturn TURN/STUN for WebRTC, no per-minute relay fees](https://meetrix.io/store/coturn/)
-    [Grafana Dashboards and alerts for Prometheus and more](https://meetrix.io/store/grafana/)
-    [Mattermost Team chat, a self-hosted Slack alternative](https://meetrix.io/store/mattermost/)
-    [RustDesk Remote desktop AMI, a TeamViewer alternative](https://meetrix.io/store/rustdesk/)
-    [OpenVPN Encrypted remote access, no per-user fees](https://meetrix.io/store/openvpn/)

[Browse all products](https://meetrix.io/store/)
