Skip to content
CERALIVEPublic
forked from irlserver/srtla

About

SRT link aggregation — bonds multiple network interfaces (LTE, WiFi) for resilient live video streaming; upstream-aligned fork with irlserver improvements

Resources

Stars

1 star

Watchers

0 watching

Forks

 
 

Repository files navigation

SRTLA Receiver (srtla_rec)

Overview

srtla_rec is an SRT transport proxy with link aggregation. SRTLA is designed to transport SRT traffic over multiple network links for capacity aggregation and redundancy. Traffic is balanced dynamically depending on network conditions. The primary application is bonding mobile modems for live streaming.

Note: This is a fork of the original SRTLA implementation by BELABOX. The original server component (srtla_rec) was marked as unsupported by BELABOX.

Features

  • Support for link aggregation across multiple network connections
  • Automatic management of connection groups and individual connections
  • Robust error handling and timeouts for inactive connections
  • Logging of connection details for easy diagnostics
  • Improved load balancing through ACK throttling
  • Connection recovery mechanism for temporary network issues

Requirements

  • C++11 compatible compiler
  • CMake for the build process
  • spdlog library
  • argparse library

Assumptions and Prerequisites

SRTLA assumes that:

  • Data is streamed from an SRT sender in caller mode to an SRT receiver in listener mode
  • To benefit from link aggregation, the sender should have 2 or more network links to the SRT listener (typically internet-connected modems)
  • The sender needs to have source routing configured, as SRTLA uses bind() to map UDP sockets to specific connections

Installation

# Clone the repository
git clone https://github.com/OpenIRL/srtla.git
cd srtla

# Build with CMake
mkdir build
cd build
cmake ..
make

Usage

srtla_rec runs as a proxy between SRTla clients and an SRT server:

./srtla_rec [OPTIONS]

Command Line Options

  • --srtla_port PORT: Port to bind the SRTLA socket to (default: 5000)
  • --srt_hostname HOST: Hostname of the downstream SRT server (default: 127.0.0.1)
  • --srt_port PORT: Port of the downstream SRT server (default: 4001)
  • --verbose: Enable verbose logging (default: disabled)
  • --debug: Enable debug logging (default: disabled)
  • --metrics_port PORT: Serve Prometheus metrics over HTTP on this port (default: 0, disabled)
  • --metrics_bind ADDR: Numeric address the metrics endpoint binds to (default: 127.0.0.1)
  • --metrics_detail: Also export per-connection metrics (default: disabled)

Example

./srtla_rec --srtla_port 5000 --srt_hostname 192.168.1.10 --srt_port 4001 --verbose

How It Works

  1. srtla_rec creates a UDP socket for incoming SRTLA connections.
  2. Clients register with srtla_rec and create connection groups.
  3. Multiple connections can be added to a group.
  4. Data is received across all connections and forwarded to the SRT server.
  5. ACK packets are sent across all connections for timely delivery.
  6. Inactive connections and groups are automatically cleaned up.

Technical Details

SRTLA implements a protocol for packet transmission over multiple network connections, aggregating the data and making it available to the SRT protocol. The implementation is based on the following core mechanisms:

  1. Connection Group Management: The software organizes connections into groups, with each group corresponding to an SRT stream. This enables support for multiple simultaneous SRTLA senders with a single receiver.

  2. Packet Tracking: The code tracks received packets with sequence numbers and periodically sends SRTLA-ACK packets back to confirm receipt.

  3. Two-phase Registration Process:

    • Sender (conn 0): SRTLA_REG1 (contains sender-generated random ID)
    • Receiver: SRTLA_REG2 (contains full ID with receiver-generated values)
    • Sender (conn 0): SRTLA_REG2 (with full ID)
    • Receiver: SRTLA_REG3
    • Additional connections follow a similar pattern
  4. Error Handling: The receiver can send error responses:

    • SRTLA_REG_ERR: Operation temporarily failed
    • SRTLA_REG_NGP: Invalid ID, group must be re-registered
  5. Connection Cleanup: Inactive connections and groups are automatically cleaned up after a configurable timeout (default: 10 seconds).

  6. Load Balancing through ACK Throttling: The server controls ACK frequency to influence the client's connection selection without requiring client-side modifications.

  7. Connection Recovery Mechanism: Connections that show signs of recovery after temporary outages are given a chance to stabilize again.

The implementation uses epoll for event-based network I/O, allowing efficient handling of multiple simultaneous connections.

Enhanced Load Balancing and Recovery

This version of SRTLA includes improvements to address two key issues in the original implementation:

Problem 1: Connections with Issues Had No Recovery Path

In the original implementation, connections with temporary problems were completely disabled. In this enhanced version:

  • Connections showing signs of recovery enter a "recovery mode"
  • These connections receive more frequent keepalive packets for a set period (5 seconds)
  • After successful recovery, they are fully reactivated for data transmission
  • Recovery attempts are abandoned after a certain time if unsuccessful

This functionality allows connections to "heal" after brief disruptions (e.g., due to network issues) rather than remaining completely disabled.

Problem 2: Unbalanced Connection Utilization

In the original implementation, load was unevenly distributed across available connections. The new implementation:

  • Introduces a monitoring and evaluation system for connection quality
  • Checks connection quality every 5 seconds based on:
    • Bandwidth (kbits/s) and performance ratio (actual vs expected bandwidth)
    • Packet loss rate
    • Connection performance relative to median bandwidth
  • Assigns error points to each connection based on these metrics
  • Calculates a quality weight for each connection (10% to 100%)
  • Controls ACK packet frequency based on connection quality
    • Good connections receive ACKs more frequently
    • Poor connections receive ACKs less frequently
  • Indirectly influences the window size in the client and thus connection selection

The result is better data distribution, with more stable connections carrying more load than problematic ones, without requiring client modifications.

Technical Implementation Details

ACK Throttling

The central innovation of this solution is ACK throttling for load distribution. It's based on the following principles:

  1. The SRT/SRTLA client (srtla_send) selects connections based on a score derived from the window size and in-flight packets.
  2. The window size in the client is adjusted when ACKs are received.
  3. By selectively throttling ACK frequency, we can indirectly control how quickly the window grows in the client.
  4. This causes the client to prefer better connections without requiring changes to the client code.

Connection Quality Assessment

Connection quality is assessed by measuring and analyzing:

  • Bandwidth Performance: The system calculates a performance ratio by comparing actual bandwidth to expected bandwidth. Poor performance relative to expectations leads to more error points
  • Packet Loss: Higher loss rates lead to more error points
  • Dynamic Bandwidth Evaluation: Connections are evaluated against either median bandwidth (for good connections) or minimum threshold (for poor connections)
  • Grace Period: New connections receive a 10-second grace period before penalties are applied

The weight levels are:

  • 100% (WEIGHT_FULL): Optimal connection
  • 85% (WEIGHT_EXCELLENT): Excellent connection
  • 70% (WEIGHT_DEGRADED): Slightly impaired connection
  • 55% (WEIGHT_FAIR): Fair connection
  • 40% (WEIGHT_POOR): Severely impaired connection
  • 10% (WEIGHT_CRITICAL): Critically impaired connection

Recovery Mechanism

The recovery functionality works as follows:

  1. A connection that receives data again after being marked inactive is placed in recovery mode
  2. In this mode, keepalive packets are sent more frequently (every 1 second)
  3. If the connection remains stable for a short period (5 seconds), it is considered recovered
  4. If recovery does not occur within the time window, the recovery attempt is aborted

Configuration Parameters

The following parameters can be adjusted to optimize behavior:

  • KEEPALIVE_PERIOD: Interval for keepalive packets during recovery (1 second)
  • RECOVERY_CHANCE_PERIOD: Period during which a connection can attempt to recover (5 seconds)
  • CONN_QUALITY_EVAL_PERIOD: Interval for evaluating connection quality (5 seconds)
  • MIN_ACCEPTABLE_TOTAL_BANDWIDTH_KBPS: Minimum total bandwidth for acceptable streaming quality (1000 kbps)
  • GOOD_CONNECTION_THRESHOLD: Threshold for considering a connection "good" (50% of max bandwidth)
  • CONNECTION_GRACE_PERIOD: Grace period in seconds before applying penalties (10 seconds)
  • Various weight levels (WEIGHT_FULL, WEIGHT_EXCELLENT, WEIGHT_DEGRADED, WEIGHT_FAIR, WEIGHT_POOR, WEIGHT_CRITICAL)

Limitations

  • The error point thresholds are static and could be dynamically adjusted to better adapt to different network situations
  • Performance ratio calculations are based on bandwidth expectations that may need tuning for different network environments

SRT Configuration Recommendations

The sender should implement congestion control using adaptive bitrate based on the SRT SRTO_SNDDATA size or measured RTT.

Monitoring

--metrics_port serves Prometheus text exposition over HTTP. Every path returns the same body.

./srtla_rec --metrics_port 9997
curl http://127.0.0.1:9997/metrics

The main epoll loop answers the scrape, so this costs no extra thread and takes no locks.

The endpoint listens on 127.0.0.1 by default because nothing authenticates it. --metrics_bind :: binds every interface for both IPv4 and IPv6, which is what a container needs, and --metrics_bind 0.0.0.0 binds IPv4 only. Put a firewall in front of either one. The address has to be numeric: --metrics_bind localhost logs an error at startup and leaves the endpoint off while the receiver keeps running.

Always exported: traffic counters (srtla_packets_received_total, srtla_forwarded_packets_total, srtla_srt_packets_received_total and others), registration outcomes (srtla_group_registrations_total, srtla_group_registrations_rejected_total{reason}), teardowns (srtla_groups_removed_total{reason}), auth throttling (srtla_auth_failures_total, srtla_auth_sources_blocked), NAK handling, connection recovery, send errors, and live gauges for groups and connections.

Divide srtla_packets_received_total by srtla_recv_batches_total for the receive loop fill ratio. A climbing ratio means recvmmsg() is returning fuller batches as the loop approaches saturation. That is the first number worth checking when a receiver feels slow.

--metrics_detail adds per-connection series (srtla_conn_*: bytes, packets, loss, weight, error points, RTT, window, in-flight, sender bitrate, idle time) labelled group, the group's local SRT port, and remote, the client address. These are what explain bonding behavior instead of just reporting it. Every reconnect creates a new remote label value, so watch series cardinality before enabling this on a busy receiver.

docs/grafana-dashboard.json covers every exported metric across 20 panels. Import it and pick your Prometheus data source. The per-connection panels sit in a collapsed row at the bottom and stay empty unless --metrics_detail is on.

Socket Information

srtla_rec creates information files about active connections under /tmp/srtla-group-[PORT]. These files contain the client IP addresses connected to a specific socket.

CERALIVE layer

This checkout is CERALIVE's fork of irlserver/srtla. Everything above this section is upstream's README, unchanged. src/ is byte-identical to the upstream commit the fork sits on; CERALIVE adds only the build policy, CI and test layer described here. See AGENTS.md for the maintainer contract.

Receiver only

cmake --install puts exactly one binary in bin/: srtla_rec. The C srtla_send target still builds (upstream keeps it and CI checks that it links), but it is not installed and it is not what CERALIVE devices run. The device-side sender is the Rust srtla-send-rs, which is CLI-compatible with the C sender (SRT_LISTEN_PORT SRTLA_HOST SRTLA_PORT IPS_FILE) and speaks the same wire protocol, extended keepalive included. The Verify receiver-only install step in build-check.yml fails the build if anything else lands in bin/.

No releases: build from source at a git ref

This repo cuts no GitHub releases, no .deb, no tarballs and no tags. Every consumer builds the receiver from source at a commit or branch it chooses:

git clone https://github.com/CERALIVE/srtla.git
cd srtla
git checkout <ref>          # a branch, or a 40-hex commit for reproducibility
git submodule update --init # deps/argparse
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build
cmake --install build --prefix /usr/local   # installs bin/srtla_rec only

Upstream's build-and-push.yml is kept as-is: a push to main publishes ghcr.io/ceralive/srtla:<sha> and :latest, which is the one prebuilt form of the receiver. Server images that want the receiver either pull that image or compile it at image-build time from a pinned ref. There is nothing else to fetch.

project(srtla_rec VERSION 1.0.0) in CMakeLists.txt stays exactly upstream's. Nothing is released, so there is nothing to bump.

Test harness

Unit and handler tests live in tests/ (GoogleTest, fetched by CMake). They compile against the unmodified upstream src/; a test that would need a fork-only seam is not ported, and the drop list sits at the end of tests/CMakeLists.txt.

cmake -B build && cmake --build build && ctest --test-dir build --output-on-failure

-DSRTLA_BUILD_TESTS=OFF skips them. -DBUILD_COMPAT_TESTS=ON additionally builds the compat instruments (srt-sink, ext-ka-probe) under tests/compat/.

Compatibility harness

tests/compat/ is a Docker-based interop matrix that runs this receiver against the ecosystem's senders (BELABOX, irlserver's Rust sender, CERALIVE's Rust sender, a Moblin conformance mock) and the CERALIVE sender against the ecosystem's receivers (BELABOX, OpenIRL, go-srtla, go-irl). tests/compat/matrix.yaml is the single registry of pins and pairs; every third-party implementation is addressed by an immutable pin:. The CERALIVE sender uses ref: main; the libsrt registry and build-helper default use the published srt-v1.5.7+ceralive.2 tag. Completed A/B documents and evidence retain their original branch names and resolved SHAs as historical provenance; the hosted-jitter lane also keeps its exact legacy SHA pins.

bash tests/compat/run-matrix.sh --validate-only      # schema + invariants, no Docker
bash tests/compat/run-matrix.sh --tier blocking      # the pairs that gate a PR
python3 tests/compat/lib/ab-verdict.py --selftest    # the frozen A/B rules

The same harness carries two pre-registered A/B campaigns under tests/compat/scenarios/ab-*.yaml (periodic-NAK gating on the receiver-side libsrt, and recovery-keepalive cadence). Their decision rules are frozen text; ab-verdict.py recomputes the winner from a committed rows.json. They are never run by CI, only on a quiesced bench host. The privileged netem scenarios self-skip (exit 77) without CAP_NET_ADMIN. See docs/COMPATIBILITY.md.

CI

main is the canonical and default branch; legacy preserves the former canonical history. The build gate is named Build Check (build-check.yml), not ci.yml.

Workflow Trigger What it proves
build-check.yml push / PR to main AMD64 + ARM64 build, ctest, receiver-only install
static-analysis.yml push / PR to main clang-tidy, the workflow-contract scripts, a test lane
compat-matrix.yml push / PR to main / weekly harness self-test, blocking + informational pairs, pcap replay, upstream drift
build-and-push.yml push to main upstream's GHCR image

tests/workflow-contracts.sh pins the shape of the workflows (ccache bounds, permission scopes, job graph); tests/workflow-contracts-negative.sh proves those assertions bite.

Further reading

License

This project is licensed under the GNU Affero General Public License v3.0:

  • Copyright (C) 2020-2021 BELABOX project
  • Copyright (C) 2024 IRLToolkit Inc.
  • Copyright (C) 2024 OpenIRL
  • Copyright (C) 2025 IRLServer.com

You can use, modify, and distribute this code according to the terms of the AGPL-3.0.

About

SRT link aggregation — bonds multiple network interfaces (LTE, WiFi) for resilient live video streaming; upstream-aligned fork with irlserver improvements

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages