srtla_rec is an SRT transport proxy with link aggregation. SRTLA is designed to transport SRT traffic over multiple network links for capacity aggregation and redundancy. Traffic is balanced dynamically depending on network conditions. The primary application is bonding mobile modems for live streaming.
Note: This is a fork of the original SRTLA implementation by BELABOX. The original server component (srtla_rec) was marked as unsupported by BELABOX.
- Support for link aggregation across multiple network connections
- Automatic management of connection groups and individual connections
- Robust error handling and timeouts for inactive connections
- Logging of connection details for easy diagnostics
- Improved load balancing through ACK throttling
- Connection recovery mechanism for temporary network issues
- C++11 compatible compiler
- CMake for the build process
- spdlog library
- argparse library
SRTLA assumes that:
- Data is streamed from an SRT sender in caller mode to an SRT receiver in listener mode
- To benefit from link aggregation, the sender should have 2 or more network links to the SRT listener (typically internet-connected modems)
- The sender needs to have source routing configured, as SRTLA uses
bind()to map UDP sockets to specific connections
# Clone the repository
git clone https://github.com/OpenIRL/srtla.git
cd srtla
# Build with CMake
mkdir build
cd build
cmake ..
makesrtla_rec runs as a proxy between SRTla clients and an SRT server:
./srtla_rec [OPTIONS]--srtla_port PORT: Port to bind the SRTLA socket to (default: 5000)--srt_hostname HOST: Hostname of the downstream SRT server (default: 127.0.0.1)--srt_port PORT: Port of the downstream SRT server (default: 4001)--verbose: Enable verbose logging (default: disabled)--debug: Enable debug logging (default: disabled)--metrics_port PORT: Serve Prometheus metrics over HTTP on this port (default: 0, disabled)--metrics_bind ADDR: Numeric address the metrics endpoint binds to (default: 127.0.0.1)--metrics_detail: Also export per-connection metrics (default: disabled)
./srtla_rec --srtla_port 5000 --srt_hostname 192.168.1.10 --srt_port 4001 --verbose- srtla_rec creates a UDP socket for incoming SRTLA connections.
- Clients register with srtla_rec and create connection groups.
- Multiple connections can be added to a group.
- Data is received across all connections and forwarded to the SRT server.
- ACK packets are sent across all connections for timely delivery.
- Inactive connections and groups are automatically cleaned up.
SRTLA implements a protocol for packet transmission over multiple network connections, aggregating the data and making it available to the SRT protocol. The implementation is based on the following core mechanisms:
-
Connection Group Management: The software organizes connections into groups, with each group corresponding to an SRT stream. This enables support for multiple simultaneous SRTLA senders with a single receiver.
-
Packet Tracking: The code tracks received packets with sequence numbers and periodically sends SRTLA-ACK packets back to confirm receipt.
-
Two-phase Registration Process:
- Sender (conn 0):
SRTLA_REG1(contains sender-generated random ID) - Receiver:
SRTLA_REG2(contains full ID with receiver-generated values) - Sender (conn 0):
SRTLA_REG2(with full ID) - Receiver:
SRTLA_REG3 - Additional connections follow a similar pattern
- Sender (conn 0):
-
Error Handling: The receiver can send error responses:
SRTLA_REG_ERR: Operation temporarily failedSRTLA_REG_NGP: Invalid ID, group must be re-registered
-
Connection Cleanup: Inactive connections and groups are automatically cleaned up after a configurable timeout (default: 10 seconds).
-
Load Balancing through ACK Throttling: The server controls ACK frequency to influence the client's connection selection without requiring client-side modifications.
-
Connection Recovery Mechanism: Connections that show signs of recovery after temporary outages are given a chance to stabilize again.
The implementation uses epoll for event-based network I/O, allowing efficient handling of multiple simultaneous connections.
This version of SRTLA includes improvements to address two key issues in the original implementation:
In the original implementation, connections with temporary problems were completely disabled. In this enhanced version:
- Connections showing signs of recovery enter a "recovery mode"
- These connections receive more frequent keepalive packets for a set period (5 seconds)
- After successful recovery, they are fully reactivated for data transmission
- Recovery attempts are abandoned after a certain time if unsuccessful
This functionality allows connections to "heal" after brief disruptions (e.g., due to network issues) rather than remaining completely disabled.
In the original implementation, load was unevenly distributed across available connections. The new implementation:
- Introduces a monitoring and evaluation system for connection quality
- Checks connection quality every 5 seconds based on:
- Bandwidth (kbits/s) and performance ratio (actual vs expected bandwidth)
- Packet loss rate
- Connection performance relative to median bandwidth
- Assigns error points to each connection based on these metrics
- Calculates a quality weight for each connection (10% to 100%)
- Controls ACK packet frequency based on connection quality
- Good connections receive ACKs more frequently
- Poor connections receive ACKs less frequently
- Indirectly influences the window size in the client and thus connection selection
The result is better data distribution, with more stable connections carrying more load than problematic ones, without requiring client modifications.
The central innovation of this solution is ACK throttling for load distribution. It's based on the following principles:
- The SRT/SRTLA client (srtla_send) selects connections based on a score derived from the window size and in-flight packets.
- The window size in the client is adjusted when ACKs are received.
- By selectively throttling ACK frequency, we can indirectly control how quickly the window grows in the client.
- This causes the client to prefer better connections without requiring changes to the client code.
Connection quality is assessed by measuring and analyzing:
- Bandwidth Performance: The system calculates a performance ratio by comparing actual bandwidth to expected bandwidth. Poor performance relative to expectations leads to more error points
- Packet Loss: Higher loss rates lead to more error points
- Dynamic Bandwidth Evaluation: Connections are evaluated against either median bandwidth (for good connections) or minimum threshold (for poor connections)
- Grace Period: New connections receive a 10-second grace period before penalties are applied
The weight levels are:
- 100% (WEIGHT_FULL): Optimal connection
- 85% (WEIGHT_EXCELLENT): Excellent connection
- 70% (WEIGHT_DEGRADED): Slightly impaired connection
- 55% (WEIGHT_FAIR): Fair connection
- 40% (WEIGHT_POOR): Severely impaired connection
- 10% (WEIGHT_CRITICAL): Critically impaired connection
The recovery functionality works as follows:
- A connection that receives data again after being marked inactive is placed in recovery mode
- In this mode, keepalive packets are sent more frequently (every 1 second)
- If the connection remains stable for a short period (5 seconds), it is considered recovered
- If recovery does not occur within the time window, the recovery attempt is aborted
The following parameters can be adjusted to optimize behavior:
KEEPALIVE_PERIOD: Interval for keepalive packets during recovery (1 second)RECOVERY_CHANCE_PERIOD: Period during which a connection can attempt to recover (5 seconds)CONN_QUALITY_EVAL_PERIOD: Interval for evaluating connection quality (5 seconds)MIN_ACCEPTABLE_TOTAL_BANDWIDTH_KBPS: Minimum total bandwidth for acceptable streaming quality (1000 kbps)GOOD_CONNECTION_THRESHOLD: Threshold for considering a connection "good" (50% of max bandwidth)CONNECTION_GRACE_PERIOD: Grace period in seconds before applying penalties (10 seconds)- Various weight levels (
WEIGHT_FULL,WEIGHT_EXCELLENT,WEIGHT_DEGRADED,WEIGHT_FAIR,WEIGHT_POOR,WEIGHT_CRITICAL)
- The error point thresholds are static and could be dynamically adjusted to better adapt to different network situations
- Performance ratio calculations are based on bandwidth expectations that may need tuning for different network environments
The sender should implement congestion control using adaptive bitrate based on the SRT SRTO_SNDDATA size or measured RTT.
--metrics_port serves Prometheus text exposition over HTTP. Every path returns the same body.
./srtla_rec --metrics_port 9997
curl http://127.0.0.1:9997/metricsThe main epoll loop answers the scrape, so this costs no extra thread and takes no locks.
The endpoint listens on 127.0.0.1 by default because nothing authenticates it. --metrics_bind :: binds every interface for both IPv4 and IPv6, which is what a container needs, and --metrics_bind 0.0.0.0 binds IPv4 only. Put a firewall in front of either one. The address has to be numeric: --metrics_bind localhost logs an error at startup and leaves the endpoint off while the receiver keeps running.
Always exported: traffic counters (srtla_packets_received_total, srtla_forwarded_packets_total, srtla_srt_packets_received_total and others), registration outcomes (srtla_group_registrations_total, srtla_group_registrations_rejected_total{reason}), teardowns (srtla_groups_removed_total{reason}), auth throttling (srtla_auth_failures_total, srtla_auth_sources_blocked), NAK handling, connection recovery, send errors, and live gauges for groups and connections.
Divide srtla_packets_received_total by srtla_recv_batches_total for the receive loop fill ratio. A climbing ratio means recvmmsg() is returning fuller batches as the loop approaches saturation. That is the first number worth checking when a receiver feels slow.
--metrics_detail adds per-connection series (srtla_conn_*: bytes, packets, loss, weight, error points, RTT, window, in-flight, sender bitrate, idle time) labelled group, the group's local SRT port, and remote, the client address. These are what explain bonding behavior instead of just reporting it. Every reconnect creates a new remote label value, so watch series cardinality before enabling this on a busy receiver.
docs/grafana-dashboard.json covers every exported metric across 20 panels. Import it and pick your Prometheus data source. The per-connection panels sit in a collapsed row at the bottom and stay empty unless --metrics_detail is on.
srtla_rec creates information files about active connections under /tmp/srtla-group-[PORT]. These files contain the client IP addresses connected to a specific socket.
This checkout is CERALIVE's fork of
irlserver/srtla. Everything above this section is
upstream's README, unchanged. src/ is byte-identical to the upstream commit the fork
sits on; CERALIVE adds only the build policy, CI and test layer described here. See
AGENTS.md for the maintainer contract.
cmake --install puts exactly one binary in bin/: srtla_rec. The C srtla_send
target still builds (upstream keeps it and CI checks that it links), but it is not
installed and it is not what CERALIVE devices run. The device-side sender is the Rust
srtla-send-rs, which is CLI-compatible with
the C sender (SRT_LISTEN_PORT SRTLA_HOST SRTLA_PORT IPS_FILE) and speaks the same
wire protocol, extended keepalive included. The Verify receiver-only install step in
build-check.yml fails the build if anything else lands in bin/.
This repo cuts no GitHub releases, no .deb, no tarballs and no tags. Every consumer
builds the receiver from source at a commit or branch it chooses:
git clone https://github.com/CERALIVE/srtla.git
cd srtla
git checkout <ref> # a branch, or a 40-hex commit for reproducibility
git submodule update --init # deps/argparse
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build
cmake --install build --prefix /usr/local # installs bin/srtla_rec onlyUpstream's build-and-push.yml is kept as-is: a push to main publishes
ghcr.io/ceralive/srtla:<sha> and :latest, which is the one prebuilt form of the
receiver. Server images that want the receiver either pull that image or compile it at
image-build time from a pinned ref. There is nothing else to fetch.
project(srtla_rec VERSION 1.0.0) in CMakeLists.txt stays exactly upstream's. Nothing
is released, so there is nothing to bump.
Unit and handler tests live in tests/ (GoogleTest, fetched by CMake). They compile
against the unmodified upstream src/; a test that would need a fork-only seam is not
ported, and the drop list sits at the end of tests/CMakeLists.txt.
cmake -B build && cmake --build build && ctest --test-dir build --output-on-failure-DSRTLA_BUILD_TESTS=OFF skips them. -DBUILD_COMPAT_TESTS=ON additionally builds the
compat instruments (srt-sink, ext-ka-probe) under tests/compat/.
tests/compat/ is a Docker-based interop matrix that runs this receiver against the
ecosystem's senders (BELABOX, irlserver's Rust sender, CERALIVE's Rust sender, a Moblin
conformance mock) and the CERALIVE sender against the ecosystem's receivers (BELABOX,
OpenIRL, go-srtla, go-irl). tests/compat/matrix.yaml is the single registry of pins and
pairs; every third-party implementation is addressed by an immutable pin:.
The CERALIVE sender uses ref: main; the libsrt registry and build-helper default
use the published srt-v1.5.7+ceralive.2 tag. Completed A/B documents and evidence
retain their original branch names and resolved SHAs as historical provenance;
the hosted-jitter lane also keeps its exact legacy SHA pins.
bash tests/compat/run-matrix.sh --validate-only # schema + invariants, no Docker
bash tests/compat/run-matrix.sh --tier blocking # the pairs that gate a PR
python3 tests/compat/lib/ab-verdict.py --selftest # the frozen A/B rulesThe same harness carries two pre-registered A/B campaigns under
tests/compat/scenarios/ab-*.yaml (periodic-NAK gating on the receiver-side libsrt,
and recovery-keepalive cadence). Their decision rules are frozen text; ab-verdict.py
recomputes the winner from a committed rows.json. They are never run by CI, only on
a quiesced bench host. The privileged netem scenarios self-skip (exit 77) without
CAP_NET_ADMIN. See docs/COMPATIBILITY.md.
main is the canonical and default branch; legacy preserves the former canonical
history. The build gate is named Build Check (build-check.yml), not ci.yml.
| Workflow | Trigger | What it proves |
|---|---|---|
build-check.yml |
push / PR to main |
AMD64 + ARM64 build, ctest, receiver-only install |
static-analysis.yml |
push / PR to main |
clang-tidy, the workflow-contract scripts, a test lane |
compat-matrix.yml |
push / PR to main / weekly |
harness self-test, blocking + informational pairs, pcap replay, upstream drift |
build-and-push.yml |
push to main |
upstream's GHCR image |
tests/workflow-contracts.sh pins the shape of the workflows (ccache bounds, permission
scopes, job graph); tests/workflow-contracts-negative.sh proves those assertions bite.
docs/HOW_IT_WORKS.md: protocol, registration, quality modeldocs/NETWORK_SETUP.md: source routing on the sender hostdocs/TROUBLESHOOTING.md: failure modes and diagnosticsdocs/COMPATIBILITY.md: ecosystem interop and guaranteesdocs/adr/ADR-002-srt-patch-necessity.md: the historical srt-patch A/B
This project is licensed under the GNU Affero General Public License v3.0:
- Copyright (C) 2020-2021 BELABOX project
- Copyright (C) 2024 IRLToolkit Inc.
- Copyright (C) 2024 OpenIRL
- Copyright (C) 2025 IRLServer.com
You can use, modify, and distribute this code according to the terms of the AGPL-3.0.