<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Learning in Public]]></title><description><![CDATA[Learning in Public]]></description><link>https://adiem.hashnode.dev</link><image><url>https://cdn.hashnode.com/uploads/logos/6a2a9d5b695440ecb50e316b/5b876a4d-d72a-4efc-b352-00268ea684b2.png</url><title>Learning in Public</title><link>https://adiem.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Wed, 30 Sep 2026 13:33:53 GMT</lastBuildDate><atom:link href="https://adiem.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[I Thought WebRTC Was Just "Add Some Sockets." Then I Hit RTP.]]></title><description><![CDATA[I've built backends. Designed systems. Done RAG pipelines. I thought I had a decent mental model of how real-time stuff works — sockets, events, done.
Then I started exploring WebRTC properly. Not the]]></description><link>https://adiem.hashnode.dev/i-thought-webrtc-was-just-add-some-sockets-then-i-hit-rtp</link><guid isPermaLink="true">https://adiem.hashnode.dev/i-thought-webrtc-was-just-add-some-sockets-then-i-hit-rtp</guid><dc:creator><![CDATA[Adam]]></dc:creator><pubDate>Mon, 15 Jun 2026 06:20:36 GMT</pubDate><content:encoded><![CDATA[<p>I've built backends. Designed systems. Done RAG pipelines. I thought I had a decent mental model of how real-time stuff works — sockets, events, done.</p>
<p>Then I started exploring WebRTC properly. Not the copy-paste-a-tutorial kind. The <em>actually understand what's happening</em> kind.</p>
<p>That's where things got uncomfortable. That discomfort is exactly why I'm writing this.</p>
<hr />
<h2>The "just use sockets" trap</h2>
<p>When I first thought about video calling, my brain went: <em>peer A sends data, peer B receives it, put a server in between if needed.</em> Like a chat app, but with video. How different could it be?</p>
<p>Very. Very different.</p>
<p>The first thing that breaks your mental model is that video/audio data isn't like your typical JSON payload. It's a <em>stream</em>. Continuous. Time-sensitive. Lossy by design — dropping a frame is acceptable, waiting for a retransmit is not. TCP isn't ideal here. Media is time-sensitive, so WebRTC prefers UDP because it's better to drop a frame than wait for a retransmission.</p>
<p>And once you're on UDP, you need something sitting on top of it to give you <em>some</em> structure. That's where <strong>RTP</strong> (Real-time Transport Protocol) comes in.</p>
<p>Before any media flows, peers still need a way to discover each other and exchange connection metadata. WebRTC intentionally leaves signaling undefined, which is why we often use WebSockets or <a href="http://Socket.IO">Socket.IO</a> for that part.</p>
<hr />
<h2>RTP — The thing nobody talks about enough</h2>
<p>RTP is what actually carries your media. It's a protocol that runs over UDP and gives you:</p>
<ul>
<li><p>Sequence numbers (so you know packet order, even if they arrive out of order)</p>
</li>
<li><p>Timestamps (for syncing audio and video)</p>
</li>
<li><p>Payload type (so the receiver knows how to decode it)</p>
</li>
</ul>
<p>Alongside RTP, there's <strong>RTCP</strong> — the control sibling. It sends feedback: packet loss stats, jitter, round-trip time. This is how the sender knows to back off quality or increase bitrate.</p>
<p>When I understood this, something clicked: WebRTC isn't magic. It's a stack. getUserMedia → encode → RTP → network → RTP → decode → render. Somewhere in that flow, you have to make architectural decisions. And those decisions are what makes or breaks your system at scale.</p>
<hr />
<h2>P2P — The "it just works" lie</h2>
<p>WebRTC natively supports P2P (peer-to-peer). Two browsers, negotiate via SDP (Session Description Protocol), exchange ICE candidates, done. No server needed for media.</p>
<p>This is great for two people. And it genuinely works.</p>
<p>But the moment you have <strong>more than two peers in a call</strong>, P2P becomes a problem.</p>
<p>Say you have a 6-person call. In full mesh P2P:</p>
<ul>
<li><p>Every peer connects to every other peer</p>
</li>
<li><p>That's 5 upload streams per person</p>
</li>
<li><p>6 people × 5 streams = 30 connections total</p>
</li>
</ul>
<p>Your users are uploading 5 simultaneous video streams. On their home WiFi. While their OS is also doing background updates. You see the issue.</p>
<p>P2P mesh is a bandwidth and CPU nightmare at scale. There had to be a better way.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a2a9d5b695440ecb50e316b/545df37d-b4de-48c0-9863-92b86c045ef3.png" alt="" style="display:block;margin:0 auto" />

<hr />
<h2>MCU — The old-school approach</h2>
<p><strong>MCU (Multipoint Control Unit)</strong> is the server-heavy solution. The server receives all streams, <em>decodes</em> them, <em>mixes</em> them into one combined stream, <em>re-encodes</em> it, and sends that single stream to everyone.</p>
<p>From the client's perspective: bliss. One upload stream, one download stream.</p>
<p>From the server's perspective: absolute carnage. Decoding and re-encoding video in real-time is computationally brutal. You're paying for that CPU hard. Also, the server now has access to the decoded media — which is a privacy consideration.</p>
<p>MCU made sense in the era of expensive codec hardware. In today's world, SFUs are often the preferred choice, unless you specifically need server-side mixing, broadcasting, or recording workflows.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a2a9d5b695440ecb50e316b/958db168-8786-4a37-97d1-d103917a1c85.png" alt="" style="display:block;margin:0 auto" />

<hr />
<h2>SFU — The architecture that made things click</h2>
<p><strong>SFU (Selective Forwarding Unit)</strong> is the sweet spot.</p>
<p>The server receives streams from all participants, but <em>doesn't decode or mix anything</em>. It just <strong>forwards</strong> the RTP packets selectively to the right peers.</p>
<p>Every client still uploads once (to the SFU). But instead of receiving one mixed stream (MCU) or N-1 streams (P2P), they receive forwarded streams from the SFU — which can be optimized, filtered, and scaled intelligently.</p>
<p>This unlocks things like:</p>
<ul>
<li><p><strong>Simulcast</strong> — send multiple quality layers, SFU picks the right one per receiver based on their bandwidth</p>
</li>
<li><p><strong>Selective forwarding</strong> — don't send the presenter's video to someone who has it minimized</p>
</li>
<li><p><strong>Recording</strong> — tap into the RTP stream server-side without decoding</p>
</li>
</ul>
<p>The SFU doesn't touch your media. It just routes packets. This is both the efficiency win and the privacy win.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a2a9d5b695440ecb50e316b/19905c95-fd0d-479a-b560-bd85dd1bd38f.png" alt="" style="display:block;margin:0 auto" />

<hr />
<h2>So why mediasoup?</h2>
<p>When I started implementing an SFU myself, I hit a wall fast.</p>
<p>You're not just forwarding UDP packets. You have to:</p>
<ul>
<li><p>Handle WebRTC signaling (SDP, ICE)</p>
</li>
<li><p>Manage DTLS handshakes</p>
</li>
<li><p>Track SSRCs and stream lifecycles</p>
</li>
<li><p>Handle simulcast layers</p>
</li>
<li><p>Manage bitrate estimation and RTCP feedback loops</p>
</li>
<li><p>Deal with NACK, PLI and FIR recovery mechanisms</p>
</li>
</ul>
<p>At this point, you realize you're not building a feature anymore — you're implementing years of networking research.</p>
<p><strong>mediasoup</strong> solves this. It's a Node.js library with a C++ core that handles all the low-level WebRTC/RTP machinery and exposes a clean API for routing, rooms, and transports. You focus on <em>who</em> talks to <em>whom</em>. It handles <em>how</em> the packets get there.</p>
<p>I didn't reach for mediasoup because it was the trending tool. I reached for it because I understood the problem it was solving — and that understanding only came after going deep enough to feel the pain.</p>
<p><a href="https://github.com/versatica/mediasoup">mediasoup</a></p>
<img src="https://cdn.hashnode.com/uploads/covers/6a2a9d5b695440ecb50e316b/4afa0da2-932a-40e3-8ef5-db0d95955283.png" alt="" style="display:block;margin:0 auto" />

<hr />
<h2>The pattern I noticed</h2>
<p>Every time I hit a new layer in this stack — RTP, RTCP, ICE, DTLS, simulcast — my instinct was to abstract it away. But each layer was there for a reason. The complexity isn't accidental.</p>
<p>P2P fails at scale → SFU fixes the topology problem SFU is hard to implement correctly → mediasoup handles the low-level correctly You still need to design rooms, routing, signaling → that's your job as the backend engineer</p>
<p>Understanding <em>why</em> a tool exists is a different thing from knowing <em>how</em> to use it. I'm glad I took the longer road here.</p>
<p>If you're building anything with real-time media and you're still at the "just WebRTC it" stage — go one level deeper. Understand RTP. Feel the P2P pain. Then look at your options.</p>
<p>The architecture will make a lot more sense.</p>
<p>-Adam</p>
]]></content:encoded></item></channel></rss>