Content Delivery Network Blog

The RTMP Handshake Explained Step by Step

Written by BlazingCDN | Sep 29, 2026, 7:33:00 AM

The RTMP handshake is three fixed-size exchanges that happen before any stream data moves. The client sends C0 (1 byte, protocol version 3) and C1 (1,536 bytes). The server replies with S0, S1 and S2 (3,073 bytes). The client then sends C2 (1,536 bytes), for 6,146 bytes in total. If a live ingest connection stalls before the stream starts, the last packet in a capture shows which layer failed: TCP, the handshake itself, the AMF command phase, or the first media frames. This playbook isolates the failing layer in about 20 minutes, using tcpdump, tshark and FFmpeg on the encoder or ingest host.

What happens in the RTMP handshake, packet by packet

The sequence is strict. Each side sends fixed-length blobs, and each side waits for specific bytes before it proceeds. That makes it easy to read in a packet capture.

  1. C0 (1 byte): the RTMP version. The value 0x03 means plain RTMP. The value 0x06 means RTMPE, the legacy encrypted variant. The client sends C0 together with C1 in a single write.
  2. C1 (1,536 bytes): a 4-byte timestamp, 4 bytes that are zero in the simple handshake or carry a client version in the digest handshake, and 1,528 bytes of random data.
  3. S0 and S1 (1 plus 1,536 bytes): the server's version byte and its own timestamp and random block. The server can send these as soon as C0 arrives.
  4. S2 (1,536 bytes): an echo of C1. It holds the client's timestamp, the time the server read C1, and the client's random bytes. Most servers send S0, S1 and S2 as one 3,073-byte burst.
  5. C2 (1,536 bytes): the client's echo of S1. Many encoders coalesce C2 with the first chunk-stream messages in the same TCP segment.

Simple vs. complex (digest) RTMP handshake

The complex handshake embeds an HMAC-SHA256 digest inside C1 and S1. The digest is placed at an offset computed from four bytes of the block, and it is keyed with the well-known Adobe player and media server strings. Legacy Flash Media Server deployments and some older ingest stacks require the digest. Most modern servers accept either form.

A client that sends zeros in bytes 4 through 7 to a server that validates the digest gets its socket closed right after C1. The reverse case also happens: a strict client can reject an S2 that does not echo C1 exactly. FFmpeg's librtmp-free implementation sends the digest form by default, which makes it a useful known-good reference client.

What the RTMP connection does after the handshake

The handshake only proves that both ends speak RTMP. Publishing requires a second phase of AMF0 command messages carried over the chunk stream. The default chunk size in this phase is 128 bytes, until one side sends Set Chunk Size.

  1. The client sends Set Chunk Size, often 4,096 bytes. It then sends the connect command with app, tcUrl and flashVer.
  2. The server replies with Window Acknowledgement Size, Set Peer Bandwidth, a Stream Begin user control event, and then _result carrying NetConnection.Connect.Success.
  3. The client sends releaseStream, FCPublish and createStream. The server returns _result with a message stream ID.
  4. The client sends publish with the stream key and the type "live". The server answers onStatus with NetStream.Publish.Start.
  5. The client sends the @setDataFrame onMetaData message, then the AVC sequence header, the AAC AudioSpecificConfig, and the first keyframe.

A healthy RTMP publish needs roughly five network round trips before the first video byte arrives. One is for the TCP handshake, one is for the C0 and C1 to S0, S1 and S2 exchange, and about three are for connect, createStream and publish. Over an 80 ms path in 2026 that totals about 400 ms, plus one more round trip for RTMPS over TLS 1.3. An ingest connection that has not reached NetStream.Publish.Start within 2 seconds is stalled, not slow.

Prerequisites

  • Shell access to the encoder host, the ingest host, or a span port between them.
  • tcpdump and tshark (Wireshark 3.x or later, which includes the rtmpt dissector).
  • An FFmpeg build with libx264 and the native AAC encoder.
  • A test stream key. Do not debug with your production key while a real event is live.

How to debug an RTMP connection that stalls before the stream starts

Step 1: Capture the ingest session

Run this command on whichever host you control. RTMPS on port 443 is encrypted, so reproduce against plain RTMP on port 1935 if the ingest allows it. The capture should contain only the RTMP port, so that it stays small.

tcpdump -i any -nn -s 0 -w rtmp-ingest.pcap tcp port 1935

Step 2: Reproduce with a known-good FFmpeg publisher

This rules out your encoder's configuration. The command generates a synthetic 720p30 source with a 2-second keyframe interval (-g 60). Debug logging prints "Handshaking..." and then each command that the server acknowledges.

ffmpeg -v debug -re -f lavfi -i testsrc2=size=1280x720:rate=30 -f lavfi -i sine=frequency=1000 -c:v libx264 -preset veryfast -g 60 -c:a aac -f flv rtmp://INGEST_HOST/APP_NAME/STREAM_KEY

Replace INGEST_HOST with the ingest hostname, APP_NAME with the application path (often "live"), and STREAM_KEY with your test key. If FFmpeg reaches Publish.Start but your encoder does not, the problem is in the encoder, not the network.

Step 3: Decode the RTMP handshake with tshark

The rtmpt dissector labels each phase in the Info column. Look for the last line that succeeded.

tshark -r rtmp-ingest.pcap -Y rtmpt

A clean session reads in this order: Handshake C0+C1, Handshake S0+S1+S2, Handshake C2, then connect, _result, createStream, _result, publish, and onStatus. Whatever appears after the last good line is where the stall is.

Step 4: Check the first payload byte

When tshark shows no RTMP at all, inspect the raw hex of the first data segments.

tcpdump -r rtmp-ingest.pcap -nn -X -c 6 tcp port 1935 and greater 100

The first client payload byte tells you what the client actually sent:

  • 0x03 means correct plain RTMP.
  • 0x16 is a TLS ClientHello, so the client is speaking RTMPS to a plain RTMP listener.
  • 0x50 is the letter P from a PROXY protocol header, which means a load balancer is prepending it to a server that does not expect it.

Step 5: Rule out a path MTU black hole

C0 and C1 together are 1,537 bytes, which is larger than a 1,460-byte MSS. The server's 3,073-byte reply fills about three full-size segments. That means the RTMP handshake is often the first full-size traffic on a new path. The SYN exchange succeeds, and then large segments disappear inside VPN, GRE or PPPoE tunnels.

ping -M do -s 1472 -c 3 INGEST_HOST

If this fails while a 1,300-byte payload succeeds, enable MTU probing on the encoder host. This lets the kernel recover when ICMP messages are filtered.

sysctl -w net.ipv4.tcp_mtu_probing=1

RTMP debugging matrix: symptom, cause, fix

Last packet seenLikely causeFix
SYN retransmits, no SYN-ACKPort 1935 blocked by an egress firewallOpen the port, or switch to RTMPS on 443
C0+C1 retransmitted, no S0Path MTU black holeEnable MTU probing or clamp the MSS on the tunnel
C0+C1, then an immediate FIN or RSTWrong first byte (TLS or PROXY header), or a digest check failedMatch the URL scheme to the listener; align PROXY protocol on both ends; test with FFmpeg
connect, then _errorWrong app name or tcUrl (NetConnection.Connect.Rejected)Correct APP_NAME; check token or auth parameters
publish, then onStatus BadName or a silent closeInvalid key, or the key is already publishingRotate the key; kill the stale session on the ingest side
Publish.Start, but the player stays blackMissing AVC sequence header, or a keyframe interval above 4 secondsSet the keyframe interval to 2 seconds; confirm the sequence header is sent first

If the capture never shows S0, the fault is in the network path. If the server closes the connection after C1, the fault is protocol framing. Any stall after Publish.Start is an encoder problem.

Validation: what a healthy RTMP handshake and publish look like

Rerun Steps 1 through 3 after the fix. The tshark output should show onStatus NetStream.Publish.Start within 2 seconds of the SYN on paths under 100 ms. Then video and audio messages should arrive continuously, with timestamps that increase monotonically. On the server side, the stream should appear in your ingest statistics endpoint with a nonzero input bitrate within one keyframe interval.

Rollback

The diagnostic steps change nothing, so only the fixes need reverting. Set net.ipv4.tcp_mtu_probing back to 0 with the same sysctl command. Restore the saved encoder profile, and revert any change to the PROXY protocol setting on the load balancer and the ingest server together. Delete the pcap files as well. They contain your stream key in cleartext.

Trade-offs and tuning

Capturing on a busy ingest node costs CPU and disk. Run captures with a BPF filter for a single encoder IP, not the whole port. RTMPS protects stream keys but hides the handshake from passive capture. Keep a plain RTMP test path in staging for exactly this reason.

Clamping the MSS to 1,360 bytes adds roughly one percentage point of header overhead, which is negligible at 6 to 8 Mbps contribution bitrates. Encoder chunk sizes of 4,096 bytes or more reduce per-message overhead. Very large chunk sizes, however, delay audio interleaving on constrained uplinks. If an idle load balancer sits in front of the ingest, raise its TCP idle timeout above the ingest server's own timeout (60 seconds by default in nginx-rtmp) so the two do not race each other.

For more protocol walkthroughs like this one, browse the live streaming and RTMP debugging guides on the BlazingCDN engineering blog.

FAQ: RTMP handshake and RTMP connection debugging

How many bytes are exchanged in the RTMP handshake?

Each side sends 3,073 bytes in the RTMP handshake, for a total of 6,146 bytes. The client sends C0 (1 byte), C1 (1,536 bytes) and C2 (1,536 bytes). The server sends S0, S1 and S2 with the same sizes. These sizes are fixed, so any different length in a capture points to a non-RTMP payload.

Why does an RTMP connection drop right after the handshake starts?

An immediate close after C0 and C1 almost always means the server rejected the first bytes it received. Common causes are TLS sent to a plain RTMP port, a PROXY protocol header the server does not expect, or a failed digest check in the complex handshake. Inspect the first payload byte to tell these apart.

What port does RTMP use, and how does RTMPS differ?

Plain RTMP uses TCP port 1935. RTMPS wraps the same handshake and chunk stream in TLS, usually on port 443. RTMPS adds one round trip with TLS 1.3, or two with TLS 1.2, before C0 is sent. It also encrypts the stream key, which makes passive packet-level RTMP debugging impossible without terminating TLS.

Why is the stream black after NetStream.Publish.Start?

The RTMP connection is working, but the decoder has nothing to start from. Most black-screen cases come from a missing AVC sequence header or a long keyframe interval. A 10-second GOP means the ingest cannot produce a playable segment for up to 10 seconds. Set the keyframe interval to 2 seconds.

Run the five-round-trip test on your ingest path this week

Capture one publish from your production encoder location and measure the time from SYN to NetStream.Publish.Start. Divide that by the measured RTT. A healthy plain RTMP session lands near 5 RTTs, and RTMPS lands near 6. Anything above 8 means something is retransmitting or waiting, and the tshark timeline will show which phase is responsible. Record the baseline, add it to your pre-event checklist, and rerun the test whenever the network path, load balancer or encoder firmware changes.