The RTMP handshake is three fixed-size exchanges that happen before any stream data moves. The client sends C0 (1 byte, protocol version 3) and C1 (1,536 bytes). The server replies with S0, S1 and S2 (3,073 bytes). The client then sends C2 (1,536 bytes), for 6,146 bytes in total. If a live ingest connection stalls before the stream starts, the last packet in a capture shows which layer failed: TCP, the handshake itself, the AMF command phase, or the first media frames. This playbook isolates the failing layer in about 20 minutes, using tcpdump, tshark and FFmpeg on the encoder or ingest host.
The sequence is strict. Each side sends fixed-length blobs, and each side waits for specific bytes before it proceeds. That makes it easy to read in a packet capture.
The complex handshake embeds an HMAC-SHA256 digest inside C1 and S1. The digest is placed at an offset computed from four bytes of the block, and it is keyed with the well-known Adobe player and media server strings. Legacy Flash Media Server deployments and some older ingest stacks require the digest. Most modern servers accept either form.
A client that sends zeros in bytes 4 through 7 to a server that validates the digest gets its socket closed right after C1. The reverse case also happens: a strict client can reject an S2 that does not echo C1 exactly. FFmpeg's librtmp-free implementation sends the digest form by default, which makes it a useful known-good reference client.
The handshake only proves that both ends speak RTMP. Publishing requires a second phase of AMF0 command messages carried over the chunk stream. The default chunk size in this phase is 128 bytes, until one side sends Set Chunk Size.
A healthy RTMP publish needs roughly five network round trips before the first video byte arrives. One is for the TCP handshake, one is for the C0 and C1 to S0, S1 and S2 exchange, and about three are for connect, createStream and publish. Over an 80 ms path in 2026 that totals about 400 ms, plus one more round trip for RTMPS over TLS 1.3. An ingest connection that has not reached NetStream.Publish.Start within 2 seconds is stalled, not slow.
Run this command on whichever host you control. RTMPS on port 443 is encrypted, so reproduce against plain RTMP on port 1935 if the ingest allows it. The capture should contain only the RTMP port, so that it stays small.
tcpdump -i any -nn -s 0 -w rtmp-ingest.pcap tcp port 1935
This rules out your encoder's configuration. The command generates a synthetic 720p30 source with a 2-second keyframe interval (-g 60). Debug logging prints "Handshaking..." and then each command that the server acknowledges.
ffmpeg -v debug -re -f lavfi -i testsrc2=size=1280x720:rate=30 -f lavfi -i sine=frequency=1000 -c:v libx264 -preset veryfast -g 60 -c:a aac -f flv rtmp://INGEST_HOST/APP_NAME/STREAM_KEY
Replace INGEST_HOST with the ingest hostname, APP_NAME with the application path (often "live"), and STREAM_KEY with your test key. If FFmpeg reaches Publish.Start but your encoder does not, the problem is in the encoder, not the network.
The rtmpt dissector labels each phase in the Info column. Look for the last line that succeeded.
tshark -r rtmp-ingest.pcap -Y rtmpt
A clean session reads in this order: Handshake C0+C1, Handshake S0+S1+S2, Handshake C2, then connect, _result, createStream, _result, publish, and onStatus. Whatever appears after the last good line is where the stall is.
When tshark shows no RTMP at all, inspect the raw hex of the first data segments.
tcpdump -r rtmp-ingest.pcap -nn -X -c 6 tcp port 1935 and greater 100
The first client payload byte tells you what the client actually sent:
C0 and C1 together are 1,537 bytes, which is larger than a 1,460-byte MSS. The server's 3,073-byte reply fills about three full-size segments. That means the RTMP handshake is often the first full-size traffic on a new path. The SYN exchange succeeds, and then large segments disappear inside VPN, GRE or PPPoE tunnels.
ping -M do -s 1472 -c 3 INGEST_HOST
If this fails while a 1,300-byte payload succeeds, enable MTU probing on the encoder host. This lets the kernel recover when ICMP messages are filtered.
sysctl -w net.ipv4.tcp_mtu_probing=1
| Last packet seen | Likely cause | Fix |
|---|---|---|
| SYN retransmits, no SYN-ACK | Port 1935 blocked by an egress firewall | Open the port, or switch to RTMPS on 443 |
| C0+C1 retransmitted, no S0 | Path MTU black hole | Enable MTU probing or clamp the MSS on the tunnel |
| C0+C1, then an immediate FIN or RST | Wrong first byte (TLS or PROXY header), or a digest check failed | Match the URL scheme to the listener; align PROXY protocol on both ends; test with FFmpeg |
| connect, then _error | Wrong app name or tcUrl (NetConnection.Connect.Rejected) | Correct APP_NAME; check token or auth parameters |
| publish, then onStatus BadName or a silent close | Invalid key, or the key is already publishing | Rotate the key; kill the stale session on the ingest side |
| Publish.Start, but the player stays black | Missing AVC sequence header, or a keyframe interval above 4 seconds | Set the keyframe interval to 2 seconds; confirm the sequence header is sent first |
If the capture never shows S0, the fault is in the network path. If the server closes the connection after C1, the fault is protocol framing. Any stall after Publish.Start is an encoder problem.
Rerun Steps 1 through 3 after the fix. The tshark output should show onStatus NetStream.Publish.Start within 2 seconds of the SYN on paths under 100 ms. Then video and audio messages should arrive continuously, with timestamps that increase monotonically. On the server side, the stream should appear in your ingest statistics endpoint with a nonzero input bitrate within one keyframe interval.
The diagnostic steps change nothing, so only the fixes need reverting. Set net.ipv4.tcp_mtu_probing back to 0 with the same sysctl command. Restore the saved encoder profile, and revert any change to the PROXY protocol setting on the load balancer and the ingest server together. Delete the pcap files as well. They contain your stream key in cleartext.
Capturing on a busy ingest node costs CPU and disk. Run captures with a BPF filter for a single encoder IP, not the whole port. RTMPS protects stream keys but hides the handshake from passive capture. Keep a plain RTMP test path in staging for exactly this reason.
Clamping the MSS to 1,360 bytes adds roughly one percentage point of header overhead, which is negligible at 6 to 8 Mbps contribution bitrates. Encoder chunk sizes of 4,096 bytes or more reduce per-message overhead. Very large chunk sizes, however, delay audio interleaving on constrained uplinks. If an idle load balancer sits in front of the ingest, raise its TCP idle timeout above the ingest server's own timeout (60 seconds by default in nginx-rtmp) so the two do not race each other.
For more protocol walkthroughs like this one, browse the live streaming and RTMP debugging guides on the BlazingCDN engineering blog.
Each side sends 3,073 bytes in the RTMP handshake, for a total of 6,146 bytes. The client sends C0 (1 byte), C1 (1,536 bytes) and C2 (1,536 bytes). The server sends S0, S1 and S2 with the same sizes. These sizes are fixed, so any different length in a capture points to a non-RTMP payload.
An immediate close after C0 and C1 almost always means the server rejected the first bytes it received. Common causes are TLS sent to a plain RTMP port, a PROXY protocol header the server does not expect, or a failed digest check in the complex handshake. Inspect the first payload byte to tell these apart.
Plain RTMP uses TCP port 1935. RTMPS wraps the same handshake and chunk stream in TLS, usually on port 443. RTMPS adds one round trip with TLS 1.3, or two with TLS 1.2, before C0 is sent. It also encrypts the stream key, which makes passive packet-level RTMP debugging impossible without terminating TLS.
The RTMP connection is working, but the decoder has nothing to start from. Most black-screen cases come from a missing AVC sequence header or a long keyframe interval. A 10-second GOP means the ingest cannot produce a playable segment for up to 10 seconds. Set the keyframe interval to 2 seconds.
Capture one publish from your production encoder location and measure the time from SYN to NetStream.Publish.Start. Divide that by the measured RTT. A healthy plain RTMP session lands near 5 RTTs, and RTMPS lands near 6. Anything above 8 means something is retransmitting or waiting, and the tshark timeline will show which phase is responsible. Record the baseline, add it to your pre-event checklist, and rerun the test whenever the network path, load balancer or encoder firmware changes.