Content Delivery Network Blog

What Is HLS? HTTP Live Streaming Explained

Written by BlazingCDN | Sep 30, 2026, 1:11:02 PM

HTTP Live Streaming (HLS) is an adaptive bitrate video delivery protocol, specified in RFC 8216, that cuts a stream into short media segments, lists them in plain-text M3U8 playlists, and lets the player fetch everything over ordinary HTTP. The player switches between renditions of different bitrates segment by segment, so any HTTP server or CDN can deliver it.

Apple introduced HLS in 2009, and the IETF published it as RFC 8216 in 2017. Apple's HLS authoring guidance recommends 6-second segments. At that length, classic HLS plays roughly 18 to 30 seconds behind live, while Low-Latency HLS (LL-HLS) cuts that to about 2 to 5 seconds with partial segments.

How HTTP live streaming works: segments, playlists and the player loop

HTTP live streaming works in three layers. The packager writes media segments, and media playlists list those segments in order. A multivariant playlist lists every available rendition. The player downloads the multivariant playlist and picks a rendition. It then loops: it reloads the media playlist, fetches new segments and re-evaluates bandwidth before each request.

Segments

A segment is a self-contained chunk of encoded video, usually 2 to 6 seconds long. Each segment starts on a keyframe, so the encoder's GOP (group of pictures) length must divide evenly into the segment duration. Otherwise the player cannot switch renditions cleanly at segment boundaries. Segments are either MPEG-TS files or fragmented MP4 (fMP4). HLS has supported fMP4 since 2016, and that support lets HLS share CMAF segments with MPEG-DASH.

Playlists

The multivariant playlist (formerly the "master" playlist) carries one entry per rendition, with its peak bandwidth, resolution and codecs. Each rendition has its own media playlist, which lists segment URIs and their durations. For video on demand, the media playlist is complete and ends with an end marker. For live, the playlist works as a sliding window: the packager appends new segments and drops old ones, and the media sequence number counts how many have been dropped.

Adaptive bitrate and timing

The adaptive bitrate logic lives entirely in the player. It measures segment download throughput and buffer level, then requests the next segment from whichever rendition fits. The protocol itself sets the timing. RFC 8216 tells clients not to start playback closer than three target durations from the end of a live playlist. It also tells them to reload the media playlist once per target duration, or after half a target duration if the playlist did not change. Those two rules explain the classic latency floor: with a 6-second target duration, the player sits at least 18 seconds behind the live edge, before encoder and packager delay are added.

LL-HLS keeps the same model and adds partial segments. These are sub-segments, typically 200 ms to 1 s long, that are published before the full segment completes. It also adds blocking playlist reloads, where the server holds a playlist request until the next part exists. The player asks for a specific part with the _HLS_msn and _HLS_part query parameters. Any CDN in the path therefore has to include those parameters in the cache key.

HLS streaming at scale: request and bandwidth math

These numbers show what the protocol costs at the edge. The worked example assumes 100,000 concurrent viewers, 6-second segments and a 5 Mb/s average rendition.

  • Edge requests: each viewer fetches one segment and one playlist every 6 seconds, which is 0.33 requests per second. Across 100,000 viewers, that comes to about 33,000 requests per second.
  • LL-HLS edge requests: with 1-second parts, each viewer makes roughly one part request and one blocking playlist request per second. That is about 200,000 requests per second, six times the classic load for the same audience.
  • Egress: 5 Mb/s multiplied by 100,000 viewers is 500 Gb/s. One viewer-hour at 5 Mb/s is 2.25 GB, so one hour of the event moves about 225 TB.
  • Origin load with shielding: with an origin shield and request coalescing, the origin sees each new segment and playlist only once per shield. Six renditions produce about 12 origin requests per 6 seconds, regardless of whether 1,000 or 1,000,000 people are watching.

The takeaway for capacity planning: request rate scales with audience divided by segment duration, while origin load scales only with rendition count. Shrinking segments to cut latency multiplies edge request volume long before it moves egress.

Where HLS sits in the video delivery stack

HLS is a delivery format, not an ingest protocol. A typical live pipeline runs through five stages:

  1. A contribution encoder pushes a feed over RTMP or SRT.
  2. A transcoder produces the bitrate ladder.
  3. A packager writes HLS segments and playlists to an origin.
  4. A CDN caches and serves them.
  5. A player on the device runs the adaptive bitrate loop.

Content protection attaches at the packager. This is either AES-128 whole-segment encryption, SAMPLE-AES, or FairPlay DRM, signaled through the key tag in the media playlist.

Caching splits cleanly by file type. Segments are immutable once written, so they take long TTLs and hit ratios near the top of any cache's range. Live media playlists change every target duration, so they need short TTLs, commonly about half the target duration. Treat them as a separate cache class with its own hit-ratio metric.

HLS vs. DASH, RTMP and WebRTC

HLS vs. MPEG-DASH: both are segment-based adaptive bitrate protocols over HTTP. DASH is an ISO standard (ISO/IEC 23009-1) with an XML manifest, while HLS uses text M3U8 playlists. With CMAF, both can reference the same fMP4 segments, so the practical difference shrinks to manifests, DRM systems and native playback on Apple devices, where HLS is the native format.

HLS vs. RTMP: RTMP is a persistent-connection protocol that survives today mainly as an ingest path from encoders to platforms. HLS replaced it for playback because stateless HTTP requests cache on any CDN and pass through corporate firewalls.

HLS vs. WebRTC: WebRTC delivers sub-second, interactive latency over UDP. However, every viewer holds a stateful session, so it does not use HTTP caches. HLS trades latency for cacheability. LL-HLS closes most of the gap for broadcast-style events, while WebRTC remains the choice when viewers talk back.

Reading an HLS video playlist, tag by tag

The table below walks through the tags you will meet when you inspect a real playlist. It shows which playlist each tag belongs to and what it controls.

Tag Playlist Example value What it controls
EXT-X-STREAM-INF Multivariant BANDWIDTH 5000000, RESOLUTION 1920x1080 One rendition in the ladder; the player's adaptive bitrate logic chooses between these entries
EXT-X-TARGETDURATION Media 6 Maximum segment duration; sets the reload cadence and the three-target-duration live offset
EXT-X-MEDIA-SEQUENCE Media 48213 Sequence number of the first listed segment; lets the player track a sliding live window
EXTINF Media 6.006 Exact duration of the next segment URI
EXT-X-KEY Media METHOD AES-128 or SAMPLE-AES Encryption method and key location for the segments that follow
EXT-X-SERVER-CONTROL Media (LL-HLS) CAN-BLOCK-RELOAD YES, PART-HOLD-BACK 3.0 Enables blocking reloads and sets how far behind live the player holds in low-latency mode
EXT-X-PART Media (LL-HLS) DURATION 1.0 A partial segment published before its parent segment completes
EXT-X-ENDLIST Media No value Marks the playlist complete, turning a live stream into video on demand and making the playlist fully cacheable

Target duration is the tag that matters most: it sets latency, playlist reload frequency and edge request rate at the same time.

Why HLS became the default format for video at scale

HLS won because it turned video into static files. Once segments are plain HTTP objects, delivery needs no special streaming servers: existing CDN caches, TLS termination and HTTP/2 multiplexing all work unchanged. Native playback on iPhone, iPad, Apple TV and Safari made HLS the one format every large service had to ship anyway. Apple's App Store guidelines also require HLS for video longer than 10 minutes streamed over cellular.

The same property means HLS scale is mostly a caching problem. What matters is how well your delivery layer handles immutable segments, short-TTL playlists and query-string cache keys for LL-HLS. BlazingCDN's HLS Streaming CDN for pre-encoded HLS, LL-HLS and DASH serves that packaged output, with origin shield and request coalescing included on every plan. Transcoding and packaging stay upstream in your own pipeline.

Common HLS misconceptions, corrected

  • "HLS means MPEG-TS." fMP4 and CMAF segments have been valid since 2016, and most new pipelines use them to share segments with DASH.
  • "Shorter segments are always better." Each segment must open with a keyframe, so 1-second segments cost compression efficiency, multiply requests and make bitrate switching noisier. Use LL-HLS parts instead of tiny segments.
  • "HLS is live only." Most HLS traffic is video on demand: a complete playlist with an end marker.
  • "HLS handles ingest." It does not. Encoders still push RTMP or SRT to a transcoder, and HLS appears only after packaging.
  • "A high overall hit ratio means playlists are fine." Segments dominate request counts and hide playlist misses. Measure the two cache classes separately.

FAQ: HTTP Live Streaming (HLS)

What is HLS used for?

HLS is used to deliver live and on-demand video over HTTP to phones, browsers, smart TVs and set-top boxes. Streaming platforms, broadcasters, online education and sports services rely on it because segments cache on any CDN, adaptive bitrate adjusts quality to each viewer's connection, and Apple devices play HLS natively without extra player code.

Is HLS better than MPEG-DASH?

Neither HLS nor MPEG-DASH is universally better; they solve the same problem with different manifest formats. HLS plays natively on Apple devices, while DASH is an open ISO standard popular on Android and smart TVs. With CMAF, one set of fMP4 segments can serve both, so many services publish HLS and DASH manifests side by side.

What latency does HLS streaming have?

Classic HLS streaming with 6-second segments typically runs 18 to 30 seconds behind live, because players start three target durations from the live edge. Low-Latency HLS uses partial segments of roughly 200 ms to 1 second and blocking playlist reloads, which brings latency to about 2 to 5 seconds while keeping delivery cacheable over HTTP.

Can a CDN cache HLS video?

Yes, a CDN caches HLS video well because segments are immutable HTTP files that can take long TTLs. Live media playlists change every target duration and need short TTLs, often half the target duration. For LL-HLS, the CDN must keep the _HLS_msn and _HLS_part query parameters in the cache key and support held playlist requests.

Measure your HLS delivery before the next live event

Pull one hour of edge logs from your last live stream and split requests into three classes: multivariant playlists, media playlists and segments. For each class, compute requests per viewer per minute and the cache hit ratio. Then compare your media playlist TTL with half your target duration. If playlist hit ratio trails segment hit ratio by a wide margin, or origin requests grow with audience size rather than rendition count, your shield or coalescing is not doing its job.

If you deliver packaged HLS or LL-HLS at volume, the BlazingCDN OTT and VOD streaming CDN page explains how segments are served. A 14-day testing period on real production traffic lets you rerun this same measurement on your own streams.