Netflix CDN Explained: How Open Connect and ISP Caches Work

Netflix CDN Explained: Open Connect and ISP Caches 2026

The Netflix CDN moves the overwhelming majority of its bytes from a server sitting inside your ISP's network, often two or three hops from your router, and it decided what to put on that server roughly 12 hours before you pressed play. That is the entire trick. Open Connect is not a faster edge than everyone else's edge; it is an edge that already knows what you are going to watch. This article walks the full request flow from client API call to first segment, breaks down fill windows and steering ranking signals, explains the ISP partnership model that makes the whole thing free for carriers, and closes with a concrete playbook for platforms that will never have Netflix's catalogue concentration.

image-2

What the Netflix CDN actually is in 2026

Netflix content delivery network is a purpose-built, single-tenant delivery fabric. There is no multi-tenant edge, no shared cache, no generic PoP hosting third-party origins. Every appliance serves one catalogue, which is why the cache economics work: a few tens of terabytes of flash can hold effectively the entire high-demand working set for a given metro.

The split is clean. The control plane runs in AWS and handles authentication, the discovery UI, licensing, manifest generation, and steering decisions. The data plane is Open Connect Appliances (OCAs) deployed in two flavours: embedded inside ISP networks, and clustered at internet exchange points. As of 2026 the program spans thousands of ISP partners globally, with tens of thousands of appliances in the field. Playback bytes essentially never come from AWS.

That separation is the architectural point most write-ups miss. Netflix did not build a CDN. It built a scheduling system that happens to have storage attached.

How a Netflix playback request actually flows

Describe the sequence as a diagram in your head: client at the left, AWS control plane top-centre, a ranked set of OCAs at the right, ISP boundary drawn as a vertical line the video bytes never cross.

  1. Client to control plane. The app authenticates and requests playback for a title. The control plane resolves entitlement, DRM licence, and the available encode ladder (AV1, HEVC, VP9, H.264 depending on device class).
  2. Steering computation. The control plane takes the client's public IP, the resolver IP with EDNS Client Subnet where the resolver supports it, and the BGP-derived ASN. It intersects that with the current health and fill state of every OCA eligible to serve that subscriber.
  3. Ranked URL list returned. The client receives an ordered list of OCA hostnames, typically three or more, each already known to hold the requested title. Ranking is not just "nearest".
  4. Direct HTTPS fetch. The client opens TLS to the top-ranked OCA and pulls segments over HTTP. Adaptive bitrate logic runs entirely client-side.
  5. Failover. On connection failure, sustained throughput collapse, or HTTP error, the client walks down the ranked list without another control-plane round trip. Re-steering only happens on session refresh.

The consequence for capacity planning: a cache miss in the Netflix CDN is not a latency event visible to the user. The steering layer simply does not hand you a URL for an appliance that lacks the file. Misses are converted into a planning problem upstream.

What signals rank one OCA above another

  • Content presence. Binary gate. No file, no candidate.
  • Network proximity. Same-ASN embedded appliance beats IXP cluster beats transit-reachable cluster.
  • Appliance health. Disk error rates, in-flight connection count, CPU and NIC saturation, reported continuously.
  • Headroom against peak. An appliance approaching its egress ceiling is de-ranked before it starts degrading, not after.
  • ISP-declared policy. Partners can express preferences about which peering interface or site should absorb traffic.

Fill windows: the part that makes the Netflix CDN work

Every OCA has a nightly fill window, computed per-appliance against that network's local traffic curve. A typical window opens after the regional evening peak collapses and closes before the morning ramp, commonly a six to ten hour block. During that window the appliance pulls new and re-ranked content from a peer OCA in the same cluster, or from a regional fill source, or from origin as last resort. Peer fill is strongly preferred because it keeps bytes off transit.

What gets filled is decided by a popularity model that predicts, per metro, the next day's demand. Inputs include recent regional viewing, release calendars, promotional placement in the UI (Netflix knows what it will merchandise before you do), and cross-title similarity. Because the UI drives a large share of plays, Netflix's demand prediction is partly self-fulfilling. That is a structural advantage no third-party CDN can replicate.

Well-provisioned embedded sites sustain offload above 95% as of 2026. The remaining few percent is long-tail catalogue and unpredicted spikes, absorbed by IXP clusters.

Live changes the fill math

Netflix's live event slate through 2025 and into 2026 broke the pre-fill assumption outright. You cannot pre-position a boxing match. Live traffic on the Netflix CDN behaves like a conventional CDN workload: a tiered fill from encoder to regional cache to edge, with hit ratio determined by concurrency rather than prediction. The 2024 Paul-Tyson event and the NFL Christmas games exposed exactly this, and the subsequent engineering work went into hierarchical live fill and stricter admission control rather than into more storage. Expect the split-brain architecture — predictive for VOD, tiered-and-reactive for live — to persist.

OCA hardware and the software stack in 2026

ClassStorageSustained egressRole
Storage OCA~36 HDDs, 16–20 TB each~90 GbpsDeep catalogue, long tail
Flash OCANVMe, ~500 TB usable~160 Gbps and aboveHot titles, peak absorption
Global/IXP clusterMixed fleetAggregate multi-TbpsFill source, overflow, unembedded ISPs

Networking is multiple 100 GbE, with 400 GbE on newer builds. The software is a stripped FreeBSD image: asynchronous sendfile, kernel TLS so encryption happens without copying payload into userspace, and NUMA-aware disk-to-NIC pathing. Netflix engineers have publicly demonstrated 800 Gbps-class serving from single boxes in lab conditions; production appliances are deliberately provisioned well below that ceiling so a failed neighbour can be absorbed.

The ISP partnership model, and why carriers say yes

Netflix gives the hardware away. The ISP provides rack space, power, and IP transit-free connectivity; Netflix ships, owns, and remotely operates the appliance. The ISP gets a large fraction of its heaviest traffic class removed from paid transit and from internal backhaul. Netflix gets an edge it could never rent at that price.

Eligibility thresholds sit in the low single-digit Gbps of peak Netflix traffic for an embedded appliance; below that, the ISP peers with an IXP cluster instead. The commercial relationship is settlement-free in both directions, which has made it a reference point in every peering-dispute argument since 2014.

The cache hierarchy, end to end

AWS origin → regional or IXP OCA → embedded ISP OCA → (optionally) deeper access-layer node. Netflix has trialled very small caching nodes positioned near subscriber clusters in fibre networks, holding only locally hot titles and filling from the upstream embedded OCA. The economics only work where access-to-core links are expensive relative to the cost of a small appliance, which is a narrow but real set of deployments.

Failure modes: what breaks in a private CDN

This is the section most Netflix CDN explainers skip, and it is the one worth reading twice.

  • Silent fill failure. The worst class. An appliance reports healthy, serves old content fine, but missed last night's fill for a tentpole release. The steering layer correctly excludes it for that title, so the title's traffic concentrates on fewer appliances than planned. Symptom is not errors; it is unexpected load skew. Mitigation is redundant fill paths and per-title presence auditing, not health checks.
  • Steering misattribution. An ISP re-allocates IP space, or a resolver stops honouring EDNS Client Subnet, and subscribers get steered to a cluster three hundred kilometres away. Traffic still works. Quality drops and transit costs rise. Detect via per-ASN throughput distributions, not availability.
  • Thundering herd on release. Global simultaneous drops produce a demand spike on a title with zero viewing history. Prediction is useless; pre-positioning is mandatory. If pre-positioning slips, the shortfall lands on IXP clusters and transit.
  • Appliance loss during peak. Health telemetry de-ranks it within seconds, clients fail over on their next segment fetch. The real risk is the neighbour absorbing its load and exceeding its own headroom, so capacity planning assumes N-1 per site.
  • Live concurrency overshoot. Unlike VOD, live has no long tail to shed. Netflix's response has been admission control plus aggressive ladder trimming: drop the top rungs before you drop viewers.

Lessons for platforms that are not Netflix

The naive takeaway is "build your own CDN". For almost everyone that is wrong. The transferable lessons are cheaper than that.

  1. Treat hit ratio as a capacity metric, not a performance metric. If your architecture converts misses into user-visible latency, you have a different problem than Netflix does. Pre-warm anything with a known release time.
  2. Pre-position on your own schedule. If you know a game patch, a film, or a software release drops Tuesday 10:00 UTC, push it to edge overnight Sunday. Most commercial CDNs support prefetch or tiered-cache warming. Very few teams use it.
  3. Measure per-ASN, not per-region. Netflix's entire advantage is ASN-level placement. You can get a meaningful slice of that benefit by simply knowing which ASNs carry your worst rebuffer rates and choosing a provider with good peering there.
  4. Model egress cost per delivered hour, not per GB. Bitrate ladder decisions and egress pricing are the same decision. A 20% ladder trim at the top rung often moves cost more than any cache tuning.

On that last point: for teams doing the cost-per-TB math, the provider tier matters more than micro-optimisation. Within the high-volume, cost-at-scale set — Bunny.net, CDN77, KeyCDN, Gcore, Medianova, and Fastly for streaming-specific tooling — the differences are real. Fastly's programmable edge and log latency are genuinely strong for live workloads; Bunny.net is hard to beat on developer ergonomics at small scale. BlazingCDN's media delivery platform competes on NVMe SSD edge storage, 100% uptime, and pricing that starts at $5 per TB ($0.005 per GB) and falls to $2 per TB ($0.002 per GB) past 2,000 TB — $100/month covers 25 TB, $1,500/month covers 500 TB, $4,000/month covers 2,000 TB. For enterprises running sustained multi-petabyte VOD egress, that delivers stability and fault tolerance comparable to Amazon CloudFront at a materially lower cost per delivered hour, with flexible configuration and fast scaling during release spikes. Worth benchmarking against your current bill before the next contract renewal.

If you want the appliance-level detail behind this overview, our engineering blog covers the Open Connect deep dive and the contrasting YouTube edge model, which solves the same problem with a very different cache-admission philosophy.

FAQ

Does the Netflix CDN use AWS for video delivery?

No. AWS runs the control plane — authentication, search, recommendations, manifests, steering, billing — but playback bytes come from Open Connect Appliances inside ISPs or at internet exchanges. Origin fetch for video is a fill-path fallback, not a serving path.

Can other companies use Netflix Open Connect?

No. Open Connect is single-tenant and serves the Netflix catalogue only. ISPs can host appliances and peer with Open Connect clusters, but there is no mechanism for third-party content to be cached on them.

How does Netflix decide which OCA a viewer connects to?

The control plane returns a ranked list of appliance URLs computed from client IP, resolver subnet, ASN, appliance health, current headroom, and confirmed presence of the requested title. The client tries them in order and fails over locally without a new steering call.

Why does Netflix fill caches at night instead of on demand?

Off-peak fill uses network capacity that is otherwise idle and costs the ISP nothing incremental, while guaranteeing the content is local before demand arrives. It converts what would be a latency problem into a scheduling and capacity problem, which is far easier to engineer against.

How much Netflix traffic is served from embedded ISP caches?

Well-provisioned embedded deployments routinely exceed 95% offload as of 2026, meaning under 5% of that ISP's Netflix traffic crosses transit or peering. Networks without embedded appliances are served from IXP clusters instead, with correspondingly higher transit usage.

Does live streaming work the same way on the Netflix CDN?

No. Live cannot be pre-positioned, so it uses a hierarchical fill from encoder through regional caches to the edge, with hit ratio driven by concurrency rather than prediction. Netflix has added admission control and ladder trimming for live specifically because the VOD playbook does not apply.

Run this test before your next release

Pick your single largest scheduled content drop in the next 30 days. Before it ships, pull your CDN's per-ASN hit ratio and rebuffer rate for the first 60 minutes after the previous comparable release. Then pre-warm the same asset to edge 12 hours ahead and measure the same two numbers again. If the delta is under two percentage points, your bottleneck is not cache fill and you should stop tuning it. If it is above five, you have found free performance that costs nothing but a scheduler.

Open question for anyone running live at scale in 2026: at what concurrency does reactive tiered fill stop being cheaper than brute-force pre-positioning of the first ninety seconds of every stream? We have opinions. We would rather see your numbers.

Heavy traffic.
Light bill.

The CDN for video and large traffic