Blog
In this dedicated blog page, we invite you to explore a wealth of knowledge, expertise, and inspiration that transcends traditional boundaries.

CDN Optimization Starts With the Cache Key
Most CDN optimization work starts by raising TTLs and stops when the hit-ratio chart looks better. Then the origin egress bill does not move. The cause is almost always the cache key: the CDN is storing dozens of near-identical copies of the same object, and every one of them misses on first request. This piece covers what actually controls origin load — key cardinality, revalidation, and tiering — and the order to fix them in.
Cache hit ratio is the wrong number to run on
Cache hit ratio counts requests. Your origin bill counts bytes. The two diverge the moment object sizes vary, which is every real catalogue.
Fastly draws the distinction explicitly: origin offload is "the ratio of bytes served to end users that were cached inside the CDN (not fetched from the origin), over total bytes served to end users for the service" (Fastly blog, accessed 2026). One large video segment miss outweighs thousands of hits on thumbnails, and hit ratio will not show it. The same write-up documents a customer who disabled shielding: hit ratio stayed in the low 90s while origin load went from under 5 GiB/s to over 20 GiB/s. A metric that can miss a 4× change in origin load is not a control surface.
There is also a scaling effect people get backwards. Moving hit ratio from 90% to 95% is not a 5% improvement — the miss rate goes from 10% to 5%, so origin requests are cut in half. The last few points hold all the money, which is exactly why header changes alone rarely buy them.
What is CDN optimization? CDN optimization is the work of reducing how much traffic reaches your origin without serving stale or incorrect content. It has three levers: shrinking the cache key so fewer duplicate objects exist, setting freshness and revalidation headers so objects survive longer at the edge, and adding a tiered or shield layer so edge misses do not all reach origin.
In our migration work the first thing we ask for is not the hit-ratio dashboard. It is origin egress in bytes per hour, split by content type. That number tells you whether the problem is fragmentation, expiry or catalogue shape, and those have different fixes.
The cache key decides everything TTL cannot
A cache key is the string the CDN hashes to decide whether it already has your object. Anything in it that varies per user, per campaign or per browser multiplies your stored copies and drives every copy through a first-request miss.
Cloudflare's default key is the full URL including query string, plus the Origin header and the method-override and forwarding headers (x-http-method-override, x-forwarded-host, x-original-url and several others) (Cloudflare docs, accessed 2026).
Know which half of the fix is gated. Sort query string, Ignore query string, Cache deception armor and Cache by device type are ordinary Cache Rules settings. What Enterprise buys is the custom cache key: selecting named query parameters to include or exclude, and adding headers, cookies, host or user attributes to the key (Cloudflare docs, accessed 2026). So canonicalising parameter order is a toggle on most plans; surgically keeping ?variant= while dropping ?utm_source= is not, and below Enterprise that job belongs in your application or a Worker.
Here is what fragments a key in production, with the documented behaviour behind each:
That is the specification. The platform you are on may not implement it, and this is the single most consequential divergence in the whole article.
Cloudflare states it plainly: "By default, Cloudflare does not consider vary values in caching decisions." Vary values are respected only "when you configure the Cache Rules Vary setting, when Vary for images is configured, and when the vary header is vary: accept-encoding" (Cloudflare docs, accessed 2026). A response containing Vary: * always bypasses cache regardless of that configuration (Cloudflare docs, accessed 2026).
Read what that means for the common case. If you shipped Vary: Cookie on a cacheable page and assumed it protected per-user content, it did not fragment the cache and it did not disable it — by default the header was ignored and one stored object was served to everyone. That is a cross-user content leak, not a hit-ratio problem, and it fails in the direction that ends up in an incident report. Only Vary: * gives you the bypass. Treat Vary as a correctness control you must explicitly enable per platform, not as something the origin can assert and trust.
TTL and revalidation are a negotiation with the platform
RFC 9111 fixes the precedence. A cache calculates freshness lifetime from s-maxage if it is a shared cache, then max-age, then Expires minus Date, and only then falls back to heuristics (RFC 9111 §4.2.1, 2022). Where heuristics apply and a Last-Modified header exists, the spec encourages "no more than some fraction of the interval since that time. A typical setting of this fraction might be 10%" (RFC 9111 §4.2.2, 2022).
That precedence is the most underused lever on the list. Split the edge lifetime from the browser lifetime: hold an object at the edge for a day while browsers revalidate in five minutes, so a purge on publish reaches every user immediately and the origin still sees almost nothing.
Which header does the splitting depends on the platform, and this is where vendor-agnostic matters. On Varnish-derived stacks, Surrogate-Control takes precedence over everything else: Fastly documents its order of preference as Surrogate-Control: max-age, otherwise Cache-Control: s-maxage, otherwise Cache-Control: max-age, otherwise Expires (Fastly docs, accessed 2026). Surrogate-Control is also stripped before the response reaches the browser, which makes it the cleanest way to say "one hour at the edge, one minute in the browser" without the browser ever seeing the edge value. Use Surrogate-Control where it is honoured and s-maxage where it is not.
Platform defaults then sit on top of the spec, and they differ enough to matter:
That last row is the single most dangerous default in CDN configuration, and it is documented rather than hidden. It is also almost never mentioned in optimization guides, because it is a correctness risk rather than a performance tip.
Freshness expiry then creates a second problem: a synchronised miss. Every edge holding a copy expires at roughly the same second and every one of them fetches. Your origin sees a periodic spike with no traffic change behind it.
RFC 5861 solves that in one directive. stale-while-revalidate "indicates that caches MAY serve the response in which it appears after it becomes stale, up to the indicated number of seconds" (RFC 5861 §3, 2010), and stale-if-error allows a stale response when the origin returns 500, 502, 503 or 504 (RFC 5861 §4, 2010). AWS recommends supplementing max-age with both (AWS docs, accessed 2026).
How do you improve CDN cache hit ratio? Reduce the cache key first: strip tracking parameters, normalise query string order and case, forward only named cookies and headers, and keep Vary to the smallest set. Then extend freshness with s-maxage and add stale-while-revalidate so expiry does not force a synchronous origin fetch. Enable a tiered or shield layer last.
stale-if-error also converts a short origin outage into cached-content delivery rather than a 502 page. We treat it as an availability control, not a caching one, and it is the cheapest one available.
Then check that your platform will actually honour it. Cloudflare documents that "the stale-if-error directive is ignored if Always Online is enabled or if an explicit in-protocol directive is passed", and that those directives "include a no-store or no-cache cache directive, a must-revalidate cache-response-directive, or an applicable s-maxage or proxy-revalidate cache-response-directive" (Cloudflare docs, accessed 2026). Read that carefully: on Cloudflare, the s-maxage you just added to extend edge lifetime disables stale-if-error on the same response. The two most commonly recommended directives in CDN optimization advice cancel each other out on one of the largest CDNs, and no guide that treats headers as portable will tell you. Pick one per platform — s-maxage where you want edge/browser split and Always Online is off, stale-if-error where availability during an origin outage matters more — or use Surrogate-Control on stacks that honour it and leave s-maxage off the response entirely.
One more thing the bytes-not-requests thesis demands: a revalidation is not a miss. When an object goes stale and the origin answers a conditional request with 304 Not Modified, you paid one request and almost no bytes. That is why ETag and Last-Modified belong on every cacheable response, not just the ones you plan to soft-purge. A site with aggressive revalidation can show a mediocre hit ratio and an excellent egress bill at the same time, and it will look like a problem on the dashboard your vendor gives you.
Tiering and shielding beat every header change
If edge PoPs each fetch independently, one object costs you as many origin requests as you have caching locations. A tier in front of the origin collapses that.
CloudFront states it directly: "When you use Origin Shield, all requests from all of CloudFront's caching layers to your origin come from a single location. CloudFront can retrieve each object using a single origin request from Origin Shield, and all other layers of the CloudFront cache (edge locations and regional edge caches) can retrieve the object from Origin Shield" (AWS docs, accessed 2026).
Cloudflare's equivalent is Tiered Cache. Smart Tiered Cache, which picks a single closest upper tier per origin, is available on all plans; Generic Global Tiered Cache and Custom Tiered Cache topologies are Enterprise features, as is the Regional Tiered Cache middle layer on the Tiered Cache documentation (Cloudflare docs, accessed 2026). Verify this one against the live docs before you plan around it — Cloudflare is currently moving parts of the tiering stack under the Smart Shield brand, and the availability tables on the two sets of pages do not yet agree.
Most audits get the order wrong here. Teams spend three weeks tuning headers and never turn on the upper tier, when tiering is a toggle with a larger effect than every header change combined. Check one thing first: whether your platform bills the shield layer as its own request tier. Offload that arrives as a new line item is still worth having, but it is not free, and it belongs in the [CDN cost comparison → /blog-posts/cdn-pricing-comparison] before you commit.
Purging is a hit-ratio decision in disguise
Purge-all at deploy time is a self-inflicted origin incident: every edge misses simultaneously on the next request for every object.
Surrogate keys are the alternative — tag responses with keys and purge by tag. Individual keys are capped at 1,024 bytes and the total header at 16,384 bytes; beyond that "the key currently being parsed and all keys following it within the same header will be ignored" (Fastly docs, accessed 2026). Tag a category page with its product identifiers and a price change purges exactly the pages that display it.
Soft purge is the second half. It marks content outdated while keeping it usable — "stale objects remain available to use in some circumstances while Fastly fetches a new version from origin" (Fastly docs, accessed 2026) — and requires either ETag/Last-Modified on origin responses or stale_while_revalidate configured together with stale_if_error. Without that groundwork it does close to nothing.
The optimization sequence, in the order that pays
- Measure origin egress in bytes per hour, split by content type. Not hit ratio. If bytes are flat and requests are spiky, you have an expiry problem; if both are high, you have a key problem.
- Dump the actual cache key your platform computes for your top 20 URLs by origin fetches. Most fragmentation is visible in one look.
- Strip tracking parameters and normalise query string order and case before the key is computed.
- Reduce Vary to the minimum set, remove Vary: Cookie from anything you intend to cache, and confirm whether your platform honours Vary at all — on Cloudflare it is off by default and must be enabled in Cache Rules. Never rely on Vary alone to keep per-user content out of a shared cache.
- Split cache behaviours by path so static assets never carry cookies or session headers. This, not Vary, is what actually protects authenticated content.
- Split edge lifetime from browser lifetime — Surrogate-Control: max-age where it is honoured, otherwise s-maxage long with max-age short — and add stale-while-revalidate. Add stale-if-error only after checking the platform: on Cloudflare an applicable s-maxage or must-revalidate, or Always Online being on, disables it (Cloudflare docs, accessed 2026).
- Put ETag or Last-Modified on every cacheable response so stale objects revalidate into a 304 instead of a full refetch.
- Enable tiered caching or origin shield, then re-measure origin bytes — this is the step with the largest single effect.
- Replace purge-all with tag-based soft purge and wire it into the deploy pipeline.
- Re-baseline after seven days. Cache fill has to complete across the footprint before any number means anything.
What breaks in production
Logged-in HTML cached at the edge. A minimum TTL above zero overriding no-cache, no-store and private (AWS docs, accessed 2026) is how one user's account page gets served to another. This is the failure that ends careers, and it is a two-field misconfiguration.
Vary trusted as an access control. The same outcome by a different route. Cloudflare does not consider vary values by default (Cloudflare docs, accessed 2026), so a page shipping Vary: Cookie is stored as one object and served to every visitor. The dangerous part is that this looks better on your dashboard — hit ratio goes up — right up to the point someone sees another user's session. Path-based cache separation is the control; Vary is a hint your platform may ignore.
"We raised the TTL and nothing happened." On Cloudflare, HTML and JSON are not cached by default (Cloudflare docs, accessed 2026). The TTL was raised on content the CDN was never storing.
Compression variance removed instead of normalised. "Shrink the cache key" and "drop Accept-Encoding" read as the same advice and are not. Removing the variance without removing the variation ships a Brotli-encoded body to a client that asked for identity, or an uncompressed body to everyone. Normalise the header — CloudFront's cache policy compression settings collapse it to br,gzip (AWS docs, accessed 2026) — and never strip Vary: Accept-Encoding from a response whose body actually varies by encoding.
Soft purge that does nothing. Without ETag/Last-Modified or stale_while_revalidate and stale_if_error configured together, soft purge has nothing to serve from (Fastly docs, accessed 2026).
Header advice copied between platforms. The s-maxage plus stale-if-error pairing is the clearest case — correct on CloudFront, self-cancelling on Cloudflare (Cloudflare docs, accessed 2026). Caching headers are a specification; honouring them is a per-vendor decision. Test the response, do not trust the article, including this one.
Hit ratio improves, bill does not. Almost always fragmentation on large objects: the small files got better and the heavy ones still miss. Check bytes, not requests.
Cache fill measured too early. Any change to the cache key invalidates the entire stored set, so the first day after always looks worse. Judging it on day one is how good changes get reverted.
The ceiling no vendor will mention, and how to decide
Your hit ratio has a ceiling set by your catalogue, not your configuration.
If most of your objects are requested once — a large media library, user-generated storage, a long-tail download archive — edge caching adds a hop and a fill cost and cannot reach a high hit ratio at any TTL. The honest fixes are a tiered topology that concentrates fills, a shorter path between origin and users, or a provider whose pricing model suits long-tail delivery rather than one that prices as if you were a news site.
Every platform's optimization guide is written to make its own cache look effective. None will tell you that your content shape is the binding constraint, because the conclusion is sometimes "this is the wrong CDN for your workload" — see [how to switch CDN providers without downtime → /blog-posts/cdn-shutdown-migration-playbook] for what that migration costs. We say it because we sell across providers rather than one of them.
So take the branch that matches your measurement:
- Origin bytes high, hit ratio high. Fragmentation on large objects. Fix the cache key, starting with query strings and Vary.
- Origin bytes spiky on a fixed interval. Synchronised expiry. Add stale-while-revalidate and stagger TTLs across object classes.
- Origin requests high, bytes moderate. Every PoP fetching independently. Turn on tiering or shield before touching a single header.
- Hit ratio structurally low across everything. Catalogue shape, not configuration. Re-examine topology and provider fit, and read [where egress fees actually land → /blog-posts/cloud-egress-fees-comparison] before you renegotiate.
- Hit ratio fine, bill still wrong. The problem is the rate card, not the cache. Start from [optimizing CDN costs → /blog-posts/optimize-cdn-costs-strategies-and-best-practices].
What is a good cache hit ratio? There is no universal target, because the ceiling depends on how often your objects are re-requested. A media catalogue with long-tail access cannot reach the same ratio as a news site serving the same assets to everyone. Judge a hit ratio against your own baseline and against origin bytes, not against a published benchmark.
We run this sequence across whichever CDNs a client already uses, and the deliverable is the origin-bytes delta, not a hit-ratio chart. Having the cache key on your top-20 origin fetches dumped and read before you change anything is the place to start — and if the answer turns out to be topology rather than headers, a [multi-CDN routing setup → /blog-posts/multi-cdn-strategy] is the next conversation, not another round of TTL edits.
FAQ
Does raising TTL always improve cache hit ratio? No. TTL only matters for objects the CDN is already storing under a stable key. If tracking parameters or Vary headers are fragmenting the key, a longer TTL preserves more duplicate copies rather than fewer, and the first request for each still reaches origin. Fix key cardinality before touching freshness values.
What is the difference between cache hit ratio and origin offload? Cache hit ratio measures the share of requests served from cache. Origin offload measures the share of bytes served to users that did not come from origin (Fastly blog, accessed 2026). Because object sizes vary, the two can move in opposite directions. Origin offload tracks your egress bill; hit ratio does not.
Should I use max-age or s-maxage? Both. s-maxage applies to shared caches and takes precedence over max-age in them (RFC 9111 §4.2.1, 2022). Setting s-maxage long and max-age short holds objects at the edge while keeping browsers honest, so a purge on publish reaches every user without your origin absorbing the expiry.
Is origin shield worth enabling on a small site? Usually yes on request volume, and it should be checked on cost. A shield collapses many independent PoP fetches into a single origin request per object (AWS docs, accessed 2026). The caveat is that some platforms bill the shield layer as its own request tier, so confirm the rate card before assuming the offload arrives free.
Why did our cache hit ratio drop after a configuration change? Any change to the cache key invalidates every stored object, so the entire footprint refills from origin. That refill period looks identical to a regression. Hold the change for at least a full traffic cycle — seven days on most sites — before comparing, and compare origin bytes rather than hit ratio.
Can a CDN cache a page my application marked no-cache? Yes. CloudFront documents that a minimum TTL above zero causes caching "even if the Cache-Control: no-cache, no-store, or private directives are present in the origin headers" (AWS docs, accessed 2026). Audit minimum TTL on every cache behaviour that can serve authenticated paths before you trust your origin headers to protect you.

When Your CDN's IP Range Gets Blocked
A single blocked Hetzner IP address took 325 unrelated domains offline in Italy. One of them belonged to a Portuguese hosting provider, which lost email connectivity with its Italian customers for 16 days. Neither had anything to do with the pirated football stream the block was aimed at (Sommese, Sperotto, van der Ham, Affinito and Prado, 90th Minute: A First Look to Collateral Damages and Efficacy of the Italian Piracy Shield, University of Twente, peer-reviewed, CNSM 2025). On average, domains caught in that collateral damage stayed unreachable for around 320 days. CDN IP blocking is not a censorship story — it is a reachability risk sitting inside every anycast CDN contract, and almost nobody prices it.
The mechanic: anycast means sharing reachability with strangers
A hyperscale CDN serves thousands of customers from the same anycast IP ranges. That is the design — it is what makes anycast fast and what makes DDoS absorption possible. It also means the routable address your customers resolve to is shared with tenants you did not choose and cannot audit.
When a regulator blocks by IP rather than by domain, everything on that address goes with it. Your site does not need to be the target. It only needs to be a neighbour.
The clearest public example of the mechanic is a Cloudflare Community thread from July 2026: a site owner found their domain unreachable from Russia because it resolved to a shared Cloudflare IP that had been added to the state blocklist over a completely unrelated site sharing the same address. The browser simply spun — no error, no notification, connection silently dropped.
You cannot fix this by being compliant. Your legal standing is irrelevant to a /24 null route.
The Twente study put names to who actually gets hit. Of 7,114 domains affected during the study window, researchers manually confirmed 510 as legitimate sites with no connection to streaming — hotels, restaurants, retail shops, an accountant, a car mechanic, a nunnery, and a telehealth missionary program among them. In three separate cases, blocking one IP address caused 60 collateral blocks each. Mail server IPs were blocked in 782 cases and nameserver disruption in 397 — meaning the damage extends past the website to email and DNS.
And there is a second-order effect that should concern anyone leasing IPv4 space. Of the 10,918 blocked addresses, 24% were linked to leased address space, and 453 addresses were leased for the first time after they had already been blocked. The new tenant inherits an address that is silently unreachable from Italy, with no notification and no obvious cause. A further 250 addresses were re-leased to different companies while still under block. If you lease IPv4 space for a CDN origin or a mail server, check the block status of the range in your target markets before you announce it — the study's authors describe the polluted address space they leave behind as one of the system's lasting harms.
Four jurisdictions, one pattern
This is not a Russia story. It is a structural consequence of IP-level enforcement, and it is happening across regulatory contexts with very different intentions.
Sources: TorrentFreak (2025); RIPE Labs and APNIC Blog (2025) on the Piracy Shield study; TechRadar (2025); The Record (2024, 2025); DNS at Risk.
Note what Italy, Spain and Austria have in common: these are EU member states pursuing copyright enforcement, not censorship. The intent does not change the blast radius. Any business with customers in those markets carries the same exposure as one serving a heavily censored market — which is why this belongs in an infrastructure risk register rather than a politics conversation.
The Austrian outcome is the most useful signal in the table: a regulator that tried IP-level blocking on shared infrastructure and then banned it, because the collateral damage was indefensible. That is the direction of travel, but it arrives jurisdiction by jurisdiction, years apart.
Blocking happens at four layers, and each has a different fix
Teams conflate these, then apply the wrong countermeasure and conclude nothing works.
The ECH case deserves its own note because it is the single most common misdiagnosis in this space. Cloudflare enables Encrypted Client Hello by default; Russian networks have blocked TLS connections using the ECH extension since 2024. The result is a site that becomes unreachable in waves, which owners routinely misread as "Cloudflare is fully banned." It usually isn't. With ECH turned off, connections fall back to plain SNI and typically stop being dropped — a zone-level configuration change, not a migration.
Check ECH before you plan a CDN migration. It is free, reversible, and resolves a meaningful share of cases that look like IP blocking but aren't.
Why you find out last
There is no notification. No appeal path that operates on infrastructure timescales. In Italy, the blocklist behind Piracy Shield is not published — AGCOM has repeatedly refused freedom-of-information requests for it, leaving only single-IP lookup tools with no bulk export. Cloudflare told Human Rights Watch in May 2025 that it is generally unable to identify or confirm government-directed blocking and had received no notice from any Russian entity about the reported disruptions.
Read that carefully: the regulator will not tell you, and your CDN cannot tell you. Your first signal is a traffic drop from one country, and if that country is not a top-three market you may not notice for weeks.
Monitoring from inside the affected market is the only reliable detection. Synthetic checks from your own CI or from a US/EU monitoring region will pass while real users in the blocked market see nothing. If you have meaningful revenue from a market with an active blocking regime, you need a probe that actually resolves and connects from inside it.
What no CDN vendor will tell you
A large CDN cannot publish this article. Cloudflare cannot write "our IP ranges are subject to state-level blocking in several markets, so consider a provider with a smaller footprint." Akamai cannot either. The structural incentive runs the other way: their answer to concentration risk will always be more of their own product. That is not dishonesty, it is positioning — but it means the SERP for this problem is written entirely by parties who cannot recommend the fix.
The fixes that exist are gated behind Enterprise. Cloudflare's BYOIP, where Cloudflare announces IP space you lease or own across its locations, is Enterprise-only. The traditional onboarding path ran up to four to six weeks, involving addressing teams, network engineering, legal, and a Letter of Agency — Cloudflare has since launched a self-serve BYOIP API that removes much of that manual process. Either way, BYOIP is a posture you adopt in advance, not a lever you pull during an outage.
Not every "dedicated IP" product solves this problem, and the distinction matters. Cloudflare's Dedicated CDN Egress IPs sound like the answer and are not: they govern traffic from Cloudflare to your origin, for origin allowlisting and firewall lockdown, not the address your users resolve to. Buying them changes nothing about whether a regulator can reach you. The controls that affect the user-facing address are BYOIP, leased static IPs with address maps, and the certificate option below.
And leased static IPs solve only half the problem. They are still Cloudflare addresses. If a neighbour on your shared IP gets you blocked, a static address fixes it. If a regulator blocks Cloudflare's ranges wholesale — as happened across eastern Russia in March 2025 — a static Cloudflare IP goes down with everything else. Only address space that is yours, or a provider whose ranges are not listed, survives range-level blocking. Match the control to which of the two failure modes you actually face.
There is one cheap partial mitigation almost nobody uses. Business and Enterprise customers can reduce the number of Cloudflare IPs their domain shares with other customer domains by uploading a Custom SSL certificate. That does not give you a dedicated address, but it shrinks the neighbourhood — fewer strangers on your IP means lower probability of inheriting someone else's block. For a Business-plan site with exposure to an active blocking regime, it is the highest ratio of risk reduction to effort available.
And the honest limit on all of this: if a block is a lawful order directed at your own content in a market where you operate, the answer is legal and compliance work, not a different IP range. Infrastructure diversity solves collateral damage — being caught by someone else's enforcement. It does not solve being the target, and any provider who pitches it that way is selling you a problem rather than a solution.
What breaks in production
You migrate CDN and inherit a new set of neighbours. Moving from one hyperscale anycast provider to another swaps one shared-IP exposure for a different one. If the new provider's ranges are already listed in your target market, you have executed a migration for nothing. Check the destination's IP ranges against the market's blocklist status before cutover, not after.
Your monitoring says everything is fine. Covered above, and it is the most common failure. Region-blind synthetic monitoring is how a two-week outage in a secondary market becomes a quarterly revenue surprise.
Lowering TTLs after you need them is too late. If your records sit at 86,400 and a market goes dark, your fastest possible failover is a day away. Pre-lowered TTLs on anything exposed to a blocking regime are cheap insurance — the same discipline covered in our [CDN migration playbook → /blog/cdn-shutdown-migration-playbook].
Regional CDNs solve reachability and introduce other constraints. A provider with local presence in a market is far less likely to be blocked there, and often has better last-mile performance. It will also usually have a smaller global footprint, different feature coverage, and its own compliance obligations in that jurisdiction. That is a real trade, not a free win.
"Not blocked yet" is a depreciating asset. Selecting a provider purely because regulators have not reached it is a position with a shelf life. The durable property is not obscurity — it is how fast you can move when it changes, which is a procurement and architecture question rather than a provider-selection one.
The mitigation ladder
Work down this list; each step costs more than the one above it.
- Check ECH first. If the symptom is intermittent unreachability in a specific market, disable ECH on the zone and re-test before doing anything else. Free, reversible, resolves a large share of misdiagnosed cases.
- Deploy in-market monitoring. You cannot manage exposure you cannot see. One probe resolving and connecting from inside each market that matters.
- Upload a Custom SSL certificate if you are on a plan that supports it, to reduce the number of domains sharing your IPs.
- Pre-lower TTLs on any hostname serving a market with an active blocking regime.
- Keep a second CDN configured and warm, carrying real traffic at a small percentage rather than sitting idle with stale config and expired certificates.
- Evaluate a regional provider for markets where reachability is a revenue-critical requirement, accepting the feature and footprint trade honestly.
- Consider BYOIP or leased static IPs if you have Enterprise scale and the market exposure justifies it — static IPs for neighbour risk, BYOIP for range-level risk. Start the process before you need it.
The decision framework
Three questions decide how much of this ladder applies to you.
- Do you have revenue from a market with an active IP-blocking regime? If no, steps 1 and 2 are still worth doing and the rest is over-engineering. If yes, you need at least through step 5.
- Is the exposure collateral or direct? Collateral — someone else's content on your shared IP — is an infrastructure problem with infrastructure fixes. Direct — a lawful order about your content — is a legal problem, and no CDN change resolves it.
- How fast can you currently switch providers? Not "do you have a second CDN," but: certificates provisioned, config parity verified, TTLs low, contract already in place. If the honest answer is measured in weeks, that is the gap to close, regardless of which provider you use today.
INXY brokers CDN capacity across multiple providers — including regional and specialist networks that do not appear on hyperscaler comparison pages — on a single contract and a single invoice. The practical consequence is that adding or switching a provider is a configuration decision rather than a procurement cycle, which is precisely the capability this risk requires. We will also tell you when your problem is ECH rather than IP blocking and costs nothing to fix, because a migration you did not need is not a sale worth making.
If you have a market where traffic dropped and you are not sure why, send us the affected hostnames and the market. [Request an infrastructure audit → /book-a-demo]. To review delivery options: [CDN, streaming and data delivery → /data-delivery].
This article describes publicly reported regulatory and network conditions for general information. It is not legal advice — obligations differ by jurisdiction and by the nature of the content being served, and decisions with regulatory consequence should be reviewed by qualified counsel in the relevant market.
FAQ
Why is my site blocked in a country when I have done nothing wrong? Hyperscale CDNs serve many customers from shared anycast IP ranges. When a regulator blocks by IP address rather than by domain, every site on that address becomes unreachable regardless of its content. Documented cases include a single blocked IP disrupting 325 unrelated domains in Italy.
Is Cloudflare blocked in Russia? Not fully. Sites behind Cloudflare load from Russia but unreliably, due to ECH connection blocking, shared-IP collateral damage, and targeted blocks of individual resources. Roskomnadzor has officially recommended against foreign CDNs, which is a recommendation rather than a prohibition, but a clear regulatory signal.
What is ECH and why does it break my site? Encrypted Client Hello is a TLS extension that encrypts the server name during the handshake. Cloudflare enables it by default, and some national networks block TLS connections using it. This produces intermittent unreachability that looks like IP blocking. Disabling ECH on the zone usually restores connections.
How do I know if my CDN's IP ranges are blocked somewhere? Monitor from inside the affected market. Synthetic checks running from US or EU regions will pass while real users in the blocked market see nothing. Neither the regulator nor your CDN will notify you — in Italy the blocklist itself is not published, and Cloudflare has stated it generally cannot confirm government-directed blocking.
Does switching CDN providers fix IP blocking? Only if the destination's IP ranges are not already listed in that market. Moving between two hyperscale anycast providers swaps one shared-IP exposure for another. Regional providers with local presence are less likely to be blocked in their own market, at the cost of smaller global footprint and different feature coverage.
Can I get dedicated IP addresses from a CDN? Some providers offer it at Enterprise tier. Cloudflare's BYOIP announces IP space you lease or own across its network, and leased static IPs with address maps fix which addresses your domain resolves to. Note that Dedicated CDN Egress IPs are a different product governing Cloudflare-to-origin traffic, not the user-facing address.

EU Data Act: Your Cloud Exit Gets Free in January 2027
From 12 January 2027, a cloud provider serving EU customers cannot charge you to leave. Not a reduced fee, not a cost-recovery fee — nothing (Article 29(1), Regulation (EU) 2023/2854). That date is a hard deadline in a phased regime that has already been running since September 2025, and most infrastructure teams are treating it as a legal department problem. It isn't. The EU Data Act switching rules change what you should sign this year, what leverage you have at renewal, and — critically — what they don't touch on your monthly bill.
Sources: LEXIA (July 2026); Turing Law (2025); Garrigues (2025).
As of mid-2026, the full free-of-charge regime is not yet in force (LEXIA, 2026). You are currently in the middle stage: your provider may still charge you to migrate, but only at genuine cost, and only if it was disclosed in your contract before you signed.
What "reduced switching charges" actually permits
This is the part worth reading carefully, because it's where providers have room to manoeuvre until 2027.
Article 29(3) limits interim charges to costs that do not exceed those incurred by the provider and directly linked to the switching process concerned. In practice that means bandwidth to transfer the data, technical assistance time, and use of specific export tools are chargeable — while profit margins, general infrastructure costs, and unrelated penalties are not (Garrigues, 2025). Recital 89 further indicates that certain costs cannot be passed to the customer, such as those arising from services the provider itself outsources.
The practical test: ask your provider to itemise the switching charge. If the line items don't map to bandwidth, engineering hours, and tooling, the charge is likely not compliant with the interim regime — and you can say so in writing before you pay it.
What survives 2027 — and this is the part nobody leads with
The headline is "switching becomes free." Four categories of charge survive that headline, and together they determine whether your bill actually changes.
Early termination penalties remain legal. Article 29(4) preserves proportionate early termination penalties for fixed-duration contracts, and Recital 89 backs this. The Act does not clearly define what degree of penalty is permissible (Lexology, 2025). So a three-year commit with an exit penalty is still a three-year commit with an exit penalty in 2027 — the switching fee goes away, the termination penalty does not. Providers must take care that early termination penalties are not characterised as switching charges, but that is a line drawn case by case, not a bright rule (Penningtons, 2026).
Multi-cloud egress stays chargeable. Article 34(2) permits charging for data egress in in-parallel use — running services across two providers simultaneously — provided the charge does not exceed the actual cost. Recital 99 addresses multi-cloud deployment strategies specifically (Penningtons, 2026). If your architecture runs continuously across two providers, January 2027 does nothing for that line item. This is the single most common misreading of the Act among engineering teams, and it matters because multi-cloud egress is frequently the larger number.
Standard service fees are untouched. The Act regulates the cost of leaving, not the cost of staying.
And the cost will look for somewhere to live. Industry observers expect providers to restructure pricing during the transition window, moving residual switching costs into base subscription fees (Global Law Experts, 2026). A prohibition on a specific charge does not remove the commercial instinct behind it.
The three largest providers moved ahead of the deadline with free-exit programmes — Google on 11 January 2024, AWS on 5 March 2024, Microsoft on 13 March 2024 (UK Competition and Markets Authority, Appendix N). Those programmes are one-time and conditional on actually leaving. They are not a discount you can apply to a running workload.
The contract terms the Act now forces — check yours against these
Chapter VI imposes specific contractual requirements. These are the clauses that should now be in any in-scope contract, and their absence is a flag:
- Notice period of no more than two months for the customer to initiate a switch (Article 25(2)(d)).
- A transitional period of 30 days, extendable where technically infeasible.
- A data retrieval period of at least 30 calendar days, starting after the termination of the transitional period (Penningtons, 2026).
- Pre-contract disclosure of standard service fees, early termination penalties, and any reduced switching charges — Article 29(4) requires this before you sign, and Article 29(6) requires it to be publicly available on the provider's website or another easily accessible place.
- Disclosure where switching is highly complex, costly, or impossible without significant interference in your data or service architecture (Article 29(5)).
- For IaaS: functional equivalence — the provider must supply capacity, documentation, technical support and, where necessary, tools to facilitate the transition (Articles 30 and 34(2)).
- For SaaS: open interfaces free of charge, equally available to all customers and to the destination provider, plus compatibility with open interoperability specifications or, failing that, a structured, commonly used, machine-readable export format (Article 30(5)).
Point 5 deserves attention. A provider disclosing that switching from its service is "highly complex" is making a compliance statement — and simultaneously handing you a written admission of lock-in that you can use at renewal. It is the first clause INXY reads when reviewing a client's existing provider agreements, because it tells you what the provider already knows about its own portability.
What no vendor will tell you
Your existing contract is probably non-compliant, and that is leverage, not a problem. The majority of Chapter VI obligations already apply, and existing contracts are likely to be non-compliant, with regulatory enforcement and customer scrutiny expected to increase through 2026 (Penningtons, March 2026). Providers are working through contract remediation right now. A customer who arrives at a renewal conversation already knowing which clauses are missing is negotiating from a materially stronger position than one who doesn't — because the provider has to fix those clauses regardless, and would rather do it as a concession than as a correction.
Throttling and performance limitations count as obstacles. Legal analysis of the Act flags that technical and contractual policies which obstruct switching — performance throttling during migration, exclusivity terms, penalties — need to be progressively adapted to the portability framework (Garrigues, 2025). Article 28 requires good faith cooperation and prohibits obstacles to switching. If your provider's export path is technically available but throttled to the point of impracticality, that is a compliance question, not just an engineering annoyance.
Enforcement is national, not centralised. Enforcement is delegated to Member States, which must implement penalties that are effective, proportionate and dissuasive. Customers can lodge complaints with the relevant national supervisory authority, and have a right to judicial remedy if that authority fails to act (Lexology, 2025). This means the practical strength of your position varies by which Member State's authority has jurisdiction — worth knowing before you rely on it.
What breaks in production
Treating the free-exit programme as a migration budget. It covers data transfer out. It does not cover dual-running two environments, engineering time, or the period where you pay both providers. The regulation removes a toll; it does not fund a project.
Assuming the Act covers your provider. The obligations apply to providers of "data processing services." Providers should evaluate whether their offerings fall within scope (Lexology, 2025) — and so should you, before building a switching plan on the assumption that a niche or non-EU-established vendor is captured.
Confusing the 2027 date with your renewal date. The free-switching regime applies from 12 January 2027 including to contracts already in force. But if you sign a three-year fixed term in 2026 with an early termination penalty, that penalty survives 2027 intact. The date that constrains you is your own commit length.
Planning a switch without the 30-day retrieval window. The minimum retrieval period starts after the transitional period ends. Teams that plan the cutover but not the retrieval window discover the data export clock hasn't started when they thought it had.
What to do this quarter
- Inventory which contracts are in scope and when each renews. The renewal calendar, not the regulatory calendar, is your action timeline.
- Request itemised switching charges from any provider currently quoting them. If the items don't reduce to bandwidth, engineering hours and tooling, challenge them in writing under Article 29(3).
- Audit each contract against the seven-point clause list above. Missing clauses are both a compliance gap for the provider and a negotiating lever for you.
- Separate your exit cost from your run cost. Model them independently — the Act addresses one and not the other, and conflating them produces a business case that won't survive scrutiny.
- If you run multi-cloud, model that egress separately. Article 34(2) leaves it chargeable. Any savings forecast that assumes 2027 zeroes it out is wrong.
- Check which Member State authority has jurisdiction over your provider relationship, before you need to rely on it.
The decision framework
The Act changes the cost of leaving, which changes the value of staying — but only if you act on it.
- If you are signing new in the next 12 months: push for short commit terms. The switching fee is disappearing anyway; the early termination penalty is what will actually hold you, so that is the clause to negotiate hardest.
- If you are renewing: arrive with the clause audit. Compliance remediation is work the provider owes you regardless — trade it for terms rather than accepting it as a favour.
- If you are already planning to leave: the interim regime means you may still be charged, but only at cost, and only if disclosed pre-contract. Check both conditions before paying.
- If you run multi-cloud in parallel: none of the above materially reduces your ongoing egress. Treat that as a delivery-architecture problem, covered in our [breakdown of cloud egress pricing → /blog/cloud-egress-fees-comparison].
INXY brokers cloud, dedicated and colocation capacity across multiple providers on a single contract, which means the exit path is a procurement question rather than a legal one — we can move a workload between providers we already hold agreements with, without you running a fresh contract negotiation each time. If you want your current provider contracts checked against the Chapter VI clause list before your next renewal, send them over. [Request an infrastructure audit → /book-a-demo]. To review capacity options: [dedicated and cloud servers → /hosting-solutions].
This article describes regulatory requirements for general information. It is not legal advice — INXY is an infrastructure consultancy, not a law firm, and contract decisions with legal consequence should be reviewed by qualified counsel in the relevant Member State.
FAQ
Does the EU Data Act ban egress fees? It bans switching charges, including egress charged for the purpose of moving to another provider, from 12 January 2027. It does not ban egress charges for ongoing use, and Article 34(2) explicitly permits charging at cost for egress in in-parallel multi-cloud deployments. Your monthly running bill is unaffected.
When do the EU Data Act switching rules take effect? Chapter VI has applied since 12 September 2025, requiring cost transparency. From 12 January 2026, switching penalties are prohibited and migration charges reduced to the strict minimum. From 12 January 2027, switching becomes entirely free, including under contracts already in force.
Can my cloud provider still charge me to leave in 2026? Yes, but only reduced charges not exceeding the provider's own costs directly linked to the switch — bandwidth, technical assistance time, and export tooling. Profit margin and general infrastructure costs are excluded, and the charge must have been disclosed in your contract before signing.
Do early termination penalties survive the 2027 deadline? Yes. Article 29(4) preserves proportionate early termination penalties for fixed-duration contracts, and the Act does not precisely define what level is permissible. A long commit term signed today will still carry its exit penalty after switching fees are abolished, which makes commit length the clause worth negotiating.
What contract terms does the EU Data Act require? A notice period of no more than two months, a 30-day transitional period, a data retrieval period of at least 30 calendar days after that, pre-contract disclosure of fees and penalties published accessibly, functional equivalence for IaaS, and free open interfaces with machine-readable export formats for SaaS.
Who enforces the EU Data Act? Enforcement is delegated to EU Member States, which must implement penalties that are effective, proportionate and dissuasive. Customers can complain to their national supervisory authority and pursue judicial remedy if that authority fails to act, so practical enforcement strength varies by jurisdiction.

IPv4 Addresses Are Now a Real Line on Your Hosting Bill
OVHcloud raised additional IPv4 from $2.00 to $2.40 per IP per month effective 1 April 2026 (CDNsun, 2026). AWS charges $0.005 per hour per public IPv4 address — attached or idle — which works out to roughly $3.60 a month each. On the open market, IPv4 leases run around $0.38–$0.50 per IP per month (IPXO, 2026; LARUS, 2026). None of those numbers matter on one server. Across a few hundred endpoints, load balancers and NAT gateways, IPv4 address cost stops being a rounding error and becomes a line item you have to defend in a budget review.
Why this became a cost centre
The supply side is settled and has been for years. All five Regional Internet Registries have exhausted their free pools: APNIC in 2011, RIPE NCC in 2012, LACNIC in 2014, ARIN in 2015, with AFRINIC and APNIC now rationing under community policies (APNIC; AFRINIC). The APNIC Labs address report confirms the remaining RIR pools are effectively empty — ARIN's stands at 0.0000 /8s as of September 2026 (APNIC Labs, 2026).
What changed recently is not supply — it's that providers stopped absorbing the cost. Both hyperscalers and European providers increasingly treat IPv4 as a scarce asset rather than a bundled utility (CDNsun, 2026). AWS began charging for public IPv4 in February 2024. OVHcloud's April 2026 increase applied to a limited set of bare metal servers, the VPS 2026 range, and additional IPv4 specifically. Hetzner's IPv4 pricing has moved upward across recent cycles, and its cloud calculator now surfaces an IPv6-only configuration as an explicit cost-saving option worth €0.50/month per server (CostGoat, 2026).
Transfers keep the market liquid without adding supply: roughly 33.4 million IPv4 addresses were recorded as transferred globally in 2025, up from 30.2 million in 2024, with RIPE NCC reporting over 16.7 million transferred between January and July 2026 (NRS, 2026). Transfers move existing scarcity; they do not create new addresses.
Sources: CDNsun (2026); IPXO (2026); LARUS (2026); IPv4Center (August 2026); IPbnb (2026); Atal Networks (2026).
Two mechanics the table hides. Smaller blocks cost more per address — one provider's published rates work out to roughly $0.586/IP at /24 versus $0.469/IP at /22, so scaling up the block size lowers the unit rate meaningfully (Atal Networks, 2026). And the buy-versus-lease break-even is long: at $150/month for a /24 against a $9,000–$11,500 purchase price for the same block, break-even lands around 60 months (Atal Networks, 2026). Five years is longer than most infrastructure planning horizons, which is why leasing has become the default for anything other than permanent core allocations.
Regional premiums are real. APNIC-registered blocks command lease rates above $0.60/IP/month in some markets due to supply constraints, while RIPE and ARIN blocks generally trade lower (IPXO, 2026; IPbnb, 2026).
The RIR waiting list is not a plan
ARIN's waiting list is still active in 2026, with a maximum aggregate qualification of a /22 — 1,024 addresses — and organisations already holding more than a /20 equivalent generally excluded (PubConcierge, 2026). It does distribute: on 2 July 2026, ARIN fulfilled 307 waiting-list requests using 199 cleared blocks (PubConcierge, 2026).
But there is no fixed waiting time, because fulfilment depends on which blocks become available and how they match queued requirements. Receiving space through a specified-recipient or inter-RIR transfer while on the list removes you from it. For a deployment with a date attached, the waiting list is a lottery ticket, not a procurement route.
What no vendor will tell you
The cheapest IPv4 you can find may be the most expensive thing on your network. IP reputation is the variable nobody quotes in the price. ARIN publishes the blocks it clears for waiting-list distribution and actively encourages blocklist operators to remove stale reputation data associated with previous registrants — because a range returning to the registry may previously have carried completely different traffic under a different owner (PubConcierge, 2026). Registration changes faster than third-party reputation data does. A block that is technically clean at the registry can still be listed in blocklists that haven't caught up, and you will discover this through mail deliverability failures or traffic being silently dropped by networks you don't control.
Serious lease providers screen against blocklist databases — the guidance in the market is to choose one screening against 100+ databases (IPv4Center, 2026). If a quote is meaningfully below the $0.38–$0.50 band, the question to ask is what the block was doing before you got it. Reputation screening is the step INXY runs before a block reaches a client's production traffic, precisely because the remediation cost dwarfs the price difference that made a cheap block attractive.
A registry transfer is not a working IP block. RIPE NCC notes that reverse DNS and RPKI may be affected during an inter-RIR transfer, and existing ROAs associated with resources leaving the RIPE region may be removed, requiring new routing authorisation through the receiving RIR (i.lease, 2026). A completed registry record is the start of the work, not the end of it. Budget for ROA re-creation, rDNS delegation, route objects and BGP announcement, and confirm which of these your lease or purchase actually includes — some providers bundle LOA, route object, inetnum record, RPKI/ROA configuration, rDNS delegation and WHOIS/geolocation updates; others hand you a registry entry and wish you luck.
Idle addresses bill exactly like busy ones. AWS's $0.005/hour applies per public IPv4 address whether attached to a running instance or not. Orphaned Elastic IPs from decommissioned environments are one of the most common findings in a cloud cost audit, and they are pure waste.
What breaks in production
IPv6-only looks free until something upstream isn't dual-stacked. Hetzner's own guidance suggests fronting multiple IPv6-only backends with a single IPv4 reverse proxy or load balancer (CostGoat, 2026), which is the correct pattern — but it only works if every external dependency your application calls is reachable over IPv6, and many third-party APIs and legacy partner integrations still are not. Test the full outbound dependency graph before committing to IPv6-only backends.
NAT gateways replace one cost with a bigger one. Collapsing many public IPs behind NAT removes the per-address charge and introduces per-gigabyte processing. On AWS that is $0.045/GB through the gateway, plus the gateway's own hourly charge, and internet-bound traffic pays gateway processing and standard egress — a mechanic covered in our [breakdown of cloud egress pricing → /blog/cloud-egress-fees-comparison]. For traffic-heavy workloads this is frequently worse than the IPv4 charge it replaced, so model both before switching.
Geolocation lags the transfer. A leased block registered in one region may still geolocate to its previous location in third-party databases for weeks. If you are using IP geography for content delivery, compliance routing or fraud scoring, verify geolocation accuracy before cutover rather than after.
Price increases arrive with the renewal, not the invoice. Hetzner's June 2026 cloud repricing applied only to new orders and rescales — existing servers left alone kept their old price (Safi, 2026). The corollary is that a routine resize can reprice the whole instance. Check what a rescale does to your rate before treating it as a no-op.
How to cost IPv4 properly
- Count every public address, including idle ones. Orphaned allocations from decommissioned environments bill at full rate. This is usually the fastest saving available.
- Separate addresses that genuinely need to be public from those that are public by default. Backend services behind a load balancer rarely need their own routable address.
- Compare the add-on rate against the market rate. At ~$2.40–$3.60/IP/month from providers versus ~$0.38–$0.50 on the lease market, the spread is roughly 5–8x. Above a few dozen addresses, that spread justifies the operational overhead of leasing and announcing your own block.
- Check whether you can announce your own space. Leasing only beats the add-on rate if you can run BGP and announce a /24 — the common minimum announceable and transferable block size. Without that capability, you are buying the provider's add-on whether you like it or not.
- Price the block size, not the address. Per-IP rates drop materially from /24 to /22 to /20. If your 18-month plan needs more addresses, buying the larger block now is often cheaper per address than adding incrementally.
- Verify reputation, routing readiness and geolocation before production. Blocklist status, RPKI/ROA configuration, rDNS delegation, and ASN history. A block that fails any of these costs more to remediate than the price difference that made it attractive.
- Model IPv6-only for new workloads. Not a migration of what exists — a default for what you build next. That is where the cost avoidance actually compounds.
The decision framework
- Under ~20 addresses, no BGP capability: pay the provider add-on. The operational overhead of leasing and announcing isn't worth a 5–8x spread on a small absolute number.
- Above that, with BGP capability: lease. The spread is real, the lead time is short, and month-to-month terms preserve flexibility.
- Permanent core allocation, 5+ year horizon: purchase is defensible, given a ~60-month break-even. Below that horizon it is not.
- New greenfield workload: design IPv6-only with a dual-stack ingress point, and treat IPv4 as an edge concern rather than a per-host requirement.
INXY brokers dedicated servers and colocation across multiple providers, which means the IPv4 question gets priced as part of the deployment rather than discovered as an add-on line after you've committed to the hardware — and where a workload can be architected IPv6-only behind a dual-stack edge, we will say so rather than sell you addresses you don't need. If you want your current public IP footprint costed against both provider rates and the lease market, send us the inventory. [Request an infrastructure audit → /book-a-demo]. To review capacity options: [dedicated servers, colocation and racks → /hosting-solutions].
FAQ
How much does an IPv4 address cost in 2026? Provider add-ons run roughly $2.40 per IP per month at OVHcloud and about $3.60 at AWS, billed hourly whether attached or idle. Open-market leases typically run $0.38–$0.50 per IP per month, with quoted ranges from $0.30 to $0.60 depending on block size, registry region and IP reputation.
Is it cheaper to lease or buy IPv4 addresses? Leasing is cheaper on any horizon under about five years. Purchase prices run $18–$45 per address, putting a /24 at roughly $9,000–$11,500, against lease rates that reach break-even near 60 months. Leasing also delivers in days rather than the weeks a registry-mediated transfer requires.
Why are IPv4 addresses getting more expensive? Supply is fixed — all five regional registries exhausted their free pools between 2011 and 2015. What changed is that providers stopped absorbing the cost and began billing IPv4 as a scarce asset. AWS started charging in February 2024, and OVHcloud and Hetzner both raised IPv4 pricing during 2026.
What is the minimum IPv4 block size I can announce? A /24, or 256 addresses, is the common minimum both for BGP announcement on the public internet and for transfer under several registry policies, including ARIN's minimum for transfer recipients. Smaller blocks are generally not routable globally, which sets the floor for leasing your own space.
Can I avoid IPv4 costs by going IPv6-only? Partly. IPv6-only backends fronted by a single dual-stack proxy or load balancer is the standard pattern and removes per-host IPv4 charges. It only works if every external dependency is reachable over IPv6, and many third-party APIs and legacy integrations still are not, so test the full outbound dependency graph first.
What should I check before deploying a leased IPv4 block? Blocklist and reputation status across multiple databases, geolocation accuracy, RPKI or ROA configuration, rDNS delegation, route objects, registry records, and the block's ASN and prior network associations. Registration changes faster than third-party reputation data, so a registry-clean block can still be listed elsewhere.

95th Percentile Billing vs Metered Bandwidth, Explained
A colocation quote says "$15 per Mbps, 95th percentile." A cloud invoice says "$0.09 per GB." A dedicated server listing says "20 TB included, unmetered after." None of these numbers can be compared to each other until you know what each one is actually measuring — and that gap is where 95th percentile billing gets misread most often, in both directions. This is what the four models actually do, where each one shows up, and the one traffic pattern that makes 95th percentile expensive instead of cheap.
What 95th percentile actually measures
The mechanic is older than most of the infrastructure it now bills. A provider samples your bandwidth every five minutes — 288 samples a day, roughly 8,640 across a 30-day month — then sorts all the samples from highest to lowest and discards the top 5%. Whatever sample sits just below that cutoff becomes your billed rate for the month (Wikipedia, "Burstable billing"; Kentik, 2026).
The consequence of that mechanic is the part worth internalizing: discarding the top 5% of a 30-day month gives you roughly 36 hours of unmetered headroom. You can burst to any level for a day and a half, spread across the month, and pay nothing extra for it (Wikipedia; major.io, 2026). That is the entire value proposition of the model — it absorbs short spikes without forcing you onto a higher committed rate, and it is why the model has been standard for internet transit and peering for over two decades.
Two details in the fine print change the bill more than the headline rate does. First, direction: some providers bill the 95th percentile of the higher of inbound or outbound traffic; others average the two. Billing on the maximum of the two directions, rather than the average, can push the effective rate meaningfully higher for asymmetric workloads — a video-heavy site with outbound-dominant traffic is billed differently depending on which convention the contract uses, so this is a line worth reading before signing, not after the first invoice. Second, what counts as traffic: some DDoS-protected bandwidth products exclude attack traffic from the measurement entirely and only meter legitimate post-scrubbing traffic — a distinction that matters enormously if you are DDoS-protected and assumed the meter includes everything hitting your network (Alibaba Cloud documentation, 2026).
If you have read our breakdown of [cloud egress pricing → /blog/cloud-egress-fees-comparison], the metered row is the one that piece covers in depth. This article covers the other three, which is where colocation and dedicated-server buyers actually spend their negotiating time.
The one traffic pattern that breaks the model
95th percentile billing rewards spiky traffic and punishes sustained traffic, and most people intuit it backwards.
Spiky is cheap. If your load is low most of the month with occasional bursts — a batch export job, a viral moment, a monthly report run — the discard window absorbs it. You are billed near your baseline, and the burst was functionally free.
Sustained is expensive, and here is the mechanic that surprises people: because the calculation only discards the top 5%, a workload that runs hot most of the time gets almost no benefit from the discard. Worse, a single unusually heavy day — a product launch, a Monday traffic pattern that repeats weekly, a marketing push — can set the billed rate for the entire month if it lands inside the 95% that doesn't get discarded, rather than the 5% that does (major.io, 2026; Wikipedia). Many sites see their heaviest single day of the week determine the whole month's bill, precisely because that pattern is common enough not to fall in the discarded top 5%.
This is the reverse of how metered billing works, where every byte costs the same regardless of when it moved. Under 95th percentile, when your traffic happens matters as much as how much of it there is.
Where each model actually shows up, and why
95th percentile dominates IP transit and peering because it was built for exactly that use case — bulk carriers with genuinely bursty aggregate traffic, where a fixed committed rate would either overcharge for headroom or undercharge for real peaks. It is also common in colocation contracts and DDoS-protected bandwidth, where the provider is reselling transit it bought the same way.
Metered per-GB dominates cloud hyperscalers because their cost structure and customer base look nothing like a transit carrier's — highly variable per-customer usage, no meaningful volume discard, and a billing system built to itemize everything. There is essentially no 95th-percentile cloud product at hyperscaler scale.
Unmetered and bundled dominate dedicated servers and VPS because the provider is selling you a port, not a byte count. A "20 TB unmetered after" server is really selling you a guaranteed connection speed with a soft usage ceiling attached for abuse prevention, not a precise 20 TB meter.
What no vendor will tell you
A provider that only offers a committed rate, with no 95th-percentile option, can be a signal about their network, not just their pricing philosophy. Practitioners who have run colocation deployments for years have flagged this directly: a datacenter that refuses to offer burstable billing, insisting on a flat committed rate instead, sometimes doesn't have the backbone headroom to absorb burst traffic across its customer base — burstable billing only works if the provider is genuinely oversubscribed in the statistical sense, banking on the fact that not all customers spike simultaneously (major.io, "Lessons learned: Five years of colocation," 2026). It's one of the first things INXY checks when qualifying a new bandwidth partner, and if a shortlisted provider won't do 95th percentile at all, ask why before assuming it's just a pricing preference.
The inbound-vs-outbound billing convention is the negotiating point nobody raises. A provider billing on the higher of the two directions, instead of the average, is charging more for the identical traffic pattern than a provider using the average convention — and this is rarely called out clearly in a sales conversation. For asymmetric workloads (heavy egress, light ingress, or the reverse), ask which convention applies before comparing quoted rates, because the headline $/Mbps number is not comparable across providers until you know this.
What breaks in production
Your own monitoring doesn't match the provider's meter. If you're graphing bandwidth with a different sampling interval, a different discard rule, or router counters instead of the provider's measurement point, your internal dashboard and your invoice will disagree — sometimes by a meaningful margin. Ask the provider for their raw sample export, not just the monthly summary number.
A recurring weekly pattern quietly becomes your permanent rate. A Monday-heavy traffic shape that repeats every week for a full month can settle into the 95th percentile as if it were your baseline, even though it only represents one day in seven. Reviewing daily, not just monthly, samples is the only way to catch this before it becomes a permanent line item.
"Unmetered" gets confused with "unlimited." A bundled dedicated server transfer allowance is almost always capped by port speed, and abuse thresholds exist even on plans marketed as unmetered. Sustained saturation of a "1 Gbps unmetered" port can trigger a conversation with the provider that a metered plan would never have caused, because metered billing has no concept of "too much," only "billed."
Billing method changes mid-relationship without much warning. Providers periodically retire or restrict specific burstable billing methods — Alibaba Cloud discontinued new sign-ups for its monthly 95th-percentile method for Anti-DDoS bandwidth in March 2026, moving customers toward the daily variant (Alibaba Cloud documentation, 2026). Check your contract's renewal terms for what happens if your specific billing method is deprecated.
How to pick correctly
Four questions, in order:
- Is your traffic genuinely spiky, or does it just look spiky on a monthly average? Pull daily, not monthly, granularity before assuming 95th percentile favors you. A workload that's "spiky" only because you've never looked at daily resolution is often sustained in disguise.
- What's the billing direction convention? Max of inbound/outbound, or average? This changes the effective rate for asymmetric traffic more than the headline number does.
- Do you need a hard ceiling, or genuine elasticity? Committed rate gives predictability at the cost of paying for unused headroom. 95th percentile gives elasticity at the cost of a bill that moves with your traffic shape.
- Is DDoS-protected bandwidth in the mix? If so, confirm explicitly whether attack traffic counts toward your meter. This is a materially different product depending on the answer, and it is not always obvious from the sales page.
Most of the disputes we see between infrastructure teams and providers on this topic aren't about the headline rate — they're about a billing convention nobody asked about until the first invoice. INXY brokers dedicated servers and colocation with the billing model matched to the actual traffic shape, not the default the provider happens to sell, because the "cheapest" $/Mbps number on a quote sheet is frequently not the cheapest bill once the direction convention and traffic pattern are accounted for.
If you're comparing quotes right now, send us the traffic profile — daily granularity, not monthly averages — and we'll tell you which billing model actually wins for your shape before you sign anything. [Request an infrastructure audit → /book-a-demo]. To see the underlying capacity options: [dedicated servers, colocation and racks → /hosting-solutions].
FAQ
What is 95th percentile billing? A method that samples your bandwidth every five minutes across the month, discards the highest 5% of those samples, and bills you at the rate of the next-highest sample. It gives roughly 36 hours a month of unmetered burst headroom, which makes it favorable for spiky traffic and unfavorable for sustained high usage.
Is 95th percentile billing cheaper than metered billing? It depends entirely on your traffic shape. For bursty workloads with real idle periods, 95th percentile is usually cheaper because the discard window absorbs the peaks. For sustained, consistently high traffic, metered per-GB billing can be more predictable, since 95th percentile offers little benefit when there's no meaningful low-traffic period to discard against.
Why did my bandwidth bill spike from one bad day? Under 95th percentile billing, a single unusually heavy day can set the billed rate for the entire month if it falls within the 95% of samples that aren't discarded. A weekly-recurring traffic spike, such as a Monday pattern, is especially prone to this because it repeats often enough not to land in the discarded top 5%.
What does "unmetered bandwidth" actually mean? It typically means there is no per-gigabyte charge up to the limit of your port speed, rather than truly unlimited transfer. The real ceiling is almost always the connection speed itself, and sustained saturation can trigger provider intervention even on plans marketed as unmetered.
Does 95th percentile billing measure inbound or outbound traffic? It depends on the provider's convention. Some bill on whichever direction (inbound or outbound) is higher in a given sample; others bill on the average of the two. For asymmetric traffic patterns, this convention can materially change the effective cost, so it is worth confirming before comparing quotes across providers.
Does DDoS attack traffic count toward my 95th percentile bill? It depends on the product. Some DDoS-protected bandwidth services explicitly exclude attack traffic from the billing measurement and only meter legitimate traffic that reaches your origin after scrubbing. This is not universal, so confirm it directly rather than assuming it from the marketing copy.

Data Center Power, Not GPUs, Now Picks Your Site
For two years, the story of data center expansion was GPU scarcity: hyperscalers and neoclouds fighting over allocation, facilities sitting half-built while silicon shipped late. That constraint has eased — TSMC has doubled its CoWoS advanced packaging capacity multiple times since 2024, and GPU shipment volumes have scaled with it (TrendForce, 2024). What has not eased is the grid. Interconnection queues, transformer manufacturing, and utility capital planning move on cycles measured in years, and none of them accelerated to match electricity demand (Inflect, 2026). In 2026, data center power availability, not chip supply, decides where you can deploy and how fast — and that constraint now reaches standard hosting and colocation buyers who have never touched a GPU.
The numbers behind the bottleneck
Securing grid power for a new data center in 2026 typically takes 24 to 72 months depending on market and load size, with some constrained regions quoted at 5 to 7 years (Inflect, 2026). Northern Virginia, Phoenix, and Dallas specifically run 4 to 7 years (Sightline Climate, cited by Bloomberg, May 2026).
The delivery risk compounds the wait. Goldman Sachs Research estimated in 2026 that only 60% of next year's scheduled capacity will arrive on time, dropping to 50% over the following two years — and Bloomberg separately reported Sightline Climate's estimate that 30–50% of the roughly 16 GW planned for 2026 in the US will be delayed or canceled due to power constraints (both cited May–July 2026).
Texas shows how concentrated the demand has become: ERCOT is managing a 410 GW large-load queue, with data centers comprising 87% of it, while CenterPoint Energy reported a 700% jump in large-load interconnection requests in a single year — from 1 GW to 8 GW (EnkiAI; Hanwha Data Centers, 2026).
What this means if you don't run GPUs
The instinct is to treat this as an AI-infrastructure story. It isn't only that. Power availability, not floor space, is now the binding constraint on new colocation supply in every major US market, according to Cushman & Wakefield's Data Center Power and Lease Pricing Outlook (cited across multiple 2026 industry reports). That constraint prices standard enterprise racks the same way it prices GPU racks, because both are drawing from the same interconnection queue and the same transformer supply chain.
The pricing evidence is consistent across independent sources. US wholesale colocation for 250–500 kW deployments averaged $195.94–$196.25 per kW per month in H2 2025, up roughly 6.5–6.6% year-over-year, with 3–10 MW requirements up 12.5% as competition intensified for large contiguous power blocks (Cushman & Wakefield data, cited across multiple 2026 reports; Brightlio, 2026). Retail — the single-cabinet deployments most hosting buyers actually lease — moved faster still: Lightyear's own platform data showed all-in per-kW pricing rising roughly 17–21% in just the second half of 2025 (Lightyear, 2026).
The market has also split by deal size in a way that matters more than price. Roughly 90% of new colocation capacity under construction is in the 200 MW to 1 GW hyperscale range, not the 2–4 MW halls a typical enterprise needs (Datacenter World, June 2026). North American vacancy sits near record lows — around 1–2% depending on the source — with most capacity under construction already pre-committed before a cabinet is installed (JLL data, cited by Inflect and Datacenter World, 2026). A deployment that once took 6 to 12 months to secure now needs 18 to 24 months of lead time in primary markets.
What no vendor will tell you
The honest version of this story isn't "AI is stealing your power" — it's that the people selling you colocation are caught in the same squeeze you are, and most won't say so directly. Peter Feldman, CEO of QTD Systems, a New York colocation provider focused on traditional enterprise workloads rather than AI, described the position plainly: "Non-AI or hyperscale colocation saw steady sales; pricing for operators held [with] maybe small upticks, but most incremental gains were countered by rapidly rising power costs. Equipment, replacement, and upgrade costs have skyrocketed due to tariffs and AI construction consuming all the equipment for new construction" (Data Center Knowledge, February 2026).
Read that carefully: a provider that isn't chasing AI customers at all is still absorbing rising input costs, because AI buildout is consuming the same transformers, switchgear, and construction capacity that any data center needs — and that cost pressure has to land somewhere eventually, even on a facility that has never hosted a GPU. If your provider's rate hasn't moved yet, that is a timing gap, not evidence the pressure doesn't apply to you.
The efficiency metric that stopped being the main story
Power Usage Effectiveness — the ratio of total facility power to IT equipment power — has been the industry's standard efficiency benchmark for two decades. Global average PUE sat at 1.54 in the Uptime Institute's 2025 survey, essentially flat since around 2020, while leading hyperscale operators have pushed fleet-wide PUE down to 1.09 (Uptime Institute; Google Data Centers, 2025).
PUE still matters for operating cost and compliance — Germany's Energy Efficiency Act requires new data centers commissioned from July 2026 to hit 1.2 or below. But it answers a different question than the one buyers are now asking. A facility with a best-in-class 1.1 PUE cannot open at all without a grid interconnection, and no amount of cooling efficiency shortens a four-year utility queue. Efficiency optimizes the power you can get; it doesn't get you the power in the first place.
The regulatory response, and why it won't fix 2026
On 18 June 2026, FERC issued show-cause orders under Section 206 of the Federal Power Act to all six FERC-jurisdictional regional grid operators, instructing each to justify or revise its large-load interconnection rules — a faster, more targeted move than a standard rulemaking, following an October 2025 DOE directive (White & Case; American Action Forum, June 2026). It may shorten future queues in the regions it touches.
It does not retroactively move anyone already in a queue. Projects with existing applications wait behind whatever framework their grid operator adopts, on whatever timeline that operator sets. For a deployment decision made this quarter, the FERC action is a signal about direction, not a lever that changes your timeline.
What to do with this if you're buying colocation or dedicated capacity now
Ask about grid interconnection status before floor space. A provider quoting available cabinets without confirming power headroom is quoting you half an answer. Confirmed interconnection, not listed vacancy, is the number that determines whether your deployment date is real — it's the first thing INXY verifies with any facility before recommending it to a client.
Lead time is now 18 to 24 months in primary markets, not 6 to 12. If your growth planning still assumes the shorter window, the gap will show up as a missed deployment date, not a budget overrun.
Secondary and tertiary markets carry real tradeoffs, not just lower prices. Power-advantaged regions outside the traditional hubs (Northern Virginia, Silicon Valley, Phoenix) often have shorter interconnection queues precisely because demand hasn't concentrated there yet — but latency, connectivity density, and carrier presence vary accordingly. The right tradeoff depends on whether your workload is latency-sensitive or throughput-oriented.
A workload that doesn't need dedicated colocation power at all sidesteps the queue entirely. Delivery-layer capacity — CDN, managed DNS, cloud compute sized to the actual workload — doesn't carry the multi-year interconnection risk that a dedicated cage does, because you're not the one waiting on the utility.
We broker colocation and dedicated capacity across multiple facilities and markets, which means INXY can tell you which shortlisted sites have confirmed power today rather than a queue position and a promise — a distinction that determines whether your deployment date is real. [Request an infrastructure audit → /book-a-demo] before you commit to a facility on the strength of its floor plan alone. To review capacity options directly: [dedicated servers, colocation and racks → /hosting-solutions].
FAQ
Why is data center power harder to get than GPUs in 2026? GPU supply eased as TSMC scaled advanced packaging capacity multiple times since 2024, easing the earlier allocation crunch. Grid capacity did not scale at the same pace, because interconnection queues, transformer manufacturing, and utility planning all operate on multi-year cycles that chip production improvements don't affect.
How long does it take to get a new data center connected to the grid in 2026? Typically 24 to 72 months depending on market and load size, with 4 to 7 year waits in the most constrained primary markets such as Northern Virginia, Phoenix, and Dallas. Secondary markets with less concentrated demand generally offer shorter queues.
Does the power shortage affect standard hosting, or only AI data centers? It affects both, because standard and AI colocation compete for the same grid interconnection queue and equipment supply chain. Cushman & Wakefield data shows power availability, not floor space, is now the binding constraint on new colocation supply in every major US market, and pricing for standard deployments has risen accordingly.
What is PUE and does it matter for site selection? Power Usage Effectiveness measures total facility power against IT equipment power, with lower numbers indicating less overhead. It remains relevant for operating cost and compliance, but it doesn't address whether a site can secure grid power at all — a highly efficient facility still can't open without an interconnection, so PUE answers a different question than power availability does.
Will the June 2026 FERC orders fix data center interconnection delays? The orders instruct regional grid operators to justify or revise their large-load interconnection rules, which may shorten future queues in the affected regions. They do not retroactively accelerate projects already waiting in a queue, so deployments planned for this year should not assume near-term relief from the ruling.
How far in advance should I secure colocation capacity now? Roughly 18 to 24 months in primary markets, up from the 6 to 12 months that was standard a few years ago. This shift reflects both record-low vacancy and the share of new capacity that is pre-committed before construction completes.

