Breakpoint

How Cloudflare removed 100 TB from its DNS cache

Cloudflare's public DNS resolver caches 250 billion answers at once, so every byte a cache entry wastes costs more than 250 GB of memory across the fleet.…

cloudflare··PT2M10S

video loads only when you press play

Cloudflare's public DNS resolver caches 250 billion answers at once, so every byte a cache entry wastes costs more than 250 GB of memory across the fleet.…

Cloudflare's public DNS resolver caches 250 billion answers at once, so every byte a cache entry wastes costs more than 250 GB of memory across the fleet. Five changes to how entries are laid out in memory cut each one from 953 bytes to 420, freed roughly 100 terabytes of RAM, and made the cache faster at the same time.

  • At 250 billion cached answers, one wasted byte becomes 250 GB across the fleet.
  • Compact, contiguous wire-format records avoided enum padding and scattered allocations.
  • The redesign cut an entry from 953 to 420 bytes while improving insertion and lookup performance.

Cloudflare just found one hundred terabytes of memory, and no they didn't just find sticks of RAM sitting around somewhere. It was hiding in the cache of their public DNS resolver, where every wasted byte costs two hundred and fifty gigabytes, because the cache holds two hundred and fifty billion answers at once. Let's break down how they managed to slim each answer down, byte-by-byte. Each one starts at nine hundred and fifty three bytes, mostly a bundle of lists: the records, plus some bookkeeping. Those lists came in Rust Vectors, each reserving spare room for items that might arrive later. But a cached answer is written once and never modified again, so that room was pure packaging. Swapping every Vector for a boxed slice, which can't grow, clawed back sixty-four bytes, plus all the spare slots. Then there's the name: every record spells out the domain it belongs to, which is usually the exact name the entry is filed under anyway, so the owner became an Option that's usually empty, with the name borrowed back at read time. The biggest waste was hiding in the records themselves. Each record's data is a Rust enum, and an enum is always the size of its largest variant, which here is a hundred and forty four bytes, when the most common record needs four. That's a hundred and twenty bytes of padding on most of their traffic. You could box the big variants off to the heap and let the slots shrink, but then the data scatters, and every read chases pointers across memory. So they stored the records as raw bytes instead, just as they travel on the wire, packed end to end in one buffer, and the cache got smaller and faster in the same move. All together, each answer fell from nine hundred and fifty three bytes to four hundred and twenty, handing back about half a kilobyte per entry. Multiply that by two hundred and fifty billion, and there's the hundred terabytes we mentioned earlier, the RAM of a hundred and thirty servers. Furthermore, they actually managed to increase cache throughput by forty three percent, while lookup latency dropped by nineteen percent. Ultimately, the memory was always in the racks, and hopefully we start seeing more tech companies make similar optimizations, especially during a memory shortage as big as this one.

This explainer is based on How we saved 100 terabytes of memory by optimizing 1.1.1.1's DNS cache by Cloudflare ↗. The original reporting and technical work belong to its publisher.