Breakpoint

Why rolling dashboards miss the cache at Netflix scale

A dashboard showing "the last 3 hours" refreshes every 10 seconds. Almost all of

netflix··PT2M3S

video loads only when you press play

A dashboard showing "the last 3 hours" refreshes every 10 seconds. Almost all of

A dashboard showing "the last 3 hours" refreshes every 10 seconds. Almost all of

  • Moving time windows miss whole-result caches even when almost every minute is unchanged.
  • Netflix cached granularity-aligned time buckets and fetched only the uncovered tail.
  • Older buckets receive longer TTLs because late-arriving data becomes less likely with age.

Your dashboard just asked for 3 hours it already had. It does that every 10 seconds, and so does everyone else's, so 30 people on one dashboard means nearly two hundred queries a second, all asking for almost exactly the same 3 hours. You'd think a cache would catch that, and Netflix had one, but the window keeps moving. The last 3 hours at ten o'clock and the last 3 hours a minute later aren't the same question, so nothing matches and all 3 hours get computed again. And it won't touch anything still filling up, because it would rather be right than fast. But look at what actually changed between those two questions. Almost nothing. The data from 2 hours ago is settled, it isn't going to move again, and only the last few minutes are genuinely new. So they stopped caching the answer, and started caching the minutes. 3 hours, cut into one minute pieces, is a hundred and eighty small results, each one stored on its own. When the window shifts, you keep the ones you've got and ask only for the sliver on the end. Two hours fifty from memory, 10 minutes from the database. What makes that possible is the key. They hash what's being asked, the filters, the maths, the grouping, and leave the time out of it, so the same question over any window lands in the same place. Which leaves the one real question: how long do you keep a piece? And the answer is the good bit. The longer it's been sitting there, the longer you keep it. Something from 30 seconds ago might still change, because events turn up late, but something from 30 minutes ago is finished. So the newest pieces live 5 seconds, and every extra minute of age doubles that, all the way up to an hour. Now eighty four percent of the data comes back without touching the database, a third fewer queries get through, and the slow ones came down by two thirds. And the part that matters most: it barely notices how many people are watching. One viewer or a hundred, it does about the same work. The database never got faster. It just stopped being asked for what it had already answered.

This explainer is based on Stop Answering the Same Question Twice: Interval-Aware Caching for Druid at Netflix Scale by Netflix ↗. The original reporting and technical work belong to its publisher.