Breakpoint

Netflix engineering blog

Leave the Class Path in the Rearview Mirror
Danny Thomas introduces ja and its composable Java command-line tools for module-aware dependency resolution, formatting, symbol search, and documentation. The post explains Maven coordinate discovery, explicit runtime access authorization, dependency integrity hashes, and safer annotation processing for agent-friendly Java development.
MAPS: Netflix’s Multimodal Asset Personalization at Scale
Netflix describes MAPS, a production system that uses CLIP and MediaFM multimodal embeddings to personalize artwork and video previews while reducing cold-start problems. The post covers model consolidation, propensity-weighted offline evaluation, query-aware ranking, shared embedding infrastructure, and a linear-probe proxy for selecting embeddings efficiently.
A Tale of Two Flink Autoscalers
Netflix compares its legacy external Flink autoscaler with the Apache Flink Autoscaler, explaining how operator-level throughput modeling enables finer-grained scaling for stateful DAGs. The post details the Temporal-based architecture, metric collection fixes, safety checks, and lessons from operating autoscaling across more than 30,000 jobs.
Modeling Device Capabilities for Analytics
Netflix describes how it models device capabilities at scale to support analytics and feature management across a diverse device ecosystem. The post explains cumulative and histogram tables used to capture device states and aggregate capability distributions for use cases like 4K Ultra HD, spatial audio, and cloud gaming.
GenRec: Towards LLM-Native Recommendation at Netflix
Netflix introduces GenRec, an LLM-backed recommendation ranker that adapts an internal foundation model for large-scale personalization. The post details its two-phase training process, context engineering approach, reward-weighted objectives, serving optimizations, and online gains versus a mature production ranker.
In-House LLM Serving at Netflix
Netflix describes its in-house LLM serving stack built on Triton and vLLM, integrated into its broader JVM-based production serving system. The post explains key architectural decisions around engine selection, packaging, HTTP API compatibility, deployment strategies, observability, and constrained decoding at scale.
Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned
A deep dive into the engineering challenges of building a real-time service dependency map at Netflix scale, covering the architecture, production bottlenecks, and tuning work required to make it reliable. The post explains streaming-first ingestion, multi-stage aggregation, backpressure, and lessons learned from Kafka lag, hot nodes, memory pressure, and reactive streams complexity.
How Netflix Simplified Batch Compute with Kueue
Netflix describes migrating its managed batch compute system from custom queuing and scheduling logic to Kubernetes-native Kueue. The post covers the original tenant hierarchy, why Kueue was chosen, the migration approach, and the resulting improvements in fairness, preemption, and utilization.
Thinking Fast & Slow for a Personalized Notification System
This post presents a hierarchical notification system that separates long-term pacing decisions from real-time message selection. Netflix describes how the “slow” policy sets personalized messaging frequency while the “fast” policy picks the best message at each opportunity, improving both member experience and engagement.
Data Projects: Managing Data Assets at Netflix Scale
Data Projects introduces a durable project-level abstraction for managing access and identity across Netflix’s data platform. It groups related assets like tables and workflows under a single umbrella, replacing fragile per-asset ACLs and human-tied workflow identities with a project-owned Netflix application identity.
A Human-Augmenting Agentic Workflow for Causal Inference
Netflix shares an agentic workflow for observational causal inference that uses an actor-critic structure, rigorous diagnostics, and human review to improve analysis quality. The post also details an open-source OCI agent, case studies on trimming for overlap, and evaluations showing stronger results than one-shot prompting.
Dynamic Repartitioning for Time Series Workloads
This post explains how Netflix reduced the impact of wide Cassandra partitions for TimeSeries workloads by introducing automated re-partitioning. It covers table-level time slice re-partitioning, per-ID dynamic partition splitting, checkpointing, validation, and read-path diversion using Bloom filters.
High-Throughput Graph Abstraction at Netflix: Part I
Netflix’s Graph Abstraction is introduced as a high-throughput property-graph platform built on top of existing KV, TimeSeries, and EVCache abstractions. The post explains its schema-driven design, storage layout for nodes and edges, caching strategies, consistency model, traversal API, and production-scale performance.
From Silos to Service Topology: Why Netflix Built a Real-Time Service Map
Netflix describes how it built Service Topology, a real-time map of service dependencies designed to help engineers troubleshoot incidents, understand blast radius, and navigate distributed systems faster. The post explains how the system combines three complementary data sources—eBPF network flows, IPC metrics, and distributed traces—into a unified, queryable graph with near real-time updates and gRPC access.
Scaling ArchUnit with Nebula ArchRules
At Netflix, the JVM Ecosystem team uses Nebula Gradle plugins to standardize builds across a large Java polyrepo. This post explains how Nebula ArchRules scales ArchUnit beyond a single repository, enabling shared rule libraries, automated reporting, and organization-wide enforcement of API lifecycle and other code quality policies.
Democratizing Machine Learning at Netflix: Building the Model Lifecycle Graph
Netflix describes how it built a Metadata Service and Model Lifecycle Graph to unify fragmented machine learning metadata across teams and domains. The system ingests events from multiple source systems, normalizes them into a shared entity model, and enriches relationships to support discovery, lineage, impact analysis, and graph-based exploration in the AIP Portal.
State of Routing in Model Serving
This post explains how Netflix’s centralized ML model serving platform routes traffic to the correct model instance at scale while preserving a simple abstraction for client services and researchers. It introduces Switchboard, a routing layer that supports context-aware model selection, experimentation, and lifecycle management, and then describes how Netflix evolved the design into Lightbulb to reduce latency and eliminate the routing service from the critical request path. The article highlights the separation of routing metadata from model-serving configuration and the use of Envoy to handle cluster-level request routing more efficiently.