The post explains how Netflix attests Apache Spark workloads running on Amazon EMR and exchanges cloud IAM roles for short-lived internal X.509 identities. It details dual-claim corroboration, role sharding, executor credential fan-out, renewal, and the trust trade-offs behind the design.
Netflix engineering blog
Danny Thomas introduces ja and its composable Java command-line tools for module-aware dependency resolution, formatting, symbol search, and documentation. The post explains Maven coordinate discovery, explicit runtime access authorization, dependency integrity hashes, and safer annotation processing for agent-friendly Java development.
Netflix describes MAPS, a production system that uses CLIP and MediaFM multimodal embeddings to personalize artwork and video previews while reducing cold-start problems. The post covers model consolidation, propensity-weighted offline evaluation, query-aware ranking, shared embedding infrastructure, and a linear-probe proxy for selecting embeddings efficiently.
Netflix compares its legacy external Flink autoscaler with the Apache Flink Autoscaler, explaining how operator-level throughput modeling enables finer-grained scaling for stateful DAGs. The post details the Temporal-based architecture, metric collection fixes, safety checks, and lessons from operating autoscaling across more than 30,000 jobs.
Netflix describes how it models device capabilities at scale to support analytics and feature management across a diverse device ecosystem. The post explains cumulative and histogram tables used to capture device states and aggregate capability distributions for use cases like 4K Ultra HD, spatial audio, and cloud gaming.
Netflix introduces GenRec, an LLM-backed recommendation ranker that adapts an internal foundation model for large-scale personalization. The post details its two-phase training process, context engineering approach, reward-weighted objectives, serving optimizations, and online gains versus a mature production ranker.
Netflix describes its in-house LLM serving stack built on Triton and vLLM, integrated into its broader JVM-based production serving system. The post explains key architectural decisions around engine selection, packaging, HTTP API compatibility, deployment strategies, observability, and constrained decoding at scale.
A deep dive into the engineering challenges of building a real-time service dependency map at Netflix scale, covering the architecture, production bottlenecks, and tuning work required to make it reliable. The post explains streaming-first ingestion, multi-stage aggregation, backpressure, and lessons learned from Kafka lag, hot nodes, memory pressure, and reactive streams complexity.
Netflix shares two research explorations for more controllable AI video editing: Vera, a layered video diffusion model, and VOID, a physically-plausible video inpainting model. The post explains how both approaches aim to preserve source footage while making specific edits or deletions with greater precision and realism.
Netflix describes migrating its managed batch compute system from custom queuing and scheduling logic to Kubernetes-native Kueue. The post covers the original tenant hierarchy, why Kueue was chosen, the migration approach, and the resulting improvements in fairness, preemption, and utilization.
This post presents a hierarchical notification system that separates long-term pacing decisions from real-time message selection. Netflix describes how the “slow” policy sets personalized messaging frequency while the “fast” policy picks the best message at each opportunity, improving both member experience and engagement.
The post describes how Netflix uses production schedule data and predictive modeling to estimate delivery risk for content launches. By predicting time-to-delivery for media assets, the team can fill ETA gaps, improve accuracy versus manual schedules, and reduce launch misses.
Data Projects introduces a durable project-level abstraction for managing access and identity across Netflix’s data platform. It groups related assets like tables and workflows under a single umbrella, replacing fragile per-asset ACLs and human-tied workflow identities with a project-owned Netflix application identity.
Netflix shares an agentic workflow for observational causal inference that uses an actor-critic structure, rigorous diagnostics, and human review to improve analysis quality. The post also details an open-source OCI agent, case studies on trimming for overlap, and evaluations showing stronger results than one-shot prompting.
This post explains how Netflix reduced the impact of wide Cassandra partitions for TimeSeries workloads by introducing automated re-partitioning. It covers table-level time slice re-partitioning, per-ID dynamic partition splitting, checkpointing, validation, and read-path diversion using Bloom filters.
Netflix’s Graph Abstraction is introduced as a high-throughput property-graph platform built on top of existing KV, TimeSeries, and EVCache abstractions. The post explains its schema-driven design, storage layout for nodes and edges, caching strategies, consistency model, traversal API, and production-scale performance.
Netflix describes how it built Service Topology, a real-time map of service dependencies designed to help engineers troubleshoot incidents, understand blast radius, and navigate distributed systems faster. The post explains how the system combines three complementary data sources—eBPF network flows, IPC metrics, and distributed traces—into a unified, queryable graph with near real-time updates and gRPC access.
At Netflix, the JVM Ecosystem team uses Nebula Gradle plugins to standardize builds across a large Java polyrepo. This post explains how Nebula ArchRules scales ArchUnit beyond a single repository, enabling shared rule libraries, automated reporting, and organization-wide enforcement of API lifecycle and other code quality policies.
Netflix describes how it built a Metadata Service and Model Lifecycle Graph to unify fragmented machine learning metadata across teams and domains. The system ingests events from multiple source systems, normalizes them into a shared entity model, and enriches relationships to support discovery, lineage, impact analysis, and graph-based exploration in the AIP Portal.
This post explains how Netflix’s centralized ML model serving platform routes traffic to the correct model instance at scale while preserving a simple abstraction for client services and researchers. It introduces Switchboard, a routing layer that supports context-aware model selection, experimentation, and lifecycle management, and then describes how Netflix evolved the design into Lightbulb to reduce latency and eliminate the routing service from the critical request path. The article highlights the separation of routing metadata from model-serving configuration and the use of Envoy to handle cluster-level request routing more efficiently.