Breakpoint

Meta engineering blog

Bringing Private Processing to Meta AI Glasses
Meta explains how Private Processing extends confidential computing to AI glasses, using trusted execution environments, remote attestation, anonymous routing, and encrypted in-boundary storage. The design addresses stateful AI, non-targetability, operational observability, and independent verification while keeping personal context inaccessible to Meta.
ZGateway: Learnings from Putting a Proxy in Front of ZippyDB
Meta explains why it placed ZGateway in front of ZippyDB to replace a fragile client-to-database connection mesh with a bounded, centrally managed proxy tier. The post details connection fan-in reduction, cross-client batching and coalescing, admission control, caching, load balancing, and cross-region resilience.
An Organizational Second Brain: Building an AI That Learns From Experts
Meta describes an AI “second brain” that combines structured, auditable knowledge files with composable reasoning recipes and a human-guided self-improvement loop. Expert feedback becomes minimal, regression-tested edits without model retraining, preserving institutional expertise while reducing assessment time and preventing regressions.
MetaRoCE: A New RDMA Transport Built for AI-Scale Ethernet
Meta introduces MetaRoCE, an open RDMA transport designed for AI-scale Ethernet. The protocol moves congestion and ordering intelligence to NIC endpoints, enabling out-of-order delivery, native multipathing, loss tolerance, and graceful recovery, with benchmark results showing improved throughput and resilience over RoCEv2.
GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model
Meta’s Generative Ads Recommendation Model (GEM), the foundation model behind ads recommendations across Instagram and Facebook, now trains at LLM scale on several thousand of the latest-generation GPUs. The post explains how Meta doubled end-to-end training efficiency to 20–25% Model FLOPs Utilization while scaling training FLOPs 4x.
Meta’s AI Storage Blueprint at Scale
Over the past several years, model capabilities and training dataset sizes have grown rapidly, and Meta frames storage as a critical part of keeping AI innovation fast and cost-effective. The post appears to discuss the storage architecture needed to support frontier AI workloads at scale.
10 Years of Meta’s Commitment to Python
Meta highlights its 10th consecutive year sponsoring the Python Software Foundation and its long-term support for the Python ecosystem. The post emphasizes Python’s importance across Meta’s engineering stack and its role in the company’s open-source efforts.
How Meta Engineered Ultra-Narrow Batteries for AI Glasses
Smart glasses like Ray-Ban Meta and Oakley Meta Vanguards need batteries that can power cameras, speakers, AI workloads, and even a display while fitting into the temple arms. This post explores how Meta engineered ultra-narrow batteries to meet those constraints.
Adopting AV1 for Real-Time Communication (RTC) at Scale
Meta shares the multi-year effort behind adopting AV1 for real-time communication, including codec selection, device eligibility, rate control, and error resilience. The post covers the technical and operational challenges of deploying AV1 at scale and the improvements made to call quality.
Lights Out, Systems On: Validating Instant Power Loss Readiness
We’re introducing Instantaneous PowerLoss Storm, a new testing paradigm within Meta’s infrastructure for handling and mitigating instant or zero-notice power loss in our data centers. The post explains how Meta built readiness to tolerate instant failures using defense-in-depth strategies, the tradeoffs involved, and how the team validated that readiness.
Reel Friends: Building Social Discovery that Scales to Billions
On its face the new Friend Bubbles feature looks simple enough: it highlights Reels your friends have watched and reacted to. The article frames the feature as deceptively straightforward and emphasizes the deep engineering work required to make social discovery scale.