Breakpoint

Pinterest engineering blog

Beyond Two Towers: Launching the 3-Tower Engagement Co-Train Model (Part 2)
Pinterest engineers describe a three-tower co-training architecture that combines cached two-tower scoring with cross-attention refinement for large-scale ad ranking. The post details multi-task modeling, BF16 and serving optimizations, two-stage candidate selection, and online gains achieved within tight latency and cost constraints.
Evolving Pinterest’s Embedding Retrieval Platform
Pinterest explains how it evolved the Manas embedding retrieval platform for tens of billions of vectors. The post details quantization trade-offs, SSD-based SPANN serving, SIMD optimizations, and multi-embedding late-interaction retrieval, including measured impacts on recall, latency, memory, CPU usage, and cost.
Building Pinterest’s VLM Serving Stack on NVIDIA Dynamo
Pinterest details how it built a production vision-language model serving platform on NVIDIA Dynamo, vLLM, Blackwell GPUs, and Kubernetes. The post covers disaggregated serving, multimodal KV-aware routing, cache offloading, projection embeddings, realistic agentic benchmarking, and substantial latency improvements.
Becoming an AI Team
Pinterest’s John Grass explains how AI changes engineering-team operating models, expanding capacity while making prioritization, strategic thinking, and problem definition more important. He outlines evolving manager responsibilities, practical uses of AI agents in infrastructure, and leadership practices for experimentation, learning, and psychological safety.
Scaling Conditional Learned Retrieval for Pinterest Home Feed
Pinterest engineers explain how Conditional Learned Retrieval scales personalized home-feed candidate generation beyond two-tower retrieval. The post covers condition-aware sequence modeling, semantic IDs, unified routing, training optimizations, GPU serving, and an 85% reduction in p90 latency.
Securing Infrastructure at Scale: Introducing Pinterest’s Resource Provisioner Pipeline (RPP)
Managing infrastructure as code at Pinterest’s scale requires strong security guardrails, and this post introduces Resource Provisioner Pipeline (RPP) as a centralized Terraform execution engine. It explains how RPP uses dual controls, workspace-to-role mapping, OIDC-based role chaining, and GitHub Actions to safely plan and apply AWS infrastructure changes across many repositories.
Achieving Near-Linear Training Scalability for Pinterest’s Foundation Models
Pinterest describes how it scaled embedding-heavy foundation model training from poor multi-node performance to near-linear throughput across 2, 4, and 8 nodes. The post walks through a sequence of optimizations including quantized communications, balanced sharding, bandwidth-aware embedding changes, and 2D parallel topology redesign, plus supporting infrastructure upgrades.
Making User-Sequence Data More Cost-Efficient, Faster, and Easier to Use
Pinterest redesigned its user-sequence platform to make enriched event sequences cheaper to run, faster to extend, and easier to debug. The post covers a one-definition, many-runtimes architecture, configuration-as-code, a shared execution engine, lambda-style batch and streaming paths, and columnar time-partitioned storage.
Enhancing Ad Relevance: Integrating Real-Time Context into Sequential Recommender Models
Pinterest Engineering describes a contextual sequential two-tower recommender designed to improve ad relevance by combining historical user behavior with real-time page context. The post explains the model architecture, synthetic context-based training, and hybrid offline/online inference flow, and reports meaningful gains in candidate survival, relevance, recall, and ROAS.
Optimizing ML Workload Network Efficiency (Part I): Feature Trimmer
Guangtong Bai, Shantam Shorewala, Chi Zhang, Neha Upadhyay, and Haoyang Li describe how Pinterest reduced root-to-leaf ML serving network bottlenecks by first enabling compression and then introducing Feature Trimmer to send only the features actually used by each model. The post covers model-signature-based allowlisting, deployment synchronization, versioned fallback, and the resulting latency, bandwidth, and infrastructure cost savings.
From Clicks to Conversions: Architecting Shopping Conversion Candidate Generation at Pinterest
Pinterest Engineering describes how it built a dedicated shopping conversion candidate generation model to better optimize for sparse, noisy offsite conversion signals instead of relying primarily on engagement-based retrieval. The post details its data design, feature engineering, and architecture changes, including a parallel DCN v2 + MLP cross-layer design and a shift to a unified multi-task setup with advertiser-level loss. These changes delivered measurable gains in recall, conversion volume, CTR, and RoAS for Pinterest’s shopping ads.
Finding zombies in our systems: A real-world story of CPU bottlenecks
Pinterest describes a three-month investigation into intermittent Ray training job crashes caused by AWS ENA network driver resets. The root cause turned out to be CPU starvation from a buildup of zombie memory cgroups created by a crashlooping ECS agent on the base image, which was fixed by disabling the agent and rebooting hosts.
Scaling Recommendation Systems with Request-Level Deduplication
Pinterest explains how request-level deduplication reduces redundancy across recommendation storage, training, and serving by storing request-level data once instead of once per item. The post covers storage compression, training regressions from sorted data, and the use of SyncBatchNorm to recover model quality while preserving efficiency gains.
Performance for Everyone
Pinterest built a unified, low-effort measurement system for user-perceived latency by embedding Visually Complete logic into a base UI class so surfaces built on top automatically report perceived latency. On Android this inspects the view tree and specialized media view interfaces to determine when content is visually complete; the approach was later extended to iOS and web to provide broad performance visibility and protection.
Evolution of Multi-Objective Optimization at Pinterest Home feed
This post describes the evolution of Pinterest’s feed-level multi-objective optimization from a DPP-based diversification to a PyTorch-served Sliding Spectrum Decomposition (SSD), plus additions like soft-spacing and richer embedding signals. The changes improved diversity, serving latency, and the ability to incorporate new quality and semantic signals (e.g., PinCLIP, Semantic IDs) while migrating more logic into model servers for easier iteration.
Building an MCP Ecosystem at Pinterest
Pinterest built an internal ecosystem around the Model Context Protocol (MCP) consisting of cloud-hosted, domain-specific MCP servers and a central registry. The post explains architecture, a unified deployment pipeline, integrations into internal surfaces (chat, IDEs), and governance including JWT and mesh-based auth plus human-in-the-loop controls. It also covers observability, security review processes, and measured impact of the system.