Pinterest explains how it restores reliable partition-completeness signals in a streaming database ingestion system. The design uses event-time statistics, mergeable t-digest sketches, Iceberg snapshot metadata, and non-regressing watermarks implemented through extensible Flink sink hooks.
Pinterest engineering blog
Pinterest engineers describe a three-tower co-training architecture that combines cached two-tower scoring with cross-attention refinement for large-scale ad ranking. The post details multi-task modeling, BF16 and serving optimizations, two-stage candidate selection, and online gains achieved within tight latency and cost constraints.
Pinterest explains how it evolved the Manas embedding retrieval platform for tens of billions of vectors. The post details quantization trade-offs, SSD-based SPANN serving, SIMD optimizations, and multi-embedding late-interaction retrieval, including measured impacts on recall, latency, memory, CPU usage, and cost.
Pinterest details how it built a production vision-language model serving platform on NVIDIA Dynamo, vLLM, Blackwell GPUs, and Kubernetes. The post covers disaggregated serving, multimodal KV-aware routing, cache offloading, projection embeddings, realistic agentic benchmarking, and substantial latency improvements.
Pinterest’s John Grass explains how AI changes engineering-team operating models, expanding capacity while making prioritization, strategic thinking, and problem definition more important. He outlines evolving manager responsibilities, practical uses of AI agents in infrastructure, and leadership practices for experimentation, learning, and psychological safety.
Pinterest engineers explain how Conditional Learned Retrieval scales personalized home-feed candidate generation beyond two-tower retrieval. The post covers condition-aware sequence modeling, semantic IDs, unified routing, training optimizations, GPU serving, and an 85% reduction in p90 latency.
This post introduces Pinner Progression, a recommendation-system initiative focused on retention rather than only short-term engagement. It explains User Interest Clusters (UICs), a stateful user-interest representation built from clustered engaged Pins, and how UICs are integrated across retrieval, ranking, and blending to improve weekly active user growth and session depth.
Managing infrastructure as code at Pinterest’s scale requires strong security guardrails, and this post introduces Resource Provisioner Pipeline (RPP) as a centralized Terraform execution engine. It explains how RPP uses dual controls, workspace-to-role mapping, OIDC-based role chaining, and GitHub Actions to safely plan and apply AWS infrastructure changes across many repositories.
Pinterest describes how it scaled embedding-heavy foundation model training from poor multi-node performance to near-linear throughput across 2, 4, and 8 nodes. The post walks through a sequence of optimizations including quantized communications, balanced sharding, bandwidth-aware embedding changes, and 2D parallel topology redesign, plus supporting infrastructure upgrades.
Pinterest redesigned its user-sequence platform to make enriched event sequences cheaper to run, faster to extend, and easier to debug. The post covers a one-definition, many-runtimes architecture, configuration-as-code, a shared execution engine, lambda-style batch and streaming paths, and columnar time-partitioned storage.
The post describes a testing harness for measuring how reliably AI agents load and invoke a custom engineering skill. It then shares several techniques for improving skill invocation rates, including better frontmatter, stronger wording, and AGENTS.md guidance.
Pinterest Engineering describes a contextual sequential two-tower recommender designed to improve ad relevance by combining historical user behavior with real-time page context. The post explains the model architecture, synthetic context-based training, and hybrid offline/online inference flow, and reports meaningful gains in candidate survival, relevance, recall, and ROAS.
Guangtong Bai, Shantam Shorewala, Chi Zhang, Neha Upadhyay, and Haoyang Li describe how Pinterest reduced root-to-leaf ML serving network bottlenecks by first enabling compression and then introducing Feature Trimmer to send only the features actually used by each model. The post covers model-signature-based allowlisting, deployment synchronization, versioned fallback, and the resulting latency, bandwidth, and infrastructure cost savings.
Pinterest Engineering describes how it built a dedicated shopping conversion candidate generation model to better optimize for sparse, noisy offsite conversion signals instead of relying primarily on engagement-based retrieval. The post details its data design, feature engineering, and architecture changes, including a parallel DCN v2 + MLP cross-layer design and a shift to a unified multi-task setup with advertiser-level loss. These changes delivered measurable gains in recall, conversion volume, CTR, and RoAS for Pinterest’s shopping ads.
This post explains Pinterest’s MIQPS algorithm, which learns which URL query parameters are important for content identity and which are safe to strip. It describes how Pinterest groups URLs by parameter patterns, tests parameters against rendered content IDs, and uses anomaly detection to prevent regressions in URL normalization.
Pinterest describes a three-month investigation into intermittent Ray training job crashes caused by AWS ENA network driver resets. The root cause turned out to be CPU starvation from a buildup of zombie memory cgroups created by a crashlooping ECS agent on the base image, which was fixed by disabling the agent and rebooting hosts.
Pinterest explains how request-level deduplication reduces redundancy across recommendation storage, training, and serving by storing request-level data once instead of once per item. The post covers storage compression, training regressions from sorted data, and the use of SyncBatchNorm to recover model quality while preserving efficiency gains.
Pinterest built a unified, low-effort measurement system for user-perceived latency by embedding Visually Complete logic into a base UI class so surfaces built on top automatically report perceived latency. On Android this inspects the view tree and specialized media view interfaces to determine when content is visually complete; the approach was later extended to iOS and web to provide broad performance visibility and protection.
This post describes the evolution of Pinterest’s feed-level multi-objective optimization from a DPP-based diversification to a PyTorch-served Sliding Spectrum Decomposition (SSD), plus additions like soft-spacing and richer embedding signals. The changes improved diversity, serving latency, and the ability to incorporate new quality and semantic signals (e.g., PinCLIP, Semantic IDs) while migrating more logic into model servers for easier iteration.
Pinterest built an internal ecosystem around the Model Context Protocol (MCP) consisting of cloud-hosted, domain-specific MCP servers and a central registry. The post explains architecture, a unified deployment pipeline, integrations into internal surfaces (chat, IDEs), and governance including JWT and mesh-based auth plus human-in-the-loop controls. It also covers observability, security review processes, and measured impact of the system.