Breakpoint

Instacart engineering blog

Agentic Machine Learning Modeling at Instacart
Instacart examines how AI agents can automate and improve machine-learning modeling through adaptive hypothesis generation, experimentation, and evaluation. The post presents delivery-time and catalog-attribute case studies with measured gains, while detailing safeguards for data leakage, evaluation errors, scale mismatch, multiple testing, and human oversight.
Blueberry: Force Multiplier For The On-Call Engineer
Instacart explains how it built Blueberry, a Slack-native reasoning harness that helps on-call engineers triage alerts, test theories, and preserve operational knowledge. The post covers its durable queue, composable MCP tool surfaces, parallel investigations, team-specific specialization, evidence-grounded workflows, and production results.
Variance Reduction Below the Randomization Grain
Instacart presents an order-level CUPED method for reducing variance in marketplace experiments that require coarse, region-level randomization. By aggregating pre-treatment machine-learning predictions to the experimental grain, the approach reduced variance by 18–40% across 10 experiments and shortened runtimes by about one-third while preserving unbiased estimates.
Leveraging PyFixest for High-Cardinality Marketplace Modeling at Instacart
Instacart explains how high-cardinality fixed effects make conventional regression computationally expensive in marketplace experiments. The post derives the Frisch–Waugh–Lovell and alternating-projections approaches, then benchmarks PyFixest against Statsmodels and scikit-learn, demonstrating major gains in speed, memory efficiency, and estimator precision.
From Scoring to Spelling: Rebuilding Ads Retrieval at Instacart
Instacart details its migration from product-ID scoring to generative retrieval using Semantic IDs, RQ-VAE codebooks, autoregressive decoding, and beam search. The new GPU serving stack, built with TensorRT-LLM, Triton, and Go, increased candidate volume while reducing latency and improved click-through, add-to-cart, and catalog diversity metrics.
Semantic IDs: Product Understanding at Scale
Instacart explains how it uses residual vector quantization and contrastive learning to turn product embeddings into hierarchical semantic IDs. The post covers taxonomy-guided sampling, precision-versus-discovery embedding strategies, intrinsic evaluation, failure modes, and applications in recommendations, generative retrieval, and catalog quality.
How AI Changes the Role of Applied Scientists
Instacart’s Economics team analyzes three years of GitHub activity to show how AI tools reshaped applied scientists’ productivity, task portfolios, and technical roles. The post examines newly feasible work, coordination costs, and why AI-native, machine-interactable platforms may outperform rigid dashboards for some workflows.
Scaling Personalized Marketing for Multi-Tenant Commerce Platforms
Instacart explains how it built a multi-tenant marketing platform for hundreds of retailers, combining isolated provider workspaces, event-driven batching, asynchronous workers, and automated template deployment. The architecture addresses data isolation, API limits, deliverability, observability, and reliable personalized messaging at high volume.
Empowering Carrot Ads with Domain Adaptive Learning
Instacart explains how Carrot Ads uses domain adaptive learning to solve the cold-start problem when onboarding new retail partners. The approach transfers embeddings and behavioral signals from the data-rich Instacart Marketplace, aligns features across domains, fine-tunes partner-specific layers, and trims features to balance prediction quality with auction latency.