Breakpoint

Lyft engineering blog

Rerouting the Stream: How Lyft Moved to the Apache Flink Operator
Lyft explains its incremental migration from a custom Flink Kubernetes operator to Apache’s operator across hundreds of streaming jobs. The post details boundary translation, stateful deployment strategies, autoscaling and autotuning trade-offs, BlueGreen fixes, Beam memory isolation, and Karpenter-driven infrastructure changes.
Metric Semantic Layer: How Lyft Governs and Scales Key Data Definitions
Lyft describes how it built an internal Metric Semantic Layer to centralize metric definitions, SQL logic, and ownership metadata in one authoritative package. The post covers governance, access methods, integrations with Amundsen and AI tooling, and how the system supports consistent metric usage across teams.
How We Built a Smarter Pickup Experience for Gated Communities
Lyft’s Mapping team redesigned the pickup flow for gated communities by adding gate-aware map data, smarter pickup spot suggestions, routing through the correct entrance, and in-app gate instruction sharing. The result is a smoother experience that reduces cancellations, wait times, and awkward back-and-forth between riders and drivers.
Predicting Rider Conversion in Sparse Data Environments with Bayesian Trees
Lyft describes a hierarchical Bayesian tree framework for predicting rider conversion in high-cardinality, sparse data contexts. The approach trains parametric models at tree nodes and uses Gaussian priors (L2 regularization toward parent parameters) to smooth predictions, enforce monotonicity, and enable robust, low-latency real-time serving.
Beyond A/B Testing: Using Surrogacy and Region-Splits to Measure Long-Term Effects in Marketplaces
This article presents a methodology to estimate market-mediated long-term effects of policy changes in a two-sided marketplace. It uses a two-step surrogacy-style approach: (1) residualized models map policy changes to shifts in negative user experiences, and (2) doubly-robust causal estimation (AIPW) maps those experiences to future outcomes, with end-to-end validation via region-split experiments and a forward-selection design for treated/control regions.
Scaling Localization with AI at Lyft
This post describes how Lyft re-architected its translation pipeline to combine LLM-generated drafts with human linguist review to scale localization. It explains the Drafter/Evaluator iterative pattern, context injection, deterministic guardrails for placeholders and formatting, prompt versioning, and multi-model experimentation to reduce translation latency from days to minutes while maintaining quality.
Trusting the Untestable: Validation and Diagnostics for the Doubly Robust Models
This post describes Lyft’s use of doubly robust estimators (AIPW) for causal inference when randomization is not possible, and emphasizes rigorous validation through confounder management and diagnostic scorecards. It covers practical issues like confounder set requirements, downsampling corrections (propensity score conversion and outcome reweighting), and platform-level safeguards to build trust in non-randomized measurement.
Lyft’s Feature Store: Architecture, Optimization, and Evolution
This article presents the architecture and operational evolution of Lyft’s Feature Store, describing batch, online, and streaming feature paths and how they support large-scale ML workflows. It covers ingestion via Spark/Hive, Airflow-generated DAGs, the online dsfeatures layer backed by DynamoDB with a ValKey cache and OpenSearch for embeddings, SDKs in Go and Python, and streaming pipelines using Flink and Kafka for low-latency feature serving.
From Python3.8 to Python3.10: Our Journey Through a Memory Leak
An engineering post about diagnosing a real-world memory leak encountered during a Python upgrade from 3.8 to 3.10. The team describes how they traced memory growth with tracemalloc-based tooling, discovered interactions between gevent/greenlet and urllib3/botocore/pynamodb, and resolved the issue by adjusting urllib3/gevent versions and deployment settings.
LyftLearn Evolution: Rethinking ML Platform Architecture
Lyft rethought LyftLearn’s architecture to reduce operational complexity: retaining Kubernetes for low-latency online model serving while migrating offline compute (training, batch jobs, notebooks) to managed AWS SageMaker. The hybrid approach aims to simplify orchestration, lower TCO, and accelerate feature development by leveraging managed services for elastic offline workloads while keeping control for online serving.
My Starter Project on the Lyft Rider Data Science Team
A new-hire writeup about a starter project using Lyft's Rider Experience Score (RES) to estimate long-term effects of rider experiences on retention. The post explains the causal inference approach (AIPW with cross-fitting), model choices (XGBoost/LightGBM), confounder selection challenges, and how the author implemented and scaled RES estimates for multiple experience signals.
Migrating Lyft’s Android Codebase to Kotlin
A post describing Lyft's multi-year effort to migrate its Android applications from Java to Kotlin. It covers motivations (conciseness, Compose adoption, coroutines), migration strategy, tooling and compiler choices, and lessons learned while converting a large, long-lived codebase.
Intern Experience at Lyft
Two former Lyft interns (Morteza Taiebat and Han Gong) describe their internship projects, the teams they worked with, and how those experiences led them to return as full-time data scientists. The post highlights causal modeling (difference-in-differences), driver productivity metrics, referral program analysis, and practical lessons for future interns.
Solving Dispatch in a Ridesharing Problem Space
This post frames rideshare dispatch as a dynamic bipartite matching problem and explains the mathematical and practical challenges of producing high-quality real-time matchings at scale. It covers graph formulations, LP relaxations, batching trade-offs, myopic vs. long-term decisions, and approaches like forecasting and rebalancing to improve dispatch outcomes.
How science inspires our ETA models
The post examines travel-time uncertainty and how micro- and macro-scale traffic patterns inform statistical ETA models. It describes empirical observations (e.g., why longer trips often have more predictable ETAs), cumulative averaging over route segments, and how these insights translate into model design for more accurate travel-time predictions.