Breakpoint

Uber engineering blog

Designing MCP Gateway Uber's MCP Management Platform
Uber describes MCP Gateway, a centralized platform for discovering, governing, and executing more than 800 MCP servers and 5,000 tools. The architecture combines an AutoCrawler control plane, protocol translation across HTTP, gRPC, and TChannel, and built-in authorization, redaction, and observability for scaling agent integrations.
Taming the ML Firehose: Scaling Feature Consistency
Uber describes a feature-logging framework that makes inference-time features the canonical source for model training, eliminating training-serving skew. The post covers selective logging, compact feature aliases, Flink stream joins, state management, profiling, and Kafka scaling techniques that improved freshness, consistency, and throughput.
How Uber Protects Against Retry Storms
Uber Engineering explains how error ownership makes retries context-aware by distinguishing errors caused by a service from downstream failures it merely propagates. The production system limits cascading retry storms across deep service graphs, reducing the maximum storm radius from 25 levels to 3 and preventing up to 9.5 million unnecessary requests during an outage.
Large-Scale Automated Dependency Analysis Across Uber's Service Mesh
Uber engineers explain how they automatically map fail-close and fail-open dependencies across a large service mesh. The approach correlates inbound and outbound failures through middleware, aggregates every request with metrics, and applies probability thresholds to identify reliability-critical call paths.
From Signals to Context: Lessons from Lavaredo Ultra Trail
Uber explains how smart event correlation adds context to observability signals, helping teams reduce alert fatigue and turn noisy alerts into actionable insights. The post connects operational monitoring practices with lessons from ultra-distance racing.
Halving the Time: How Uber Eats Rebuilt Its Search Pipeline
Uber Eats engineers detail how they cut search latency in half by optimizing the full-stack pipeline, including retrieval, hydration, ranking, ad serving, serialization, and presentation. The post explains measured techniques such as request hedging, product-level grouping, parallel rendering, GPU serving, and agent-assisted performance optimization.
Running a Software Factory Efficiently at Uber Scale
Uber engineers explain how they scaled agentic coding usage 7x while stabilizing AI spend. The post details benchmark-driven model routing, prompt-cache tuning, CLI-based MCP access, code-mode batching, context graphs, and session-level cost analytics.
Running Cost-Efficient Export Workloads at Uber
Uber explains how selective export workloads over large historical datasets can trigger expensive full-table scans, keeping objects hot in Google Cloud Storage and increasing storage, retrieval, metadata, and egress costs. The post shows that combining Hudi column statistics with predicate-column sorting can dramatically reduce the scan surface area, improve pruning, and preserve colder storage tiers for long-term savings.
Uber’s Payments Platform
Uber describes the decade-long evolution of its Payments Platform, highlighting the core design principles that made it durable at scale: immutable money orders, zero-sum accounting, strongly consistent ledger balances, and loosely coupled microservices. The post also explains how the platform remained line-of-business and payment-instrument agnostic while expanding across new Uber products, payment methods, and high-throughput ledger use cases.
Scaling Exact COUNT(DISTINCT) for High-Cardinality Non-Rollup Metrics in Distributed Data Pipelines
Uber describes how it scaled exact COUNT(DISTINCT) for high-cardinality non-rollup metrics in distributed data pipelines without relying on approximate counting. The post explains a chunked aggregation buffer strategy that avoids the JVM’s 2 GB serialization limit, eliminates out-of-memory failures, and significantly improves backfill and pipeline performance across production metric families.
GitFarm: Git® as a Service for Large-Scale Monorepos
Uber introduces GitFarm, a centralized Git-as-a-Service platform designed to remove the cost of local repository checkouts for large monorepos. The post explains the system’s gateway/backend/sandbox architecture, gRPC-based execution model, and pooling strategy, then shows how it reduced CPU, memory, and startup latency for several production workflows.
Zone-Failure-Resilient OpenSearch® at Uber
This post explains how Uber built zone-failure-resilient OpenSearch deployments by combining shard allocation awareness with its internal isolation group infrastructure on Odin. It details how forced shard allocation awareness, balanced node placement, and cluster manager quorum settings help Uber maintain search and ingestion availability during zone outages and additional node failures.
Scaling Verify with Wallet for Identity Verification at Uber
Uber describes how it scaled Apple’s Verify with Wallet API across its Identity Verification Platform to support multiple use cases with stronger privacy and less user friction. The post details the backend architecture for scoping requests, decrypting and validating wallet-based identity payloads, and maintaining trust anchor certificates at scale for compliant digital ID verification.
Cart Assistant: Agentic Grocery Shopping on Uber Eats
Uber describes Cart Assistant, an agentic grocery-shopping feature for Uber Eats that turns prompts or images into a draft cart shoppers can review and edit. The post explains the underlying multi-prompt state graph architecture, including cart planning, candidate retrieval, semantic relevance judging, quantity selection, guardrails, and evaluation-driven development to keep results grounded, safe, and performant.
Simplifying Data and Product Integrations with a Data Abstraction Layer
Uber describes a Data Abstraction Layer (DAL) designed to decouple data consumers from evolving underlying datasets and schemas. The post explains how the DAL resolves logical tables to physical sources, generates and executes queries across multiple databases, and assembles results to support flexible advertiser reporting with far faster turnaround times.
Building a File Semantic Analyzer: Guarding Outbound Data at Scale with AI
Uber describes a File Semantic Analyzer that uses Generative AI to understand the meaning and context of files leaving the organization, rather than relying on brittle keyword-based DLP rules. The system ingests and preprocesses diverse file types, applies OCR and chunking for LLM analysis, then produces summaries, extracted entities, and intent-based classifications to reduce false positives and speed up security response.
Modernizing Artifact Storage at Uber
Uber describes how it modernized artifact storage by replacing a legacy on-prem, disk-based system with a managed SaaS platform. To control egress costs and preserve correctness, the team built an internal validation proxy that uses conditional requests, improves observability, and maintains high availability across regions. The post also details the legacy system’s failure modes, the proxy’s safety mechanisms, and the performance and reliability gains achieved at scale.