Uber describes MCP Gateway, a centralized platform for discovering, governing, and executing more than 800 MCP servers and 5,000 tools. The architecture combines an AutoCrawler control plane, protocol translation across HTTP, gRPC, and TChannel, and built-in authorization, redaction, and observability for scaling agent integrations.
Uber engineering blog
Uber describes a feature-logging framework that makes inference-time features the canonical source for model training, eliminating training-serving skew. The post covers selective logging, compact feature aliases, Flink stream joins, state management, profiling, and Kafka scaling techniques that improved freshness, consistency, and throughput.
Uber Engineering explains how error ownership makes retries context-aware by distinguishing errors caused by a service from downstream failures it merely propagates. The production system limits cascading retry storms across deep service graphs, reducing the maximum storm radius from 25 levels to 3 and preventing up to 9.5 million unnecessary requests during an outage.
Uber engineers explain how they automatically map fail-close and fail-open dependencies across a large service mesh. The approach correlates inbound and outbound failures through middleware, aggregates every request with metrics, and applies probability thresholds to identify reliability-critical call paths.
Uber explains how smart event correlation adds context to observability signals, helping teams reduce alert fatigue and turn noisy alerts into actionable insights. The post connects operational monitoring practices with lessons from ultra-distance racing.
Uber Eats engineers detail how they cut search latency in half by optimizing the full-stack pipeline, including retrieval, hydration, ranking, ad serving, serialization, and presentation. The post explains measured techniques such as request hedging, product-level grouping, parallel rendering, GPU serving, and agent-assisted performance optimization.
Uber explains how it evolved from a single scaling controller to a multi-orchestrator Kubernetes control plane using a dedicated ServiceScale CRD and controller. The post details lessons from stale informer caches, terminal status fields, multi-writer races, automated healing, and large-scale rollout testing.
Uber engineers explain how M3DB’s subclustered placement algorithm reduces shard-migration blast radius, limits cross-node dependencies, and enables safer parallel maintenance. The post details its structural invariants, greedy shard-donation strategy, complexity, and operational trade-offs.
Uber engineers explain how they scaled agentic coding usage 7x while stabilizing AI spend. The post details benchmark-driven model routing, prompt-cache tuning, CLI-based MCP access, code-mode batching, context graphs, and session-level cost analytics.
Uber explains how selective export workloads over large historical datasets can trigger expensive full-table scans, keeping objects hot in Google Cloud Storage and increasing storage, retrieval, metadata, and egress costs. The post shows that combining Hudi column statistics with predicate-column sorting can dramatically reduce the scan surface area, improve pruning, and preserve colder storage tiers for long-term savings.
Uber describes the decade-long evolution of its Payments Platform, highlighting the core design principles that made it durable at scale: immutable money orders, zero-sum accounting, strongly consistent ledger balances, and loosely coupled microservices. The post also explains how the platform remained line-of-business and payment-instrument agnostic while expanding across new Uber products, payment methods, and high-throughput ledger use cases.
The post appears to examine the design and evolution of Uber’s payments platform over a decade. The body was unavailable, so this summary is based on the title alone.
Uber describes how it scaled exact COUNT(DISTINCT) for high-cardinality non-rollup metrics in distributed data pipelines without relying on approximate counting. The post explains a chunked aggregation buffer strategy that avoids the JVM’s 2 GB serialization limit, eliminates out-of-memory failures, and significantly improves backfill and pipeline performance across production metric families.
Uber introduces GitFarm, a centralized Git-as-a-Service platform designed to remove the cost of local repository checkouts for large monorepos. The post explains the system’s gateway/backend/sandbox architecture, gRPC-based execution model, and pooling strategy, then shows how it reduced CPU, memory, and startup latency for several production workflows.
This post explains how Uber built zone-failure-resilient OpenSearch deployments by combining shard allocation awareness with its internal isolation group infrastructure on Odin. It details how forced shard allocation awareness, balanced node placement, and cluster manager quorum settings help Uber maintain search and ingestion availability during zone outages and additional node failures.
Uber describes how it scaled Apple’s Verify with Wallet API across its Identity Verification Platform to support multiple use cases with stronger privacy and less user friction. The post details the backend architecture for scoping requests, decrypting and validating wallet-based identity payloads, and maintaining trust anchor certificates at scale for compliant digital ID verification.
Uber describes Cart Assistant, an agentic grocery-shopping feature for Uber Eats that turns prompts or images into a draft cart shoppers can review and edit. The post explains the underlying multi-prompt state graph architecture, including cart planning, candidate retrieval, semantic relevance judging, quantity selection, guardrails, and evaluation-driven development to keep results grounded, safe, and performant.
Uber describes a Data Abstraction Layer (DAL) designed to decouple data consumers from evolving underlying datasets and schemas. The post explains how the DAL resolves logical tables to physical sources, generates and executes queries across multiple databases, and assembles results to support flexible advertiser reporting with far faster turnaround times.
Uber describes a File Semantic Analyzer that uses Generative AI to understand the meaning and context of files leaving the organization, rather than relying on brittle keyword-based DLP rules. The system ingests and preprocesses diverse file types, applies OCR and chunking for LLM analysis, then produces summaries, extracted entities, and intent-based classifications to reduce false positives and speed up security response.
Uber describes how it modernized artifact storage by replacing a legacy on-prem, disk-based system with a managed SaaS platform. To control egress costs and preserve correctness, the team built an internal validation proxy that uses conditional requests, improves observability, and maintains high availability across regions. The post also details the legacy system’s failure modes, the proxy’s safety mechanisms, and the performance and reliability gains achieved at scale.