Breakpoint

Expedia engineering blog

Building the Operating Model for Agentic AI at Expedia Group
Expedia Group describes an operating model for scaling agentic AI across a large organization, combining shared platforms, risk-based governance, observability, and cross-team collaboration. The post explains why specialized agents, early evaluations, release tollgates, and model optionality help teams move from experimentation to reliable production systems.
The Context Tax of Agentic Development
The post explains how missing or hidden context can cause AI agents to accelerate incorrect software work rather than merely slow development. It presents a context repository using proposal documents, machine-readable STATUS.yaml files, repository instructions, and reusable agent skills to improve alignment, validation, and research, while acknowledging the ongoing challenge of context drift.
How Keras 3 Helped Modernise Expedia Group’s Lodging Ranking Stack
Expedia Group describes how a Keras 3 migration became a broader modernisation of its lodging ranking stack. The work improved training speed, enabled XLA-safe and more backend-agnostic model code, and separated preprocessing from scoring to unlock faster serving with ONNX, TensorRT, and Triton.
Focus on the Feature, Not the Fixture: GenAI powered GraphQL mocks
Explains how Expedia Group uses an LLM plus GraphQL schema-aware directives to generate contextual mock responses instead of maintaining brittle handwritten fixtures. The post walks through the mockql-rs workflow, including parsing, splitting mocked versus real fields, prompting, and merging results back into a valid GraphQL response.
How Expedia Group Builds AI That Lasts at Scale
Expedia Group outlines principles for building AI systems that create business value, scale across teams, and operate safely over time. The post emphasizes measurable outcomes, shared foundations, governance, reproducibility, and continuous monitoring for agentic and generative AI systems.
Just Because We Can Build It, Should We?
How AI changed the build-vs-buy equation for platform engineering, and why discipline matters more than ever. The post argues for composable, standards-based building blocks over proprietary, overly broad internal platforms.
Expedia’s Service Telemetry Analyzer
A system that facilitates investigation of service degradations and outages using service telemetry data and AI. The post introduces STAR, Expedia’s early prototype for multi-step diagnostic workflows over telemetry, and explains its web service, AI model usage, prompt chaining, and rollout lessons.
Reimagining Platform Engineering for an Agentic Future
Expedia Group explores how platform engineering must evolve to treat agents (AI systems) as first-class users alongside humans. The post outlines changes including agent-native CLIs (Tarmac), MCP servers, markdown-based skills, and internal agent-focused apps (Koda) to reduce friction, improve ergonomics, and provide robust guardrails for agent interactions.
Operating Trino at Scale With Trino Gateway
This post describes using a Trino Gateway to route queries, improve security, and simplify cluster management for Trino deployments at scale. It outlines common cluster types and workload segregation, and details recent contributions that add UI-based routing rules, source filtering for history, cluster health displays, and full-query viewing in the Gateway.
Managing Technical Debt: Building Habits for Long-Term Agility
A practical piece outlining a lightweight framework for managing technical debt through habits: capture pain points, maintain a dedicated tech-debt backlog, review monthly, and reserve 10–15% of sprint capacity to pay down debt. The author emphasizes engineering ownership of debt and argues that consistent small investments preserve long-term agility and team morale.
Interleaving for Accelerated Testing
Explains interleaving as a within-subject alternative to A/B testing for ranking/recommender systems that is far more sensitive (reported 10x–100x improvements). The post covers attribution of clicks/bookings, a lift metric for preference, significance testing approaches (bootstrapping vs t-test), practical gains in sensitivity, and limitations when interpreting results or assuming item independence.
The Quant Crossroads: UX Research or Data Science
Compares Quantitative UX Research (QUXR) and Data Science, highlighting that both work with large datasets but differ in goals and methods: QUXRs focus more on research design and the ‘why’ (often using R and advanced statistics) while data scientists focus on prediction, modelling and engineering (SQL, Python). The article outlines complementary skill sets and how the roles together provide a fuller understanding of user behavior for product teams.
Powering Vector Embedding Capabilities
The Machine Learning Platform team describes a centralized Embedding Store Service to manage, store, and query vector embeddings at scale. The service integrates with Feast for metadata and discoverability, supports batch and real-time ingestion, maintains online/offline stores, and provides similarity and hybrid search capabilities to simplify building embedding-based ML use cases.
A Functional Programming Alternative to the Strategy Pattern
This post contrasts the object-oriented Strategy Pattern with a functional programming alternative in Kotlin, showing how handlers can be represented as lambdas and how shared logic can be composed via higher-order functions. It compares trade-offs in maintainability, boilerplate, and hierarchy complexity to help choose the right approach for a codebase.
Reimagining Software Engineering: LLMs, MCP, and the Dawn of a New Programming Paradigm
The article explores how Large Language Models combined with the Model Context Protocol (MCP) could shift programming from writing explicit code to collaborating with operationally capable agents that can perform CRUD on real resources. It discusses the potential paradigm shift, practical examples, and considerations for making LLMs operational participants in software development.
Why You Should Prefer MERGE INTO Over INSERT OVERWRITE in Apache Iceberg
This article argues for using MERGE INTO (particularly Merge-on-Read) instead of INSERT OVERWRITE when performing row-level updates in Apache Iceberg. It explains the performance and cost benefits of MOR—faster writes, reduced I/O, and lower storage costs—while noting the importance of compaction and maintenance to avoid fragmentation and metadata bloat.
Chill Your Data with Iceberg Write Audit Publish
Describes the Write‑Audit‑Publish (WAP) workflow using Apache Iceberg branches to stage, validate, and atomically publish data changes without duplicating production tables. The post details the write/audit/publish steps, branching and tagging usage, and the cost and governance benefits—reducing duplicated work and enabling lightweight, auditable promotions to production.