Expedia Group describes an operating model for scaling agentic AI across a large organization, combining shared platforms, risk-based governance, observability, and cross-team collaboration. The post explains why specialized agents, early evaluations, release tollgates, and model optionality help teams move from experimentation to reliable production systems.
Expedia engineering blog
The post explains how missing or hidden context can cause AI agents to accelerate incorrect software work rather than merely slow development. It presents a context repository using proposal documents, machine-readable STATUS.yaml files, repository instructions, and reusable agent skills to improve alignment, validation, and research, while acknowledging the ongoing challenge of context drift.
Expedia Group describes how a Keras 3 migration became a broader modernisation of its lodging ranking stack. The work improved training speed, enabled XLA-safe and more backend-agnostic model code, and separated preprocessing from scoring to unlock faster serving with ONNX, TensorRT, and Triton.
Explains how Expedia Group uses an LLM plus GraphQL schema-aware directives to generate contextual mock responses instead of maintaining brittle handwritten fixtures. The post walks through the mockql-rs workflow, including parsing, splitting mocked versus real fields, prompting, and merging results back into a valid GraphQL response.
Expedia Group outlines principles for building AI systems that create business value, scale across teams, and operate safely over time. The post emphasizes measurable outcomes, shared foundations, governance, reproducibility, and continuous monitoring for agentic and generative AI systems.
An LLM-powered workflow analyzes Spark SQL execution plans and metrics to detect performance anti-patterns such as missing broadcast joins, skewed partitions, missing filters, oversized broadcasts, and excessive spills. The post explains how structured prompting and traceable JSON outputs helped reduce debugging time and cut runtime and cost on real workloads.
How a screen-level performance metric reshaped platform decisions, engineering ownership, and release discipline.
How AI changed the build-vs-buy equation for platform engineering, and why discipline matters more than ever. The post argues for composable, standards-based building blocks over proprietary, overly broad internal platforms.
A system that facilitates investigation of service degradations and outages using service telemetry data and AI. The post introduces STAR, Expedia’s early prototype for multi-step diagnostic workflows over telemetry, and explains its web service, AI model usage, prompt chaining, and rollout lessons.
Expedia Group explores how platform engineering must evolve to treat agents (AI systems) as first-class users alongside humans. The post outlines changes including agent-native CLIs (Tarmac), MCP servers, markdown-based skills, and internal agent-focused apps (Koda) to reduce friction, improve ergonomics, and provide robust guardrails for agent interactions.
This post describes using a Trino Gateway to route queries, improve security, and simplify cluster management for Trino deployments at scale. It outlines common cluster types and workload segregation, and details recent contributions that add UI-based routing rules, source filtering for history, cluster health displays, and full-query viewing in the Gateway.
A practical piece outlining a lightweight framework for managing technical debt through habits: capture pain points, maintain a dedicated tech-debt backlog, review monthly, and reserve 10–15% of sprint capacity to pay down debt. The author emphasizes engineering ownership of debt and argues that consistent small investments preserve long-term agility and team morale.
Explains interleaving as a within-subject alternative to A/B testing for ranking/recommender systems that is far more sensitive (reported 10x–100x improvements). The post covers attribution of clicks/bookings, a lift metric for preference, significance testing approaches (bootstrapping vs t-test), practical gains in sensitivity, and limitations when interpreting results or assuming item independence.
Compares Quantitative UX Research (QUXR) and Data Science, highlighting that both work with large datasets but differ in goals and methods: QUXRs focus more on research design and the ‘why’ (often using R and advanced statistics) while data scientists focus on prediction, modelling and engineering (SQL, Python). The article outlines complementary skill sets and how the roles together provide a fuller understanding of user behavior for product teams.
The Machine Learning Platform team describes a centralized Embedding Store Service to manage, store, and query vector embeddings at scale. The service integrates with Feast for metadata and discoverability, supports batch and real-time ingestion, maintains online/offline stores, and provides similarity and hybrid search capabilities to simplify building embedding-based ML use cases.
This post contrasts the object-oriented Strategy Pattern with a functional programming alternative in Kotlin, showing how handlers can be represented as lambdas and how shared logic can be composed via higher-order functions. It compares trade-offs in maintainability, boilerplate, and hierarchy complexity to help choose the right approach for a codebase.
An engineering post examining how Kafka Streams assigns partitions when consuming multiple topics and how sub-topology boundaries can prevent partition colocation. The team diagnosed misaligned partition assignments, then fixed it by introducing a shared Kafka Streams state store to unify sub-topologies so same-index partitions are colocated and caches are effective.
The article explores how Large Language Models combined with the Model Context Protocol (MCP) could shift programming from writing explicit code to collaborating with operationally capable agents that can perform CRUD on real resources. It discusses the potential paradigm shift, practical examples, and considerations for making LLMs operational participants in software development.
This article argues for using MERGE INTO (particularly Merge-on-Read) instead of INSERT OVERWRITE when performing row-level updates in Apache Iceberg. It explains the performance and cost benefits of MOR—faster writes, reduced I/O, and lower storage costs—while noting the importance of compaction and maintenance to avoid fragmentation and metadata bloat.
Describes the Write‑Audit‑Publish (WAP) workflow using Apache Iceberg branches to stage, validate, and atomically publish data changes without duplicating production tables. The post details the write/audit/publish steps, branching and tagging usage, and the cost and governance benefits—reducing duplicated work and enabling lightweight, auditable promotions to production.