Zalando explains the architecture of its open-source Agentic Identity Broker, which enables user-to-agent delegation and token exchange without exposing provider tokens to agents. The post covers consent models, Kubernetes integration, gateway enforcement, policy controls, and standards such as OAuth 2.0, SPIFFE, and CIBA.
Zalando engineering blog
The post explains how Zalando Marketing Services addressed budget cannibalization and SUTVA violations in a two-sided advertising marketplace. It details Budget Split and Orthogonal Concurrency, which isolate experiment variants into independent campaign budgets and enable reliable concurrent testing while accounting for residual sources of bias.
Zalando shares two and a half years of experience adopting agentic engineering across hundreds of teams. The post covers a multi-provider LLM proxy, MCP authentication, coding-agent adoption metrics, risk-based PR approval, agent skills, governance, training, and lessons for managing AI-amplified software complexity.
Zalando explains its migration from a homegrown in-memory ad-event join to Apache Flink, using keyed state, incremental checkpoints, and shadow validation to preserve events during scaling. The team details RocksDB tuning, connector and Kubernetes pitfalls, and how a custom KeyedCoProcessFunction reduced resources and infrastructure costs.
Zalando explains how it built an in-process client-side load balancer for more than a million requests per second of internal fan-out traffic. The post covers consistent-hash parity, Kubernetes discovery, N-ring cache-warming fade-in, occupancy- and latency-based bounded load, safe retries, and the operational lessons from AZ-aware routing experiments.
Zalando introduces an open-source Go translator that converts Lightstep UQL queries into PromQL through a lexer, parser, AST optimizer, and code-generation pipeline. The post explains its Web UI, REST API, SDK, and use in automating large-scale observability migrations.
Zalando explains how Skipper became a validating Kubernetes admission webhook that reuses its runtime filter and predicate registries to reject invalid ingress routes during apply. The post covers actionable error reporting, control-plane rollout strategy, metrics, feature flags, and operational outcomes.
Zalando describes a scalable search quality assurance framework that combines NER-based query clustering, LLM translation, and multimodal LLM-as-a-judge evaluation. The Airflow and Kubernetes pipelines tested 1,500 search segments per market, uncovering multilingual ranking and entity-recognition issues before launch while reducing evaluation cost and effort.
Sri Adarsh Kumar details Zalando’s migration of jackson-datatype-money into FasterXML’s official Jackson ecosystem. The case study covers module separation, dependency and package renaming, licensing, review practices, and the drop-in migration path for JSR 354 users.
Zalando explains how a Flink 1.20 stream-processing pipeline replaced chained Table API joins with a unified DataStream and custom keyed processor. The redesign reduced RocksDB state from 235GB to 56GB, shortened snapshots, improved stability, and cut AWS costs by 13%.
Zalando engineers share lessons from running a year-long internal papers-reading guild, covering paper selection, meetup formats, promotion, facilitation, and community participation. The post presents a practical blueprint for building a sustainable engineering learning group and shows how discussions informed production decisions around DynamoDB and Flink.
Zalando presents a replenishment engine that combines probabilistic LightGBM demand forecasts, an extended (R, s, Q) policy, and Monte Carlo discrete-event simulation. Backtesting across roughly two million articles shows improved GMV, availability, and fill rates, while ablation studies quantify the value of uncertainty-aware forecasting and 75th-percentile optimization.
Zalando explains how it contributed two Debezium features to make PostgreSQL logical replication safer at scale. The post details WAL growth caused by inactive databases, offset-versus-slot mismatches, new flush and recovery strategies, and the operational trade-offs behind each configuration.
Zalando details how a maintenance bug caused an internal application to issue high-cardinality Elasticsearch facet aggregations, overwhelming a production cluster and degrading search. The post explains the scatter/gather failure mode, incident mitigation, and follow-up improvements including client-aware observability, query limits, workload isolation, and runbooks.
Zalando explains its brownfield React Native migration, integrating an internal Rendering Engine with native iOS and Android applications through a package-based architecture and explicit API contracts. The post details progressive adoption, developer tooling, cross-platform UI with react-strict-dom, and lessons from deploying shared mobile experiences.
Zalando describes a multi-stage LLM pipeline that analyzes thousands of postmortems to identify recurring datastore failure patterns and guide reliability investments. The post details map-fold processing, prompt constraints, human curation, performance results, and limitations including hallucinations and surface attribution errors.
Zalando describes how it replaced fragmented partner data exports, APIs, and manual processing with a secure Delta Sharing platform. The post covers partner needs, evaluation criteria, Unity Catalog integration, token-based access, proof-of-concept limitations, platformization, and lessons for scaling governed data sharing.
Zalando details a scalable inventory optimisation system that combines probabilistic demand forecasting with Monte Carlo simulation and gradient-free optimisation. The post explains its AWS, Databricks, SageMaker, feature-store, batch, and real-time architecture, including how it processes forecasts for 5 million SKUs in under two hours.
Kanupriya Gupta recounts returning to Zalando after a four-month break and adapting to changed teams, priorities, and a modernized React Native and Swift development stack. She highlights documentation, supportive teammates, gradual re-onboarding, and deliberate prioritization as ways to regain momentum.
Zalando describes replacing a bottlenecked event-driven product-data pipeline with PRAPI, a high-performance serving layer delivering millions of requests per second with single-digit-millisecond P99 latency. The post details consistent-hash load balancing, stale-while-refresh caching, DynamoDB scaling, non-blocking I/O, and JVM tuning techniques.