Whatnot describes how a third-party script caused livestream video to fail while component dashboards remained healthy. It explains the company’s event-correlated, protobuf-based end-to-end SLO system, including user-focused measurement, Kafka and ClickHouse processing, telemetry reliability, and production lessons.
Whatnot engineering blog
Whatnot Engineering explains how it generates typed Swift bindings, tests, and lifecycle safeguards from designer-authored Rive animation assets. The approach replaces string literals and tribal knowledge with compiler-checked APIs, CI-enforced code generation, asset linting, snapshot tests, and centralized runtime behavior.
Whatnot describes how its data platform supports billions of daily marketplace events, thousands of datasets, and hundreds of developers. The post details using governed AI and Snowflake event tables to improve data discovery, observability, alerting, and platform reliability.
The author explains how a carefully maintained regex filter for Android and Gradle build output reduced AI coding-agent transcript size by 20–30% without reducing task accuracy. The post details the filter’s design, its integration into AGENTS.md, results from more than 60 runs, and practices for keeping agent instructions reliable.
Whatnot’s Chief Product Officer explains how the company operates a small, senior product organization organized around customer problems rather than engineering teams. The post details an eight-step product process emphasizing system-level thinking, rapid validation, independent launches, continuous iteration, and AI-augmented individual contribution.
Whatnot Engineering explains how its hourly ML feature pipeline degraded through silent scheduling changes, runtime growth, and noisy anomaly alerts. The post details layered monitoring, freshness SLOs, alert tiers, and TTL-based graceful degradation for reliable real-time feature serving.
Whatnot’s LLM Platform treats the model as only one component of a production system built around velocity, trust, and reliability. The authors detail output-aware prompt experimentation, a reusable tool catalog, calibrated LLM judges, data mining for evaluation, and multi-provider resiliency patterns.
Whatnot explains how it replaced manual taxonomy record edits with validated, schedulable migrations across multiple marketplace taxonomies. The approach uses batched add, move, and delete operations, dry runs, automated checks for graph inconsistencies, and rollback safeguards to keep discovery experiences aligned with rapidly changing demand.
Whatnot describes Atlas, a Python-based workflow engine that converts branching support SOPs into executable, version-controlled graphs. The post covers an LLM map-reduce formalization pipeline, layered validation harness, human review, and the trade-offs of storing workflows in code rather than a database.
Whatnot details how it scaled a live shopping event to 583,000 concurrent viewers with admission control, connection-pooling proxies, feed fallbacks, resilient video delivery, and production load testing. The post also examines event-day pressure points and the durable infrastructure improvements that followed.