Breakpoint

Datadog engineering blog

How we extended Apache DataFusion to execute one query across many machines
Datadog explains how it extended Apache DataFusion into Distributed DataFusion, an open-source Rust framework for executing interactive queries across multiple machines. The post details physical-plan distribution, network shuffles, aggregation strategies, benchmarks, and design lessons such as avoiding distribution overhead for small queries.
How we built an async-aware Python profiler
Datadog engineers explain how they reconstructed asyncio task relationships into “stacked stacks” for more meaningful Python flame graphs. They also detail replacing process_vm_readv with protected memcpy and other optimizations that cut profiler overhead by more than 60%.
Making Rust observability reliable at scale with OpenTelemetry
Datadog explains how it built dd-trace-rs, an opinionated Rust tracer on OpenTelemetry, to standardize context propagation, sampling, and service configuration. The post details upstream interoperability fixes, deferred sampling and trace-level buffering, and reports 20x lower ingestion volume with 3x more indexed spans per service.
Unbiased Java CPU profiling with JFR in JDK 25
Datadog engineers explain why JFR’s traditional ExecutionSample event can produce biased CPU profiles, especially for reactive and CPU-saturated Java workloads. The post details the trade-offs of AsyncGetCallTrace and unsupported JVM internals, and shows how cooperative stack walking and the new experimental CPUTimeSample event in JDK 25 provide safer, more accurate profiling.
How we measure data completeness at scale
Datadog engineers describe a segment-based system for measuring data completeness across distributed ingestion pipelines. The approach combines idempotent create and acknowledgment tracking, adaptive sampling, custom in-memory storage, topology metadata, and resilient deployment patterns to support incident detection and safe automated decisions.
How we migrated a live routing system using AI-assisted refactoring
Datadog describes migrating its live Stream Router from a FoundationDB key-value model to PostgreSQL and DuckDB without disrupting production traffic. The post details AI-assisted, test-driven refactoring, blue/green validation against live data, schema trade-offs, and performance gains including 3,000x faster operations and 90% lower database costs.
When failover isn’t safe: Building high-availability PostgreSQL on Kubernetes
Datadog engineers explain how a Kubernetes-hosted PostgreSQL architecture failed to fail over safely during a zonal network disruption. They detail a Patroni- and ZooKeeper-coordinated redesign using synchronous replication, configuration tuning, benchmarking, and failure testing to balance durability, availability, and write performance.
Steganography at scale: Embedding share URLs in Datadog widget screenshots
Datadog engineers explain how they embed widget metadata into screenshots using resilient pixel-level watermarks. The post covers fuzzy color encoding, Redis-backed caching, device-density and color-profile handling, web workers, batching, compression, and performance optimizations that support more than a billion watermarks per day.
How we built a real-world evaluation platform for autonomous SRE agents at scale
Datadog explains how it built a replayable evaluation platform for its autonomous SRE agent, using real incident labels, reconstructed telemetry snapshots, noisy simulated environments, and historical scoring. The platform automates label creation and validation while detecting regressions across agent versions and production scenarios.
When upserts don’t update but still write: Debugging Postgres performance at scale
Datadog engineers investigate why a PostgreSQL upsert that usually skipped updates still doubled disk writes and quadrupled WAL syncs. They use pg_walinspect and debugger traces to reveal implicit row locking, then replace ON CONFLICT DO UPDATE with a data-modifying CTE that avoids unnecessary WAL overhead while documenting its concurrency trade-offs.
How we reduced the size of our Agent Go binaries by up to 77%
Datadog explains how it reduced Agent Go binary sizes by up to 77% without removing features. The team details dependency graph analysis, package refactoring, linker dead-code elimination, reflection pitfalls, and build-tag changes that produced substantial artifact reductions.
Detecting malicious pull requests at scale with LLMs
Datadog engineers describe how they built BewAIre, an LLM-powered system that reviews pull requests for malicious intent at scale. The post covers dataset curation, prompt tuning, recursive diff chunking, false-positive reduction, adversarial testing, and production performance.