Datadog explains how it extended Apache DataFusion into Distributed DataFusion, an open-source Rust framework for executing interactive queries across multiple machines. The post details physical-plan distribution, network shuffles, aggregation strategies, benchmarks, and design lessons such as avoiding distribution overhead for small queries.
Datadog engineering blog
Datadog engineers explain how they reconstructed asyncio task relationships into “stacked stacks” for more meaningful Python flame graphs. They also detail replacing process_vm_readv with protected memcpy and other optimizations that cut profiler overhead by more than 60%.
Datadog explains how it built dd-trace-rs, an opinionated Rust tracer on OpenTelemetry, to standardize context propagation, sampling, and service configuration. The post details upstream interoperability fixes, deferred sampling and trace-level buffering, and reports 20x lower ingestion volume with 3x more indexed spans per service.
Datadog explains how it rebuilt Git serving with independent mirrors, autoscaled relays, content-addressed pack reuse, and a pack cache. The architecture handled 20× traffic growth at roughly 40 ms median latency while reducing legacy backend fetch CPU usage by three to four times.
Datadog explains how its APM team encoded a prefix trie as a JVM string constant to reduce Java startup overhead during cold, pre-JIT execution. The post details the compact encoding, matching algorithm, benchmark results, and resulting startup improvements across Java versions.
Datadog engineers explain why JFR’s traditional ExecutionSample event can produce biased CPU profiles, especially for reactive and CPU-saturated Java workloads. The post details the trade-offs of AsyncGetCallTrace and unsupported JVM internals, and shows how cooperative stack walking and the new experimental CPUTimeSample event in JDK 25 provide safer, more accurate profiling.
Datadog engineers describe a segment-based system for measuring data completeness across distributed ingestion pipelines. The approach combines idempotent create and acknowledgment tracking, adaptive sampling, custom in-memory storage, topology metadata, and resilient deployment patterns to support incident detection and safe automated decisions.
Datadog describes migrating its live Stream Router from a FoundationDB key-value model to PostgreSQL and DuckDB without disrupting production traffic. The post details AI-assisted, test-driven refactoring, blue/green validation against live data, schema trade-offs, and performance gains including 3,000x faster operations and 90% lower database costs.
Datadog engineers explain how a Kubernetes-hosted PostgreSQL architecture failed to fail over safely during a zonal network disruption. They detail a Patroni- and ZooKeeper-coordinated redesign using synchronous replication, configuration tuning, benchmarking, and failure testing to balance durability, availability, and write performance.
Datadog explains how BewAIre evolved from pull-request analysis into scalable software supply-chain protection using stacked LLM evaluations, agentic investigation, codemaps, and targeted file reads. The post details accuracy, latency, cost, package-crawling, caching, and version-diff strategies for scanning npm and PyPI dependencies.
Datadog engineers explain how they embed widget metadata into screenshots using resilient pixel-level watermarks. The post covers fuzzy color encoding, Redis-backed caching, device-density and color-profile handling, web workers, batching, compression, and performance optimizations that support more than a billion watermarks per day.
Datadog explains how it built a replayable evaluation platform for its autonomous SRE agent, using real incident labels, reconstructed telemetry snapshots, noisy simulated environments, and historical scoring. The platform automates label creation and validation while detecting regressions across agent versions and production scenarios.
Datadog engineers investigate why a PostgreSQL upsert that usually skipped updates still doubled disk writes and quadrupled WAL syncs. They use pg_walinspect and debugger traces to reveal implicit row locking, then replace ON CONFLICT DO UPDATE with a data-modifying CTE that avoids unnecessary WAL overhead while documenting its concurrency trade-offs.
Datadog details how the hackerbot-claw AI agent exploited unsafe GitHub Actions input and attempted prompt injection in open source repositories. The post explains BewAIre detection, the containment provided by least-privilege permissions, and practical safeguards for securing CI workflows and LLM-powered automation.
Reilly Wood shares lessons from building Datadog’s MCP server for AI agents, including token-efficient data formats, token-budget pagination, query-based access, toolset design, and actionable errors. The post explains how these patterns improve agent accuracy, context usage, and operating cost.
Datadog explains how it reduced Agent Go binary sizes by up to 77% without removing features. The team details dependency graph analysis, package refactoring, linker dead-code elimination, reflection pitfalls, and build-tag changes that produced substantial artifact reductions.
Guillaume Fournier distills six lessons from operating eBPF-powered workload protection across thousands of environments and kernel versions. The post examines compatibility, hook coverage, reliable kernel-data capture, cache consistency, performance costs, and safe rollout strategies for production runtime security.
Datadog engineers explain how they scaled eBPF-based file integrity monitoring to process more than 10 billion kernel events per minute. The system combines in-kernel approvers and discarders with Agent-side rule evaluation, filtering 94% of events while preserving detection coverage and reducing resource overhead.
Datadog explains how it built a multi-tenant change data capture platform using Debezium, Kafka, Temporal, and schema compatibility controls. The architecture moved search workloads off PostgreSQL, cut latency by up to 97%, and enabled reliable, customizable replication across diverse systems.
Datadog engineers describe how they built BewAIre, an LLM-powered system that reviews pull requests for malicious intent at scale. The post covers dataset curation, prompt tuning, recursive diff chunking, false-positive reduction, adversarial testing, and production performance.