Breakpoint

Data & Databases

Databases and the pipelines feeding them — Postgres, MySQL, Redis and DynamoDB, plus Spark, ETL, and warehouse and lakehouse architecture. Expect query plans, migration stories, and the occasional hard-won lesson about what stops working at a billion rows.

Real-Time Retail Intelligence: Building E-Commerce Recommendations with Lakebase and AI Search on Databricks
Databricks presents a production-grade architecture for real-time e-commerce recommendations, combining Zerobus clickstream ingestion, AI Search candidate retrieval, Lakebase online features, and Model Serving. The design covers batch and session-aware serving paths, cold-start handling, business-rule reranking, latency fallbacks, model monitoring, and continuous improvement.
Scale without limits: Multigres, OrioleDB, and dbarena
Supabase introduces Multigres for PostgreSQL connection pooling and automated failover, OrioleDB for undo-log storage without table bloat or routine VACUUM, and dbarena for reproducible provider benchmarks. The post explains their architecture, scaling trade-offs, availability model, transaction-ID handling, and reported performance results.
How we extended Apache DataFusion to execute one query across many machines
Datadog explains how it extended Apache DataFusion into Distributed DataFusion, an open-source Rust framework for executing interactive queries across multiple machines. The post details physical-plan distribution, network shuffles, aggregation strategies, benchmarks, and design lessons such as avoiding distribution overhead for small queries.
Designing Neki for performance
The post explains how Neki improves sharded PostgreSQL performance by preserving wire-format messages and decoding data lazily. It details composable representations for routing, limits, sorting, projection, joins, and aggregation, showing how byte-level reuse avoids unnecessary allocations, parsing, and serialization.
Trading off compute for memory with activation checkpointing
Jane Street describes an activation-checkpointing planner for PyTorch training that trades recomputation for lower peak memory. The approach models full-graph memory lifetimes and recomputation costs, then greedily removes saves under an absolute memory budget, outperforming existing policies across tested models.
How to choose your first Genie Agents for maximum impact
Databricks presents a five-criteria rubric for selecting high-impact Genie Agent workflows, evaluating business impact, demand, data readiness, scope, and governance. The post also explains how executive champions, metadata quality, and narrowly defined use cases influence adoption and when teams should build, refine, or defer an agent.
Materialize, Out-of-Core: Replacing Swap with a Buffer Pool
Materialize explains its shift from kernel-managed paging to an application-managed buffer pool for out-of-core processing. The design uses columnar, relocatable data, explicit residency policies, and batched asynchronous lookups to exploit the throughput of NVMe while preserving memory for latency-sensitive work.
Evolving our calendar assistant Reclaim to be AI-native without starting over
Dropbox explains how Reclaim evolved into an AI-native calendar assistant without replacing its existing scheduling foundation. The design combines an agent loop, controlled tools and context, shared Schedule Actions, MCP integration, and a Redis-backed Preview Mode so users can review calendar changes before applying them.
8 major updates to Cloudflare Observability
Cloudflare announces eight observability updates that unify logs, traces, analytics, alerts, dashboards, SQL querying, and telemetry export across its platform. The release adds end-to-end tracing, agent-accessible observability APIs, custom alerting, OpenTelemetry support, and usage-based pricing changes.
Supabase Select 2026 Recap
Supabase’s 2026 Select recap introduces code-first schema and configuration workflows, native local development, agent-integrated health checks, MCP support, and database observability tools. It also covers new scaling options, including Multigres high availability, OrioleDB, and reproducible database benchmarks.
Lakebase Postgres branch-based restores for fast recovery at scale
Databricks explains how Lakebase Postgres uses decoupled compute and storage, immutable database history, and metadata-only branches to replace copy-and-replay restores. The approach enables near-instant point-in-time recovery—even for 100 TB databases—and supports agent-driven undo and versioning workflows.
Cut your AI spend with AI Gateway's Auto Router
Cloudflare explains how AI Gateway’s Auto Router classifies request complexity and task type, then balances model quality, token pricing, cache costs, and availability to select an appropriate model. Internal benchmarks show comparable task performance at substantially lower cost than frontier-only routing.
Operate with confidence
Supabase introduces tools for coding agents to observe projects, investigate production issues, test fixes, and operate within scoped permissions. The update adds SQL-accessible logs, health checks, connection diagnostics, notebooks, safer MCP controls, and near-real-time Postgres pipelines to analytical destinations.
Handling hot shards
The post explains how tenant growth and changing access patterns can make an initially sensible shard key produce hot shards. It compares vertical scaling and tenant isolation with table-specific resharding, showing how declarative topology and online data movement can redistribute load without downtime.