Databricks presents a production-grade architecture for real-time e-commerce recommendations, combining Zerobus clickstream ingestion, AI Search candidate retrieval, Lakebase online features, and Model Serving. The design covers batch and session-aware serving paths, cold-start handling, business-rule reranking, latency fallbacks, model monitoring, and continuous improvement.
Data & Databases
Databases and the pipelines feeding them — Postgres, MySQL, Redis and DynamoDB, plus Spark, ETL, and warehouse and lakehouse architecture. Expect query plans, migration stories, and the occasional hard-won lesson about what stops working at a billion rows.
Supabase introduces Multigres for PostgreSQL connection pooling and automated failover, OrioleDB for undo-log storage without table bloat or routine VACUUM, and dbarena for reproducible provider benchmarks. The post explains their architecture, scaling trade-offs, availability model, transaction-ID handling, and reported performance results.
Datadog explains how it extended Apache DataFusion into Distributed DataFusion, an open-source Rust framework for executing interactive queries across multiple machines. The post details physical-plan distribution, network shuffles, aggregation strategies, benchmarks, and design lessons such as avoiding distribution overhead for small queries.
The post explains how Neki improves sharded PostgreSQL performance by preserving wire-format messages and decoding data lazily. It details composable representations for routing, limits, sorting, projection, joins, and aggregation, showing how byte-level reuse avoids unnecessary allocations, parsing, and serialization.
Cloudflare introduces Basin, a generally available serverless analytics platform built on Apache Iceberg and R2 Object Storage. The post explains its ingestion pipelines, managed catalog maintenance, distributed SQL engine, data portability, scalability improvements, and zero-egress pricing model.
Jane Street describes an activation-checkpointing planner for PyTorch training that trades recomputation for lower peak memory. The approach models full-graph memory lifetimes and recomputation costs, then greedily removes saves under an absolute memory budget, outperforming existing policies across tested models.
Databricks presents a five-criteria rubric for selecting high-impact Genie Agent workflows, evaluating business impact, demand, data readiness, scope, and governance. The post also explains how executive champions, metadata quality, and narrowly defined use cases influence adoption and when teams should build, refine, or defer an agent.
Amazon Aurora PostgreSQL can now query live operational data alongside Apache Iceberg and Parquet data in S3 without ETL pipelines. The post explains DuckDB integration, foreign-table setup, Glue catalog federation, query optimizations, caching, and options for materializing hot data for lower latency.
Materialize explains its shift from kernel-managed paging to an application-managed buffer pool for out-of-core processing. The design uses columnar, relocatable data, explicit residency policies, and batched asynchronous lookups to exploit the throughput of NVMe while preserving memory for latency-sensitive work.
NVIDIA announces a 64GB DGX Spark configuration for running AI agents, inference, fine-tuning and data science workloads locally. The post explains how two systems can pool 128GB of memory through NVIDIA Sync Cluster Assistant, delivering up to 1.7x performance for larger models and workloads.
Dropbox explains how Reclaim evolved into an AI-native calendar assistant without replacing its existing scheduling foundation. The design combines an agent loop, controlled tools and context, shared Schedule Actions, MCP integration, and a Redis-backed Preview Mode so users can review calendar changes before applying them.
Cloudflare announces eight observability updates that unify logs, traces, analytics, alerts, dashboards, SQL querying, and telemetry export across its platform. The release adds end-to-end tracing, agent-accessible observability APIs, custom alerting, OpenTelemetry support, and usage-based pricing changes.
Databricks makes native IP functions generally available for SQL, PySpark, and Scala, enabling parsing, validation, canonicalization, IPv4/IPv6 conversion, and CIDR containment without UDFs or regex. The Photon-optimized implementation supports high-volume network analytics and delivers up to 3.1x faster and 6.4x cheaper CIDR joins in benchmarks.
Amazon S3 Tables now support the full Apache Iceberg V3 type system, including variant, geospatial, and nanosecond timestamp types, plus deletion vectors and row lineage. The post explains how to create or upgrade tables, query semi-structured data, build incremental pipelines, and account for engine and format compatibility.
Supabase’s 2026 Select recap introduces code-first schema and configuration workflows, native local development, agent-integrated health checks, MCP support, and database observability tools. It also covers new scaling options, including Multigres high availability, OrioleDB, and reproducible database benchmarks.
Databricks explains how Lakebase Postgres uses decoupled compute and storage, immutable database history, and metadata-only branches to replace copy-and-replay restores. The approach enables near-instant point-in-time recovery—even for 100 TB databases—and supports agent-driven undo and versioning workflows.
Pinterest explains how it restores reliable partition-completeness signals in a streaming database ingestion system. The design uses event-time statistics, mergeable t-digest sketches, Iceberg snapshot metadata, and non-regressing watermarks implemented through extensible Flink sink hooks.
Cloudflare explains how AI Gateway’s Auto Router classifies request complexity and task type, then balances model quality, token pricing, cache costs, and availability to select an appropriate model. Internal benchmarks show comparable task performance at substantially lower cost than frontier-only routing.
Supabase introduces tools for coding agents to observe projects, investigate production issues, test fixes, and operate within scoped permissions. The update adds SQL-accessible logs, health checks, connection diagnostics, notebooks, safer MCP controls, and near-real-time Postgres pipelines to analytical destinations.
The post explains how tenant growth and changing access patterns can make an initially sensible shard key produce hot shards. It compares vertical scaling and tenant isolation with table-specific resharding, showing how declarative topology and online data movement can redistribute load without downtime.