Breakpoint

Databricks engineering blog

Real-Time Retail Intelligence: Building E-Commerce Recommendations with Lakebase and AI Search on Databricks
Databricks presents a production-grade architecture for real-time e-commerce recommendations, combining Zerobus clickstream ingestion, AI Search candidate retrieval, Lakebase online features, and Model Serving. The design covers batch and session-aware serving paths, cold-start handling, business-rule reranking, latency fallbacks, model monitoring, and continuous improvement.
In Capital Markets, the Buy Side Runs on NAV. Finance Protects the Fee.
The post explains how Databricks Genie One uses a governed business ontology and backend AI agents to support asset-management finance workflows. It covers overnight custodian-file ingestion, NAV reconciliation, traceable valuation data, liquidity monitoring, and fee-margin analysis with human approval retained.
How to choose your first Genie Agents for maximum impact
Databricks presents a five-criteria rubric for selecting high-impact Genie Agent workflows, evaluating business impact, demand, data readiness, scope, and governance. The post also explains how executive champions, metadata quality, and narrowly defined use cases influence adoption and when teams should build, refine, or defer an agent.
How Genie One reshapes work for finance teams
Databricks explains how Genie One gives finance teams an AI coworker grounded in governed metrics, fiscal logic, entity structures, and organizational terminology. The post outlines workflows for variance analysis, cash forecasting, close management, and spend analytics, including permission-aware automation and reusable scheduled reviews.
Lakebase Postgres branch-based restores for fast recovery at scale
Databricks explains how Lakebase Postgres uses decoupled compute and storage, immutable database history, and metadata-only branches to replace copy-and-replay restores. The approach enables near-instant point-in-time recovery—even for 100 TB databases—and supports agent-driven undo and versioning workflows.
Connecting customer context to measurable ROI with agentic marketing
Databricks explains how agentic marketing combines continuous identity resolution, customer and business context, human-defined guardrails, and incrementality measurement. It presents CustomerLake’s Profile and Campaign Agents as a governed way to turn customer signals into measurable, continuously optimized actions.
A practical guide to cost optimization with Lakebase Postgres
Databricks explains how Lakebase Postgres separates storage and compute to reduce costs through shared storage, autoscaling, branching, replicas, and scale-to-zero. It provides practical guidance on syncing only required data, choosing sync modes, sizing compute around the working set, pooling connections, and managing PITR and snapshot storage.
Introducing ai_decide: make fast decisions on your governed data
Databricks introduces ai_decide, a beta AI Function that converts unstructured text into structured probabilities, classifications, and scores. The feature is designed for lower-latency, lower-cost decisions over governed data, with SQL support for batch workloads and a REST API for real-time applications and agents.
How to scale agentic applications without creating AI sprawl
The post outlines a choice, context, and control architecture for scaling enterprise agentic applications without duplicating integrations and governance. It explains how shared business context, model and harness abstractions, permissions, observability, evaluation, and cost controls can support reliable agent fleets using Databricks capabilities.
Lakebase Search: State-of-the-art full text and vector search for Postgres
Databricks introduces Lakebase Search, a serverless Postgres search engine combining BM25 full-text search with approximate vector search. The post explains how hierarchical IVF clustering, binary quantization, storage-compute separation, and parallel index builds deliver scalable retrieval with reported 97% recall and 71 ms P99 latency on 100 million vectors.
How Databricks rolls out frontier models to 12,000 employees on Day 1
Databricks describes its playbook for giving 12,000 employees Day 1 access to frontier models while controlling cost and quality risk. The approach combines Unity Gateway and CLI-based configuration, per-user budget tiers, private benchmarks, user feedback, OpenTelemetry traces, and stratified session-cost analysis to decide which models become standard or default.
The 24 most commonly misunderstood marketing data terms
The guide explains how marketing and data engineering assign different meanings to terms such as customer, audience, real time, model, and consent. It provides practical ways to align data definitions, identity, freshness requirements, activation readiness, governance, and ownership before campaigns are built.
Manufacturing data and AI: Connecting the product value chain
The post explains how manufacturers can connect siloed product-value-chain data to trace defects, assess supplier risk, and support cross-stage decisions. It outlines a practical architecture combining federated or mirrored data, governed semantic layers, shared identifiers, streaming, and natural-language agents.
How to roll out Genie One: A step-by-step enterprise playbook
Databricks presents a phased enterprise playbook for rolling out Genie One, starting with a governed data domain and a focused pilot before expanding organization-wide. It covers semantic definitions, evaluation sets, ownership, permissions, observability, user adoption, and governance for reliable AI coworkers.
Running open-Jev in SQL on Databricks
Databricks demonstrates how to deploy the SemIf-OpenJev open-weight decision model with serverless GPUs and managed Model Serving. The workflow uses ai_query to invoke custom model APIs from SQL or Lakeflow jobs and return structured classifications and probability scores over governed data.
How I built agent-based security reviews on Databricks
The post details an agent-based security review system built with Unity Catalog, foundation models, Lakeflow Jobs, and Databricks Apps. It explains how focused agents, evidence-based risk assessment, conservative escalation, and human oversight automate routine reviews while preserving expert judgment for ambiguous or high-risk cases.