Breakpoint

Notion engineering blog

How Notion handles concurrent editing with CRDTs
Notion explains how it built a CRDT-based rich-text editing system to merge concurrent changes without losing users’ work. The post covers RGA text structures, Peritext annotations, tombstones, and text-slice search labels for preserving edits across block splits and offline synchronization.
Building Shared Memory for AI Agents in Notion
Notion explains Lore, an open-source shared memory system that stores agent experiences, tasks, decisions, procedures, and structured facts in interconnected Notion databases. The post details its MCP-based architecture, memory-maintenance strategies, and benchmark results showing when relevant memory improves agent performance—and when stale or redundant context can harm it.
Rebuilding Notion’s lexical search reindexer
Notion explains how it rebuilt its lexical search reindexer as a Spark-native pipeline, replacing ECS workers, Snowflake lookups, and live Elasticsearch writes with offline snapshot generation. The design uses embedded Elasticsearch, NVMe-backed executors, validation for document parity, and versioned alias-based catchup to achieve faster, safer, fully consistent reindexing.
Building observability into Notion’s dead-letter queue
Notion explains how it built the DLQ Explorer to query and recover failed background tasks stored in S3. The system uses Athena partition projection, controlled multi-region access, query filtering, batched re-enqueueing, and audit trails to make dead-letter queue investigations faster and safer.
Enabling Multi-Region Data Systems at Notion
Notion explains how it redesigned its data infrastructure to keep workspace data processed and stored within its home region. The post details regional data lakes, Kafka and Spark pipelines, AI embeddings, Airflow orchestration, event routing, sanitization, and Terraform-based maintainability.
Updating the design of Notion pages
Notion explains how it redesigned page spacing to create a more consistent rhythm across flexible web content. The team replaced baseline-alignment goals with standardized spacing and adjacency rules that keep list items compact while giving paragraphs more breathing room.
How we built security into Custom Agents
Notion explains how Custom Agents use least-privilege, resource-level permissions, runtime safeguards, prompt-injection mitigation, and human confirmation to secure collaborative AI workflows. It also shares lessons from testing more than 28,000 agents, including why Slack needed a constrained “read and reply” permission.
Balancing cost and reliability for Spark on Kubernetes
Justin Lee explains how Notion reduced Spark compute costs by 60–90% on Kubernetes using dynamic provisioning, bin packing, and AWS Spot Instances. The post details why large-scale interruptions caused Spark failures and how the open-source Spot Balancer distributes executors across spot and on-demand capacity to preserve reliability.
Meet Scruff, Security's New AI Teammate
Notion’s security team built Scruff, an AI teammate that uses Custom Agents, MCP integrations, alerts, runbooks, and shared notes to automate security-alert triage and investigation. The case study details the workflow and reports 84% lower median investigation time and 93% faster resolution for false positives.