Notion explains how it built a CRDT-based rich-text editing system to merge concurrent changes without losing users’ work. The post covers RGA text structures, Peritext annotations, tombstones, and text-slice search labels for preserving edits across block splits and offline synchronization.
Notion engineering blog
Notion’s post examines how the company detects breaking schema changes before they affect production, likely covering safeguards in its development and deployment workflow. The post body was unavailable, so this summary is based on the title alone.
The post reports Notion’s experience scaling vector search by 10x while reducing costs to one-tenth. The body was unavailable, so this summary is based on the title alone.
Notion explains Lore, an open-source shared memory system that stores agent experiences, tasks, decisions, procedures, and structured facts in interconnected Notion databases. The post details its MCP-based architecture, memory-maintenance strategies, and benchmark results showing when relevant memory improves agent performance—and when stale or redundant context can harm it.
Notion explains how it rebuilt its lexical search reindexer as a Spark-native pipeline, replacing ECS workers, Snowflake lookups, and live Elasticsearch writes with offline snapshot generation. The design uses embedded Elasticsearch, NVMe-backed executors, validation for document parity, and versioned alias-based catchup to achieve faster, safer, fully consistent reindexing.
Notion explains how it built the DLQ Explorer to query and recover failed background tasks stored in S3. The system uses Athena partition projection, controlled multi-region access, query filtering, batched re-enqueueing, and audit trails to make dead-letter queue investigations faster and safer.
Notion explains how it redesigned its data infrastructure to keep workspace data processed and stored within its home region. The post details regional data lakes, Kafka and Spark pipelines, AI embeddings, Airflow orchestration, event routing, sanitization, and Terraform-based maintainability.
Notion explains how it redesigned page spacing to create a more consistent rhythm across flexible web content. The team replaced baseline-alignment goals with standardized spacing and adjacency rules that keep list items compact while giving paragraphs more breathing room.
Notion explains how Custom Agents use least-privilege, resource-level permissions, runtime safeguards, prompt-injection mitigation, and human confirmation to secure collaborative AI workflows. It also shares lessons from testing more than 28,000 agents, including why Slack needed a constrained “read and reply” permission.
Justin Lee explains how Notion reduced Spark compute costs by 60–90% on Kubernetes using dynamic provisioning, bin packing, and AWS Spot Instances. The post details why large-scale interruptions caused Spark failures and how the open-source Spot Balancer distributes executors across spot and on-demand capacity to preserve reliability.
Notion’s security team built Scruff, an AI teammate that uses Custom Agents, MCP integrations, alerts, runbooks, and shared notes to automate security-alert triage and investigation. The case study details the workflow and reports 84% lower median investigation time and 93% faster resolution for false positives.