Breakpoint

Slack engineering blog

Shipyard: How We Built Slack’s Next-Generation EC2 Platform
Over the past few years, we’ve been on a journey to modernise how we run Amazon Elastic Compute Cloud (EC2) instances at Slack. The post introduces Shipyard as Slack’s next-generation EC2 platform and frames it as the next step after improving the company’s Chef infrastructure and promotion workflows.
Agentic Testing: Where Agents Fit in the E2E Testing Stack
Abstract Agent-driven end-to-end (E2E) tests add a new exploratory layer to testing, but should they replace traditional deterministic tests? We ran more than 200 agentic E2E workflows using the Playwright MCP, Playwright CLI, and agent-generated Playwright tests in test workspaces using non-production data to find out how agentic testing could fit into both our and…
Slack AI: The Path to Multi-Cloud
In early 2023, Slack faced a foundational challenge: serving Large Language Models (LLMs) at enterprise scale with the security, reliability, and performance customers expect. Over three years, the team evolved from basic infrastructure to a sophisticated multi-cloud architecture designed to withstand regional outages and support enterprise AI workloads.
Managing context in long-run agentic applications
Excerpt In complex, long-running agentic systems, maintaining alignment and coherent reasoning between agents requires careful design. In this second article of our series, we explore these challenges and the mechanisms we built to keep teams of agents working productively over long time spans. We present a range of complementary techniques that balance the conflicting requirements…
How Slack Rebuilt Notifications 📣
Slack rebuilt its notifications system to reduce overwhelm and improve clarity for teams. The article describes redesigning a complex system from the ground up to bring calm, consistency, and make staying informed feel effortless.
Streamlining Security Investigations with Agents
Slack’s Security Engineering team processes billions of security events per day and reviews alerts generated by their detection systems during on-call shifts. This post explains how they use AI-driven agents to streamline investigation workflows and reduce manual effort. It covers the architecture of their ingestion pipeline and how agents augment analyst efficiency.
Android VPAT journey
This post describes Slack’s process for completing a Voluntary Product Accessibility Template (VPAT) for their Android app, outlining how the assessment informs customers about accessibility features. It explains the third-party evaluation process and lessons learned while aligning the product with accessibility standards. The article highlights practical steps Slack took to document and improve accessibility compliance.
Build better software to build software better
Slack’s build pipeline team faced slow backend build times (around 60 minutes) which hindered developer feedback loops. This post covers the initiatives they pursued to speed up builds and improve the developer experience for Quip and Slack Canvas’s backend. It discusses tooling and process changes aimed at delivering faster, more reliable builds so engineers can iterate more quickly.
Advancing Our Chef Infrastructure: Safety Without Disruption
This post revisits Slack’s evolution of Chef infrastructure, detailing the migration from a single Chef stack to a multi-stack model and the operational challenges that followed. It explains improvements made to cookbook upload handling and other processes to increase safety without disrupting services. The article focuses on maintaining stability while modernizing configuration management practices.
Deploy Safety: Reducing customer impact from change
The Deploy Safety Program at Slack aimed to reduce customer impact hours from deployments by improving deployment practices and safety culture. Over eighteen months the program significantly reduced customer-impacting incidents through process and tooling changes. The post outlines the program’s approach, metrics, and how teams collaborated to make deployments safer.
Building Slack’s Anomaly Event Response
As attacks grow more sophisticated and faster, Slack built an anomaly event response system to close the gap between detection and remediation. The post discusses moving beyond reactive approaches to reduce time-to-response and prioritize containment and investigation. It covers the design principles and components used to automate and streamline anomaly handling.
Optimizing Our E2E Pipeline
Slack’s DevXP team improved an end-to-end (E2E) testing pipeline by reusing and optimizing existing tools to reduce build times and eliminate redundant steps. The post explains the practical changes they made to speed up tests and streamline developer workflows. The improvements resulted in faster feedback loops and more efficient engineering processes.
How we built enterprise search to be secure and private
This article describes how Slack built enterprise search capabilities with a focus on security and user privacy while leveraging Slack AI. It outlines the architecture and privacy-preserving measures used to surface insights securely for customers. The post covers trade-offs and engineering choices made to balance powerful search features with strict privacy controls.