Over the past few years, we’ve been on a journey to modernise how we run Amazon Elastic Compute Cloud (EC2) instances at Slack. The post introduces Shipyard as Slack’s next-generation EC2 platform and frames it as the next step after improving the company’s Chef infrastructure and promotion workflows.
Slack engineering blog
Abstract Agent-driven end-to-end (E2E) tests add a new exploratory layer to testing, but should they replace traditional deterministic tests? We ran more than 200 agentic E2E workflows using the Playwright MCP, Playwright CLI, and agent-generated Playwright tests in test workspaces using non-production data to find out how agentic testing could fit into both our and…
In early 2023, Slack faced a foundational challenge: serving Large Language Models (LLMs) at enterprise scale with the security, reliability, and performance customers expect. Over three years, the team evolved from basic infrastructure to a sophisticated multi-cloud architecture designed to withstand regional outages and support enterprise AI workloads.
By 2024, Slack’s data platform had accumulated 700+ SSH-based operators orchestrating critical data pipelines, including daily search indexing and analytics jobs. The post describes how Slack modernized these workflows by replacing direct SSH access to production EMR clusters with a more secure REST-based approach.
Excerpt In complex, long-running agentic systems, maintaining alignment and coherent reasoning between agents requires careful design. In this second article of our series, we explore these challenges and the mechanisms we built to keep teams of agents working productively over long time spans. We present a range of complementary techniques that balance the conflicting requirements…
Slack describes limitations of legacy hybrid network measurement tooling that combined commercial SaaS and custom-built solutions. The post outlines a move toward a scalable, open approach using Prometheus to improve network probing and ensure HTTP/3 readiness across their infrastructure.
Slack rebuilt its notifications system to reduce overwhelm and improve clarity for teams. The article describes redesigning a complex system from the ground up to bring calm, consistency, and make staying informed feel effortless.
Slack’s Security Engineering team processes billions of security events per day and reviews alerts generated by their detection systems during on-call shifts. This post explains how they use AI-driven agents to streamline investigation workflows and reduce manual effort. It covers the architecture of their ingestion pipeline and how agents augment analyst efficiency.
This post describes Slack’s process for completing a Voluntary Product Accessibility Template (VPAT) for their Android app, outlining how the assessment informs customers about accessibility features. It explains the third-party evaluation process and lessons learned while aligning the product with accessibility standards. The article highlights practical steps Slack took to document and improve accessibility compliance.
Slack’s build pipeline team faced slow backend build times (around 60 minutes) which hindered developer feedback loops. This post covers the initiatives they pursued to speed up builds and improve the developer experience for Quip and Slack Canvas’s backend. It discusses tooling and process changes aimed at delivering faster, more reliable builds so engineers can iterate more quickly.
This post revisits Slack’s evolution of Chef infrastructure, detailing the migration from a single Chef stack to a multi-stack model and the operational challenges that followed. It explains improvements made to cookbook upload handling and other processes to increase safety without disrupting services. The article focuses on maintaining stability while modernizing configuration management practices.
The Deploy Safety Program at Slack aimed to reduce customer impact hours from deployments by improving deployment practices and safety culture. Over eighteen months the program significantly reduced customer-impacting incidents through process and tooling changes. The post outlines the program’s approach, metrics, and how teams collaborated to make deployments safer.
As attacks grow more sophisticated and faster, Slack built an anomaly event response system to close the gap between detection and remediation. The post discusses moving beyond reactive approaches to reduce time-to-response and prioritize containment and investigation. It covers the design principles and components used to automate and streamline anomaly handling.
Slack’s DevXP team improved an end-to-end (E2E) testing pipeline by reusing and optimizing existing tools to reduce build times and eliminate redundant steps. The post explains the practical changes they made to speed up tests and streamline developer workflows. The improvements resulted in faster feedback loops and more efficient engineering processes.
This article describes how Slack built enterprise search capabilities with a focus on security and user privacy while leveraging Slack AI. It outlines the architecture and privacy-preserving measures used to surface insights securely for customers. The post covers trade-offs and engineering choices made to balance powerful search features with strict privacy controls.