Engineering, explained visually
Short, narrated explanations of the systems behind engineering write-ups. Each page includes the video, the mechanism in plain language, a full transcript, and the original source.

Adversarial distillation: how an LLM's encrypted reasoning was replayed back out
Distillation trains a small "student" model on the outputs of a bigger "teacher". Labs do it to their own models to make the small, cheap versions. It…

asyncio hides who called your slow query. Datadog's profiler fix: memcpy and a SIGSEGV handler
A sampling profiler takes a snapshot of each thread's call stack every few milliseconds, and the function on top of the most snapshots is where the CPU…

Kernel panic logs die with the machine. Jane Street's fix: pstore, /dev/kmsg, mlockall
Jane Street stores kernel logs from every Linux host centrally, but the logs that matter most come from a failing disk or a kernel panic, which take out…

Loop engineering: why Claude's perf loop counted instructions, not milliseconds
In two weeks, more than 3,000 changes landed in claude.ai and the Claude desktop app with no customer-facing incidents or rollbacks, and the app got about…

Why your app waits seconds for an LLM to answer a yes/no question
Your code needs three typed answers about a support ticket: is the customer angry, do they get a refund, which team takes it. A chat model writes a reply…

Why activation steering can change the wrong AI answers
An AI nudged toward pears can start putting them in unrelated fairy tales. Activation steering changes internal representations, but an intervention…

Why Cloudflare's consistent hash rings ate so much RAM
Cloudflare's Pingora Backend Router accumulated hash points across weighted servers and feature-specific rings. A simple routing algorithm had become a…

Why pressing Enter makes Notion's CRDT harder
Two people edit the same sentence. One presses Enter. Where should the other's new word go?

Dario Amodei's proposal to pace the AI frontier
Dario Amodei of Anthropic argues that keeping AI safe means giving safety work time to catch up with rapidly improving models.

How OpenAI's Habitat storage scaled past connection pressure
OpenAI's storage demand grew more than tenfold annually for three years. This explainer follows Habitat's event-loop stalls, connection-pool feedback…

Why Pinterest's vector search needs so much RAM
Billions of embeddings make vector search a memory problem. Pinterest's Manas retrieval platform tackles it by changing how vectors are represented and…

Why 10% more C++ code made Figma builds 50% slower
Figma's C++ codebase grew 10% over twelve months. Build times grew 50%. Growth

How Stripe uses graph search to repair database fleets
A replicated database keeps each shard on several machines and lets exactly one

Why rolling dashboards miss the cache at Netflix scale
A dashboard showing "the last 3 hours" refreshes every 10 seconds. Almost all of

How Google automates geospatial outbreak prediction
Building a geospatial prediction model for a crisis takes specialist teams weeks: finding covariates across portals, cleaning them, guarding against…

How Cloudflare removed 100 TB from its DNS cache
Cloudflare's public DNS resolver caches 250 billion answers at once, so every byte a cache entry wastes costs more than 250 GB of memory across the fleet.…

Inside GitHub's August 17 outage and retry storm
On August 17, 2026, GitHub was down for 7 hours 47 minutes: github.com, authentication, Actions, APIs, pull requests, issues and Copilot, with roughly 20%…
How Datadog hides share links inside dashboard screenshots
Datadog renders about a billion dashboard widgets a day, and people screenshot

How Shopify moved inventory reservations from Redis to MySQL
When a buyer clicks "Complete purchase", Shopify has to guarantee the item is