Breakpoint

Kernel panic logs die with the machine. Jane Street's fix: pstore, /dev/kmsg, mlockall

Jane Street stores kernel logs from every Linux host centrally, but the logs that matter most come from a failing disk or a kernel panic, which take out…

jane street··PT1M58.867S

video loads only when you press play

Jane Street stores kernel logs from every Linux host centrally, but the logs that matter most come from a failing disk or a kernel panic, which take out…

Jane Street stores kernel logs from every Linux host centrally, but the logs that matter most come from a failing disk or a kernel panic, which take out the storage and the network those logs travel through. Workstations have no lights-out management console to fall back on. Linux Engineering intern Jacob Root built a log shipper that survives both.

  • A failure reporter cannot depend on what is failing: the shipper pins itself in RAM with mlockall, reads /dev/kmsg directly, caches DNS results in a background loop and sends plain UDP, so a dead disk cannot stall it.
  • After a kernel panic there is no storage or network stack, so Linux's pstore writes the end of the log, compressed and newest first, into the firmware's non-volatile memory, and the shipper sends it on the next boot.
  • The saved lines traced random desktop freezes to a NULL pointer dereference in the Nvidia driver, triggered by Chrome's crash reporter, on machines with no other way to keep a log.

A dying Linux machine often can't save the logs you want most. Jane Street ships kernel logs off every Linux host it runs, but a failing disk or a kernel panic takes out the very path those logs leave by. So an intern, Jacob Root, spent a summer building a shipper that keeps working while the machine is dying. The trouble is that an ordinary program leans on the disk in places you never see. Linux loads a program's code from disk a page at a time, so the shipper starts by calling mlockall, which pins all of itself in memory. The usual log reader is another binary on disk, so it reads the kernel's log buffer directly, one whole line per read. Finding the collection servers takes DNS, which reads its config from disk, so a background loop caches the addresses, and the send loop never waits on it. HTTPS would mean TLS certificates, which sit on disk too, so every line goes out as plain UDP, one copy to each collection server. A kernel panic is harder, because it leaves no storage stack and no network stack to call. What the kernel can still reach is pstore, a Linux subsystem that writes into the small non-volatile memory where the firmware lives. There's very little room, so the kernel compresses the log and writes it newest first until the space runs out, which keeps the lines closest to the crash. Two boot arguments make the kernel write there before it attempts a full crash dump, and on the next boot the shipper reassembles the fragments and sends them off. It paid off the week Jane Street's desktops kept freezing at random. The surviving lines showed Chrome's crash reporter walking a crashed process's memory into a region mapped by the Nvidia driver, where the driver dereferenced a null pointer. That pointed at a recent driver update, so they rolled it back and took the trace to Nvidia. Anything that reports a failure has to outlive it, and most of the work is finding where it quietly wouldn't.

This explainer is based on A resilient kernel log reporter by Jane Street ↗. The original reporting and technical work belong to its publisher.