Breakpoint

How OpenAI's Habitat storage scaled past connection pressure

OpenAI's storage demand grew more than tenfold annually for three years. This explainer follows Habitat's event-loop stalls, connection-pool feedback…

openai··PT1M33S

video loads only when you press play

OpenAI's storage demand grew more than tenfold annually for three years. This explainer follows Habitat's event-loop stalls, connection-pool feedback…

OpenAI's storage demand grew more than tenfold annually for three years. This explainer follows Habitat's event-loop stalls, connection-pool feedback loop, and HTTP/2 connection fan-in.

  • Rapid storage growth exposed event-loop stalls and feedback loops in connection management.
  • Reusing connections removed repeated setup work and reduced downstream pressure.
  • HTTP/2 multiplexing let many request streams share a much smaller connection pool.

OpenAI's storage demand grew more than tenfold, year after year. Habitat had to keep ChatGPT's data accessible through three years of that growth. It became a central service because library updates required coordinating dozens of deployments. They kept Python, but asyncio couldn't parallelize CPU work. Compression, encryption and response processing competed on one event loop, leaving completed database responses waiting. They measured event-loop scheduling delay directly to expose that hidden waiting. Even feature flags caused stalls: workers parsed a huge configuration on synchronized refreshes. Smaller configs, longer intervals and randomized refresh timing reduced that interference. They also capped concurrency per process and added workers, but traffic wasn't evenly distributed. The connection pool reused the most recently returned connection. Overloaded servers returned connections later, so they kept attracting new requests. Switching from last-in-first-out to first-in-first-out reuse broke that feedback loop. More workers also multiplied database connections. Envoy pooled them and multiplexed requests over HTTP two, reducing downstream connection pressure. Habitat also restricted its API to predictable operations, isolating complex queries in separate analytical systems. The platform now serves over seventy million requests a second. Scaling meant controlling how much work arrived, where it landed, and what each request could cost.

This explainer is based on Rapidly scaling online storage to serve over 1 billion ChatGPT users by OpenAI ↗. The original reporting and technical work belong to its publisher.