Breakpoint

Loop engineering: why Claude's perf loop counted instructions, not milliseconds

In two weeks, more than 3,000 changes landed in claude.ai and the Claude desktop app with no customer-facing incidents or rollbacks, and the app got about…

anthropic··PT2M5.667S

video loads only when you press play

In two weeks, more than 3,000 changes landed in claude.ai and the Claude desktop app with no customer-facing incidents or rollbacks, and the app got about…

In two weeks, more than 3,000 changes landed in claude.ai and the Claude desktop app with no customer-facing incidents or rollbacks, and the app got about 3x faster (3.1x geometric mean across thirteen p75 timings, real user monitoring, August 13 vs August 27).

  • An agent loop is only as good as its check: wall-clock time was too noisy to gate CI, so the loop used deterministic counts such as CPU instructions.
  • Each count had to prove it tracked real latency before it was trusted, then became a CI ratchet that fails any PR raising it.
  • Engineers steered rather than prompted: every thread had a human owner, user-visible changes shipped behind flags, and over 3,000 changes landed without a rollback.

Claude shipped more than 3,000 changes to claude.ai in two weeks. The app got about three times faster, with no customer incidents and no rollbacks, and nobody prompted those changes one by one: the team built a loop and let it run. That's loop engineering: instead of prompting the agent, you build the system that prompts it, and a loop is only as good as the check that says whether a change worked. For performance, the obvious check is wall-clock time, but it's far too noisy to gate CI on. So they gave Claude deterministic counts instead, like CPU instructions under Valgrind and React commits, and made it prove each one tracked real latency. In the message tree code, every message ID was being resolved three times, and cutting that to one dropped instructions by 48% and wall-clock time by 78%. Then each count became a ratchet in CI, and any PR that raised it failed. With checks like that, the loop ran thread by thread in one Slack channel. Someone would post a recording of a slow spot, like sidebar rows shifting after the page was already usable, which the existing layout shift metric missed, so Claude wrote a test that failed on any shift, red on main and green on the fix. Then it shipped PRs behind a feature flag and read the real user metrics, ratcheting the benchmark down if latency improved, or turning the flag off if it didn't, and moved on to the next slow spot. Once any number could be climbed, the highest leverage work was finding more things to measure. At the peak, more than a hundred and fifty threads ran at once, and Claude started opening threads of its own from nightly jobs. The engineers weren't writing prompts anymore, they were steering the loop. When Claude said it would put a PR up later that week, one replied: if you put it up right now, I will get it merged and deployed, please be braver. Every thread had a human owner who approved anything users could see. The team's job was never to make the app faster, it was to keep handing Claude things to count.

This explainer is based on How we made claude.ai 3x faster in two weeks by Anthropic ↗. The original reporting and technical work belong to its publisher.