Breakpoint

Adversarial distillation: how an LLM's encrypted reasoning was replayed back out

Distillation trains a small "student" model on the outputs of a bigger "teacher". Labs do it to their own models to make the small, cheap versions. It…

openai··PT2M13.933S

video loads only when you press play

Distillation trains a small "student" model on the outputs of a bigger "teacher". Labs do it to their own models to make the small, cheap versions. It…

Distillation trains a small "student" model on the outputs of a bigger "teacher". Labs do it to their own models to make the small, cheap versions. It becomes adversarial distillation when the teacher is somebody else's model and its terms of service forbid it.

  • A distillation campaign trains a student model on another lab's outputs at scale: thousands of fraudulent accounts behind proxy resellers, because one account cannot collect the millions of examples a student needs.
  • A model's reasoning is the most valuable thing to collect, so providers hide it and hand the client an encrypted block to send back with the next request, which keeps the API stateless.
  • OpenAI says operators replayed that block into a fresh conversation and had the model transcribe it, with no encryption broken. An encrypted token is only as secret as what the server will decrypt it for, so it has to be bound to its user, conversation and model.

In July, thousands of accounts tried to copy an AI's hidden reasoning. Nobody broke any encryption or touched a database, they just asked for it in the right way, and it's the clearest look yet at how a distillation campaign works. Distillation means training a small student model on the outputs of a bigger teacher, and labs do it to their own models all the time, which is where the cheap ones come from. It turns adversarial when the teacher is somebody else's, and its terms of service forbid it. A student needs millions of examples, so a campaign spreads the work across thousands of fake accounts, usually through proxy services that resell access, so banning one account just brings up another. Anthropic reported three campaigns in February, and counted 16 million exchanges across about 24,000 accounts. What a campaign wants most is the reasoning a model does before it answers, because the answer only shows where it ended up, while the reasoning shows every step, and that's far richer training data. So the labs hide it, and the API hands your client the reasoning as an encrypted blob, which your client sends back with the next request, so the server never has to store it. The catch is that the blob wasn't tied to the conversation it came from, so the operators pasted one into a fresh conversation and asked the model there to write out what it said, and since the server decrypts the blob for whichever model receives it, the model simply read it back. Any one of those requests looks ordinary, so what gave the campaign away was the pattern, 16,000 requests with the same trick from over 4,000 users in two days, which led OpenAI to a cluster of 15,000. OpenAI ties a core group to people associated with Moonshot AI, the lab behind Kimi, though it can't say they were all one actor, and it's counting attempts, not successes. The fix stops a blob being replayed by another user or another model family, and holds back streamed output that looks like leaked reasoning. It's a lesson older than AI, that encrypting what you hand a client only keeps it secret until your own server decrypts it for whoever sends it back. Nobody picked the lock, they just asked the one holding the key to read it out loud.

This explainer is based on Disrupting a coordinated model-distillation campaign by OpenAI ↗. The original reporting and technical work belong to its publisher.