An AI nudged toward pears can start putting them in unrelated fairy tales. Activation steering changes internal representations, but an intervention…
// the short version
An AI nudged toward pears can start putting them in unrelated fairy tales. Activation steering changes internal representations, but an intervention applied everywhere can disrupt answers that were already fine.
// what to take away
- A steering intervention applied everywhere can alter unrelated answers as well as its intended targets.
- A learned gate scales the intervention per token and layer, separating the correction from the decision to apply it.
- The experiments improve toxicity/capability tradeoffs, but gating adds overhead and depends on examples that distinguish when intervention is needed.
// transcript
Ask an AI to name a fruit. It might say apple. Now imagine nudging it toward pear, without changing your question. Researchers do that by adjusting the internal numbers the model uses while choosing its next words, making pear more likely. That's activation steering. The model's original weights stay the same; the nudge happens while it's producing an answer. But leave that nudge on when you ask for a fairy tale, and pears can creep into the story too. The same problem matters when steering is meant to reduce toxic language: an intervention that helps with harmful content can also disrupt perfectly harmless answers. Researchers at Apple and the Autonomous University of Barcelona added a learned dimmer to that nudge. In our fruit example, it would turn up for a fruit question and down for an unrelated story. It makes that decision for each piece of text at each affected layer, adjusting how much of the nudge gets through. That dimmer learns from examples of what should change and what should be left alone. The steering method learns the correction, while the gate learns when it's needed. That separation improved the balance between reducing toxicity and preserving language ability across three tested models. On the smaller Qwen model, adding it to one steering method brought the toxicity score from roughly ten percent to four percent. Both settings had to stay within the same limits on lost capability. The same idea also worked in image generation. They trained a model to blur pictures involving bananas. Ordinary steering blurred unrelated pictures too, while the learned gate largely preserved dogs, waterfalls and cars. But the blur still spread beyond the banana itself. That extra decision isn't free. In their detailed timing tests, generating text took roughly two to six percent more time than steering alone. And the gate depends on its examples: if it can't distinguish what needs changing, it loses that selective advantage. So this is a better tradeoff in these experiments, not a guarantee that a model will behave. But it changes what control means: knowing how to correct an answer is only half the job. The other half is knowing when to leave it alone.
// source
This explainer is based on Dynamically Scaled Activation Steering by Apple ↗. The original reporting and technical work belong to its publisher.