Case Study — Refusal Direction
Section titled “Case Study — Refusal Direction”One evidence base, three readings, and only the first of them holds. Arditi et al. (2024) ablate a single difference-in-means direction and refusal stops across 13 chat models; adding the direction back induces refusal on harmless instructions. Nobody disputes the result. What the result establishes depends entirely on which view the claim is stated under, and the inference from this vector suppresses refusal to this is the refusal mechanism crosses a view boundary without saying so.
Instrumental view — the reading that holds
Section titled “Instrumental view — the reading that holds”Under the instrumental view a mechanism is a predictive model and identity is predictive equivalence. The claim is that the direction is a lever that controls refusal, and the evidence establishes exactly that: intervene on one direction, refusal behavior changes, across thirteen models and in both directions. This is a strong instrumental result and the realism criterion has nothing to add to it.
Subspace view — the reading the evidence does not reach
Section titled “Subspace view — the reading the evidence does not reach”Under the subspace view a mechanism is a causal subspace and identity is the projector. The claim would be that refusal is organized along one dimension of the residual stream. The same evidence does not carry it. An intervention licenses a lever rather than a representational organization; no rank above one is ever ablated, so nothing rules out a higher-dimensional organization of which this direction is one coordinate; and the search enumerates single vectors only, so the result is the best one-dimensional answer to a question that was never asked in higher dimensions.
Contrastive view — a different failure
Section titled “Contrastive view — a different failure”Under the contrastive view a mechanism is defined relative to a foil, and the foil is constitutive rather than a methodological convenience. Here the foil is visible in the estimator: the direction is computed as the axis along which mean harmful and mean harmless activations differ. A harmfulness direction satisfies that estimator exactly, so harmfulness and refusal are separated by citation rather than by experiment. The confirming observation is that base models which never refuse express the same direction — which is what a contrast-defined referent looks like from outside.
Analysis
Section titled “Analysis”| View | The claim | The verdict on this evidence |
|---|---|---|
| Instrumental | The direction is a lever that controls refusal | Holds |
| Subspace | Refusal is organized along one residual-stream dimension | Does not carry — an intervention licenses a lever, not an organization |
| Contrastive | The direction is the refusal mechanism | Does not carry — the estimator is the harmful-minus-harmless axis |
Neither failing reading is a failure of the experiment. Within a row the evidence is one body of work, so the verdict changes with the conception of mechanism rather than with the amount of evidence: what fails is the transfer of one view’s evidence to another view’s claim. The repair is a declaration, not a further experiment — state which of the three propositions the paper asserts.
Further reading
Section titled “Further reading”Arditi, A., Obeso, O., Syed, A., Paleka, D., Panickssery, N., Gurnee, W., Nanda, N. “Refusal in Language Models Is Mediated by a Single Direction.” arXiv:2406.11717, 2024.
See also Probing Classifiers, which splits the same way and harder, and the Safety evidence gaps open problem.