Does steering find mechanisms?
Section titled “Does steering find mechanisms?”Arditi et al. (2024) found a direction in activation space that, when ablated, removes refusal behavior from LLMs. The result is widely cited as “finding the refusal mechanism.” Zou et al. (2023) found directions for truthfulness, power-seeking, and other behavioral traits via representation engineering. Turner et al. (2023) demonstrated activation steering for personality traits. These results are being seriously considered for deployment-time safety monitoring: find the “deception direction,” monitor it during inference, flag when it activates.
The evidence in every case is Instrumental: intervening on the direction changes behavior. This establishes that the direction is a sufficient lever for behavioral control. It does not establish that the direction is the mechanism responsible for that behavior. The distinction matters enormously for safety.
A lever can work for reasons unrelated to the mechanism. The refusal direction might disrupt a shared computational pathway that many behaviors depend on — ablating it removes refusal because it removes everything, not because it specifically targets refusal. The truthfulness direction might correlate with the actual mechanism without being the mechanism itself — a confound rather than a cause. Redundant mechanisms are another failure mode: the model might have multiple pathways for refusal, and the direction captures only one. A safety monitor built on an Instrumental finding has Instrumental-level guarantees: it predicts and controls behavior on the distribution tested. It says nothing about whether the mechanism generalizes, whether it’s the only mechanism, or whether it’s robust to adversarial inputs designed to evade monitoring.
The view gap
Section titled “The view gap”Safety monitoring requires more than Instrumental evidence. At minimum, it requires Object evidence — the mechanism exists as an identifiable, necessary component. Ideally, it requires Subspace evidence — the mechanism is a stable subspace that generalizes across distributions and is robust to measurement. The gap between what the field has (Instrumental) and what safety applications need (Object or Subspace) is the central evidence gap in AI safety interpretability.
The Schmidt Sciences 2026 “Trustworthy AI” research agenda focuses on exactly this gap: detecting deceptive behaviors (sycophancy, harmful advice, selective omission) and steering for truthfulness with tools that generalize beyond academic benchmarks. Generalization is an external validity requirement that Instrumental evidence cannot satisfy.
Mechanistic validity impact
Section titled “Mechanistic validity impact”| Criterion | Instrumental evidence | Object evidence needed | Subspace evidence needed |
|---|---|---|---|
| I2 Sufficiency | Covered (steering works) | Covered (ablation shows necessity + sufficiency) | Covered |
| I1 Necessity | Not tested | Covered (ablation) | Covered |
| I5 Rival mechanism exclusion | Impossible | Possible | Covered (subspace uniqueness) |
| E4 Cross-model generalization | Not tested | Possible but not standard | Testable via subspace stability |
| E2 Prompt generalization | Not tested | Possible but not standard | Testable via Grassmannian distance |
| I4 Specificity | Not tested | Testable via targeted ablation | Testable via subspace orthogonality |
The pattern: Instrumental evidence covers sufficiency (the lever works) but leaves necessity, alternative exclusion, and all external validity criteria untested. A safety case built on Instrumental evidence has at least five untested validity criteria. Moving to Object evidence fills necessity and makes external criteria testable. Moving to Subspace evidence covers alternative exclusion and makes external criteria measurable.
Resolution status: Scoped
Section titled “Resolution status: Scoped”Prominent safety findings — refusal direction ablation, representation engineering for truthfulness — are Instrumental evidence: they establish that a direction is a sufficient behavioral lever, not that it is the mechanism. Safety monitoring requires at minimum Object evidence (necessity plus sufficiency) and ideally Subspace evidence (stability across distributions). The gap between instrumental evidence and the safety conclusions drawn from it is the central evidence deficit in the problem survey.
Sources
Section titled “Sources”- Schmidt Sciences (2026): “Trustworthy AI” agenda §2.1–2.2 — evaluation validity, mechanistic interventions, deception detection
- Sharkey et al. (2025) §3.2: Safety monitoring and auditing (arXiv:2501.16496)
- Apollo Research (2024): Deception evaluation framework (45+ MI Projects)
- Arditi et al. (2024): Refusal in language models is mediated by a single direction
- Zou et al. (2023): Representation engineering: a top-down approach to AI transparency
- Turner et al. (2023): Activation addition: steering language models without optimization