Skip to content

Williams et al.: “Mechanistic Interpretability Needs Philosophy”

Section titled “Williams et al.: “Mechanistic Interpretability Needs Philosophy””

Williams, Oldenburg, Fierro et al. (2025) argue that MI needs philosophy as an ongoing partner — for clarifying concepts, refining methods, and navigating epistemic complexity. They illustrate this through three open problems from the MI literature, each of which maps onto view-level confusions in our framework.

Problem 1: How should we decompose networks? (§2.1)

Section titled “Problem 1: How should we decompose networks? (§2.1)”

The paper challenges what they call the assumption of “the One True Decomposition” — that there is a privileged way to carve a neural network into interpretable parts. Drawing on philosophy of science, they argue for explanatory pluralism: different decompositions serve different explanatory goals, and no single decomposition has a unique claim to being “the real one.”

This is exactly the Decomposition Identity problem. Their philosophical framing maps onto our view analysis:

Their conceptOur viewThe connection
No single correct decompositionSubspace vs ObjectObject view assumes a unique decomposition; Subspace view says the subspace is real, the basis is a choice
Explanatory pluralismPerspectivalDifferent methods project different decompositions — this is expected, not a failure
Evaluate by causal understanding, not “true structure”RoleDecompositions should be judged by functional utility, not correspondence to some Platonic structure

Problem 2: What “features” do AI systems discover? (§2.2)

Section titled “Problem 2: What “features” do AI systems discover? (§2.2)”

The paper introduces the philosophical distinction between vehicles and content of representations. The vehicle is the physical carrier (a neuron, a direction in activation space). The content is what it represents (the concept “dog,” the property “is plural”).

This maps onto our Object/Role split applied to features:

Their conceptOur viewThe connection
Vehicle (the direction/neuron)ObjectThe feature is the component
Content (what it represents)RoleThe feature is the functional role it plays
What grounds representational content?SubspaceCausal subspace (DAS/IIA) provides grounding — interventions establish that the representation actually drives behavior

They also note that “feature” is used inconsistently in MI — sometimes meaning the vehicle, sometimes the content. This ambiguity is a source of confusion in debates about SAE “true” features and probe features.

Problem 3: How can we detect deceptive behaviour? (§2.3)

Section titled “Problem 3: How can we detect deceptive behaviour? (§2.3)”

The paper argues that deception detection requires philosophical clarity on what deception is. In philosophy, deception requires intentions and beliefs on the part of the deceiver — criteria that are controversial to attribute to language models. They also make a key distinction: not all information-concealing circuits are ethically equivalent. A circuit that hides dangerous capabilities, one that protects user privacy, and one that redirects to emergency care all “conceal information” but require completely different interventions.

This connects to our Deceptive Alignment analysis:

Their conceptOur viewThe connection
Deception requires intentions/beliefsStructuralYou need gauge-invariant evidence of a computation that implements strategic reasoning, not just a “deception direction”
Behavioral detection is insufficientInstrumental ceilingCompetent deception evades behavioral detection by definition
Different “deceptive” circuits need different responsesRoleThe functional role matters — same Object-level structure, different Role-level meaning

This paper arrives at many of the same conclusions as our framework but from philosophy of science rather than measurement theory. Where we say “different views define mechanisms differently,” they say “explanatory pluralism.” Where we say “Object view vs Subspace view,” they say “vehicle vs content.” Where we say “Instrumental evidence is structurally insufficient for deception detection,” they say “behavioral approaches cannot establish deception without philosophical clarity on what deception requires.”

The convergence from independent starting points strengthens both analyses.