Skip to content

Mechanistic interpretability has produced circuits, features, and causal subspaces — each resting on an implicit answer to a prior question: what kind of thing is a mechanism? In some papers a mechanism is a set of attention heads; in others it is a named functional role such as “name mover”; in others it is a causally sufficient subspace recovered by DAS; in others it is a set of features recovered by sparse autoencoders; in others it is a training-time transition. The difference between these is ontological, and ontological differences propagate: what a mechanism is determines when two descriptions pick out the same one, what evidence can warrant a claim, and what inferences may be drawn from it. A mechanistic claim has no determinate truth-conditions until these commitments are declared.

The consequences are practical. Wang et al. establish by path patching that head 9.9 carries a large direct effect on the indirect-object logit, then label it a “name mover head” — a functional role. These are different claims under different views: the ablation supports an object claim (this component matters), while the role label requires independent operationalization. Arditi et al. show that removing a “refusal direction” suppresses refusal across 13 chat models; read instrumentally the finding is complete, but read as a claim about how refusal is organized in the residual stream it is not — two pooling choices applied to the same activations at the same layer recover directions 73° apart. Most published work answers these questions implicitly, and disagreements that are fundamentally ontological are argued as disagreements about experimental design.

We call the background assumptions behind a mechanistic claim a mechanistic view.

Five axes of a mechanistic view

A mechanistic view must answer five questions. The questions fall into distinct categories, which matter for how the answers relate to each other.

Definitional axes — what the mechanism is and when two are the same:

  1. Ontology. What kind of thing is a mechanism? Options include: a concrete component (head, neuron, feature), a functional role that can be multiply realized (Putnam 1967), a causal subspace in residual stream geometry, an invariant relational structure, a temporally extended process, or a combination of these at multiple strata. The concept of mechanism ontology in science follows Machamer, Darden & Craver (2000).

  2. Identity. When are two mechanism descriptions referring to the same mechanism? This is a form of the classical identity problem in ontology (Quine 1948): no entity without identity criteria. This matters for cross-model comparison, for deciding when two competing circuit papers are contradicting each other versus measuring different things, and for tracking a mechanism across training checkpoints. The identity criterion must fit the ontology: the right criterion for a role (functional equivalence) differs from the right criterion for a subspace (geodesic distance on a Grassmannian). When two weight configurations compute the same function, the identity question becomes one of gauge equivalence (Earman 2004).

Evidential axis — what justifies the claim:

  1. Evidence. What measurements would warrant a mechanism claim of this type? Different ontologies license different evidence. A component claim is most directly supported by ablation and patching. A subspace claim needs subspace recovery and causal tests that jointly converge. A structural claim needs measurements robust to reparameterization. Evidence and ontology are linked: an evidence type that does not track the claimed kind of object cannot confirm the claim.

    Three models of evidence are in play in the literature, corresponding to different epistemological traditions:

    • Hypothetico-deductive (HD): a mechanistic hypothesis is supported when its predictions are confirmed (Popper 1959; Hempel 1966). Activation patching is used HD-style when a researcher predicts “ablating head 9.9 should reduce accuracy on IOI” and observes the reduction.
    • Bayesian: evidence updates the probability of a mechanism hypothesis (Howson & Urbach 2006). This is the implicit model behind most “evidence accumulation” discourse: multiple independent tests move credence up. The triangulation requirement reflects the Bayesian insight that independent evidence is worth more than correlated evidence.
    • Eliminative/triangulation: evidence from structurally different domains eliminates distinct alternative hypotheses (Campbell & Fiske 1959; Woodward 2003). A weight-space result rules out one class of confounders; an activation-space result rules out a different class; convergence establishes what survives both eliminations.

    The Mechanistic Validity tier ladder is most naturally read as tracking eliminative strength: Tier 1 (Proposed) is how-possibly, Tier 4 (Triangulated) is how-actually. This site’s own verdict vocabulary is separate: view-invariant, contested, and candidate.

Representational axis — how the claim is expressed:

  1. Formalism. What mathematical language is used to represent the mechanism? A circuit diagram, a point on a Grassmannian, and a gauge orbit in a fiber bundle are different representational choices. The choice of formalism determines what can be stated precisely and what can be proved. Formalism is not itself a stance on what exists — two researchers can hold the same ontological view and express it in different formalisms — but the formalism constrains what the ontological view can commit to.

Scope axis — what the claim is about:

  1. Target. What phenomenon is the mechanism supposed to explain? A mechanism of indirect object identification, of in-context learning, of grokking, and of factual recall carry different target-level commitments. The target determines the level of description at which the explanation must be pitched (corresponding to what Craver calls “constitutive relevance” in the philosophical literature on mechanistic explanation), and therefore which evidence types are relevant and what counts as explanatory success.

The five answers are not independent. What a mechanism is determines when two are the same, which determines what formalism is needed:

Ontology    Identity    Formalism\text{Ontology} \;\longrightarrow\; \text{Identity} \;\longrightarrow\; \text{Formalism}

Under the object view, mechanisms are specific components, identity is component overlap, and the formalism is a directed graph. Under the role view, mechanisms are functional roles, identity is role equivalence, and the formalism is a functional decomposition. Possible mismatches arise between these three axes: using subspace math while holding a component-overlap notion of identity, or using patching evidence (object-level) to support a structural-level claim. These mismatches are a common source of incoherence in practice.

A view is a bundle of answers to the five axes. Formally, a mechanistic view is a 5-tuple:

σ=(O,,E,F,T)\sigma = (O, {\sim}, E, F, T)

where OO is the ontological type (the set of entities that count as mechanisms), O×O{\sim} \subseteq O \times O is an equivalence relation (the identity criterion: xyx \sim y means descriptions xx and yy refer to the same mechanism), EE is a set of evidential standards, FF is the formalism, and TT is the target phenomenon.

Coherence conditions. A view σ\sigma is coherent if and only if:

  1. Identity is grounded: {\sim} is definable using only the objects in OO and the language of FF. A view whose identity criterion involves objects not in its own ontology is incoherent.
  2. Evidence tracks identity: for every measurement type mEm \in E, the claims mm can support are expressible in terms of {\sim}-classes, not finer-grained distinctions. Evidence that distinguishes objects within a {\sim}-class is picking up on a different ontology than the one declared.
  3. Formalism is expressive: every element of OO has a representation in FF, and {\sim} is computable in FF.

Incoherent views produce contradictory demands. If a view violates condition (2) — there exists a measurement mEm \in E and objects x,yOx, y \in O with xyx \sim y but m(x)m(y)m(x) \neq m(y) — then the view simultaneously asserts that xx and yy are the same mechanism (by {\sim}) and that they are different (by mm). No experiment can satisfy both demands. This is immediate from the definitions.

A stronger consequence: incoherent views produce systematic underdetermination. Each new measurement either agrees with the identity criterion (offering no information about the discrepancy) or disagrees (introducing its own underdetermination). This may be one reason why debates between activation-patching advocates and DAS advocates are difficult to settle by more of the same evidence when the participants hold different implicit views.

Examples. Claiming O=O = component but ={\sim} = geodesic distance on Gr(k,d)\mathrm{Gr}(k,d) violates condition (1): geodesic distance requires subspace objects, not component objects. Citing activation patching as EE for a gauge-orbit identity claim violates condition (2): activation patching is head-index-dependent and therefore not sensitive to gauge-orbit membership.

The subspace view bundles:

  • OO: mechanism as causal subspace
  • {\sim}: same projector, or geodesic distance on Gr(k,d)\mathrm{Gr}(k, d) below threshold
  • EE: convergence of DAS-recovered and SVD-weight subspaces across seeds and architectures
  • FF: Grassmannian geometry, transport maps, G-SCM
  • TT: causal variables (e.g., the indirect object token in IOI)

The axes are questions. The views are answers.

  1. The same axis can take different values in different views.
  2. A view changes multiple axes at once in a coherent way. Switching from the object view to the subspace view is not just changing the ontology; it changes the identity criterion, the evidence standard, and the appropriate formalism together.
  3. Axes can be used as an audit tool. For any paper making a mechanistic claim, you can ask: which axis values is this paper assuming? Are those assumptions consistent? Are they stated?

The nine named views represent coherent combinations of axis values. But there is no requirement to adopt a named view wholesale — you can take the subspace ontology with the structural identity criterion, or the object ontology with process-view evidence standards. Such combinations are valid as long as they remain coherent.

The coherence check: does the identity criterion fit the ontology? Does the evidence track the identity criterion? If you use a subspace ontology but need to compare mechanisms across models, you need an identity criterion that works for subspaces — geodesic distance or gauge-orbit membership, not component overlap. Mixing is fine; incoherence is not.

In practice, most interpretability papers mix axes implicitly. The axis framework makes the mixing explicit so it can be evaluated.

Disagreements in mechanistic interpretability are often view-level disagreements. Two researchers looking at the same activation-patching result may disagree about what it shows, not because of different data, but because they hold different background views about what mechanism the data is evidence for.

Making the view explicit does not automatically resolve the disagreement. But it locates the disagreement precisely, which makes it possible to design experiments that distinguish between views.

A persistent problem is underdetermination: multiple distinct mechanism hypotheses can be consistent with the same evidence. An activation-patching result that shows head 9.9 is necessary is consistent with (a) head 9.9 being the mechanism, (b) head 9.9 being one realizer of a role-level mechanism, (c) head 9.9 participating in a distributed mechanism whose activity happens to be concentrated there, and (d) head 9.9 being a measurement artifact of the patching protocol. These are not equivalent claims under different descriptions; they support different predictions and require different interventions.

The realism criterion — treating a structure as view-invariant only when its support crosses at least two of the five view families, partitioned by primary failure mode — is a response to underdetermination. The single-domain evidence produces an underdetermined verdict; multi-domain convergence narrows the hypothesis space. This is a pragmatic application of the principle that independent evidence is stronger than repeated evidence from the same source.

Cross-model comparison. The claim that two models implement “the same” mechanism requires a notion of identity that works across models. Component indices (head 9.9 in model A, head 7.3 in model B) are trivially not cross-model. Functional roles may be comparable, if independently specified. Geometric structure may be, if a canonical mapping is defined. Which notion is assumed determines whether “the same mechanism” is meaningful across architectures.

Generalization. Finding a circuit invites a natural question: does its presence guarantee specific behavior on inputs the model has never seen? Under an object view, circuit presence is about specific components on a specific distribution — it says nothing about novel inputs. Under a subspace or higher view, the circuit is defined by structure that is increasingly robust: a subspace is stable under perturbation, a gauge-invariant structure survives reparameterization. The stronger the ontological commitment, the broader the class of situations to which the claim applies — if warranted by evidence.

Post-hoc explanation. In circuit discovery, components are identified first and functional roles assigned afterward — labels like “name mover” describe observed behavior, not independent predictions. Any observed behavior can be given a plausible role label, making the functional story difficult to falsify. The backup name movers in the IOI circuit (Wang et al. 2022) illustrate this: knock out the “name movers” and other heads take on the same role. A role claim becomes falsifiable when the role is predicted from independent evidence — structure, training dynamics, or cross-model transfer — rather than read off the same activations that identified the component.

Marr’s tri-level account — computational, algorithmic, implementational — organizes explanatory targets: what a system does, how it does it, what substrate realizes the process. The five axes here organize ontological commitments: what kind of object the mechanism is, when two are the same, what evidence can support the claim. These are orthogonal. Two researchers can agree that IOI is an algorithmic-level finding — a procedure involving duplicate token heads, S-inhibition heads, and name movers — and still disagree about two things: what a circuit is as a formal object, and what underlying mechanism the circuit describes.

These correspond to the nine mechanistic views — each gives a different answer.

Two senses of “mechanism.” The word “mechanism” is used here in the sense established in the interpretability literature: a computational structure inside a neural network responsible for a specific behavior. In the philosophy of science, “mechanism” carries more structure. The influential Machamer-Darden-Craver (MDC) account defines mechanisms as entities and activities organized to produce a phenomenon, with specific start conditions, end conditions, and temporal order (Machamer, Darden & Craver 2000). Craver further distinguishes how-possibly models from how-actually models, with evidence requirements increasing accordingly — a distinction that maps onto the Mechanistic Validity tier ladder (Tier 1 Proposed = how-possibly, Tier 4 Triangulated = how-actually). The interpretability sense diverges from MDC in two ways: no commitment to spatial structure, and an emphasis on Woodward’s manipulationist causation (Woodward 2003) over organized spatial layout. A related concept is constitutive relevance (Craver 2007): a component is constitutively relevant to a phenomenon if mutual manipulability holds — intervening on the component affects the phenomenon and vice versa. This is what activation patching and ablation measure.

The realism spectrum. The nine views span a realism-to-instrumentalism spectrum. The stratified view sits at the committed end: it requires evidence across multiple strata and measurement resolutions. The structural view is strongly realist in a different sense — a mechanism is the gauge orbit of weight configurations computing the same function, which exists independent of any measurement procedure. Perspectival views hold that mechanism descriptions are always relative to a measurement procedure — what exists is the procedure and its outputs (cf. Massimi 2022). Instrumentalist views hold that “mechanism” is useful shorthand for a predictive model. The contrastive view sits below the perspectival view, and the object, role, subspace, structural, and process views above it. Choosing a view is implicitly a commitment about how realist one’s conclusions are entitled to be.

The qua-problem. The qua-problem arises whenever an entity can be described under multiple predicates: head 9.9 is “the same mechanism” as head 7.2 in another model qua what? Under the component description they are different; under the role description they may be the same; under the gauge-orbit description they may be the same if the orbits are isomorphic. Every mechanistic identity claim is implicitly a qua-claim. The five axes force that implicit claim to be stated explicitly: identity qua gauge orbit (structural view), qua role (role view), qua geodesic proximity (subspace view), qua formation process (process view). There is no view-independent answer to “are these the same mechanism?” — only answers relative to a specified identity criterion.