Skip to content

Mechanistic Views

What is a mechanism?

Mechanistic interpretability has produced circuits, features, and causal subspaces — each resting on an implicit answer to a prior question: what kind of thing is a mechanism? In some papers a mechanism is a set of attention heads. In others it is a functional role (“name mover”), a causal subspace, or a training-time transition. These are different claims about different kinds of objects — and the difference determines when two descriptions pick out the same mechanism, what evidence can support the claim, and what inferences follow from it. A mechanistic claim has no determinate truth-conditions until these commitments are declared.

We define these commitments as a mechanistic view, formalizing the background assumptions behind a mechanistic claim.

Mechanistic Views is a framework defining what is a mechanism, when two mechanisms are the same, and the mathematical formalisms for representing and comparing mechanisms. We present an atlas of 9 views spanning existing interpretability claims, and a demonstration of why the views framework helps clarify open problems in the field.

What is a mechanistic view?

A mechanistic view is the set of commitments — explicit or implicit — underlying a mechanistic claim. It is formalized as five questions:

AxisQuestion
OntologyWhat kind of entity counts as a mechanism?
IdentityWhen do two descriptions refer to the same mechanism?
EvidenceWhat measurements can warrant a claim about a mechanism?
FormalismWhat mathematical language expresses the claim?
TargetWhat phenomenon is the mechanism supposed to explain?

Five axes of a mechanistic view: Ontology, Identity, Evidence, Formalism, Target

The five answers are not independent — what a mechanism is constrains when two are the same, which constrains what formalism is needed. See Framework for the formal definition, coherence conditions, and why these constraints matter.

Nine mechanistic views

The five axes admit multiple coherent answers. We organize these into nine mechanistic views, ordered by increasing ontological commitment — how much they claim about what mechanisms are, independent of any measurement:

Instrumental<Contrastive<Perspectival<{Object, Role}<Subspace<{Structural, Process}<Stratified\text{Instrumental} < \text{Contrastive} < \text{Perspectival} < \{\text{Object},\ \text{Role}\} < \text{Subspace} < \{\text{Structural},\ \text{Process}\} < \text{Stratified}

This ordering is partial, not total: Role and Object are incomparable, and the relative positions of Process and Structural depend on whether one weights temporal or algebraic structure as more committed. The endpoints are clear: the instrumental view requires only predictive utility, while the stratified view requires evidence across multiple strata and measurement resolutions. Higher-commitment views make stronger claims but require more evidence.

ViewOntologyIdentityEvidenceFormalismTarget
InstrumentalPredictive modelPredictive equivalenceForecast, intervention utilityModel theoryBehavioral prediction
ContrastiveDifference relative to foilSame contrastive patternFoil-varied patchingContrastive explanationWhy PP rather than QQ
PerspectivalMethod projectionCross-method coherenceMulti-method robustnessMeasurement algebraMethod-relative
ObjectConcrete partComponent overlapAblation, patchingDirected graphSpecific behavior
RoleFunctional roleRole equivalenceRole-specific causal testsFunctional decompositionFunctional class
SubspaceCausal subspaceSame projectorDAS/IIA (linear), subspace stabilityGrassmannian Gr(k,d)\mathrm{Gr}(k,d)Representational variable
StructuralGauge-invariant structureGauge-orbit membershipHolonomy, composition scoresFiber bundle quotientComputation class
ProcessFormation trajectorySame trajectory typeCheckpoints, formation knockoutsDynamical systemMechanism origin
StratifiedStratum pointStratum + local equivalenceParticipation ratio, localizationResolution-indexed strataResolution-relative

Relationship to Mechanistic Validity

Mechanistic Views defines what a mechanism is — what kind of object is it, when two mechanisms are the same, and what formalism expresses the claim. Mechanistic Validity evaluates the evidence behind a mechanistic claim — whether the measurement is reliable, the causal interpretation is justified, the result generalizes within its stated scope, and the description matches what was tested. The frameworks are independent, but when used together they complement each other: what kind of mechanism is being claimed, and how strong is the evidence supporting it.

Reading path

  • Framework — the five axes, coherence conditions, and why they matter
  • Views Overview — each of the 9 views with its ontology, identity criterion, evidence, formalism, and target
  • Cross-View Promotion — when a claim can be promoted to a stronger view, and what new evidence that requires
  • Methods — how 13 common interpretability methods map to implicit view commitments
  • Open Problems — how common interpretability problems can be better understood using views
  • About — about this project, how to cite it, disclaimers