Skip to content

We define nine mechanistic views — coherent positions on what a mechanism is, when two are the same, and what counts as evidence — ordered by increasing ontological commitment:

Instrumental<Contrastive<Perspectival<{Object, Role}<Subspace<{Structural, Process}<Stratified\text{Instrumental} < \text{Contrastive} < \text{Perspectival} < \{\text{Object},\ \text{Role}\} < \text{Subspace} < \{\text{Structural},\ \text{Process}\} < \text{Stratified}

This ordering is partial, and braces enclose the pairs it does not rank. Object and Role are incomparable (a role is multiply realizable while a component identity is tied to a specific model), and the relative positions of Structural and Process depend on whether one weights temporal or algebraic structure as more committed.

Higher-commitment views make stronger claims but require more evidence. The instrumental view requires only predictive utility; the stratified view requires evidence across multiple strata and measurement resolutions.

Most published interpretability work operates at the Object or Role level without stating this explicitly. This creates confusion when papers using different views disagree — the IOI circuit replication debates, the probing wars, and “how big is a circuit?” arguments are all arguably partly view-level disagreements alongside empirical ones. Making the view explicit takes one sentence and prevents a class of errors where evidence collected at one level is used to draw conclusions at another.

The views apply to both mechanisms (circuits, computations) and channels (the representations those circuits operate over). A channel is a mechanism at a different grain: the Object view says it is a specific direction or neuron, the Role view says it is whatever plays a given functional role, the Subspace view says it is a causal subspace, and so on. Whether a given channel constitutes a “real feature” is itself a view-dependent question — the Object view says yes if it is a stable direction, the Subspace view says yes if it is causally active, and the Perspectival view says the answer is relative to the detection process, and asks what survives changing it.

The views are not mutually exclusive: a single paper may use several, and convergence across views can constitute stronger evidence than any single view supplies — on one condition. A method cannot rule out its own characteristic artifact, so agreement counts only when the agreeing views fail differently. The nine views are partitioned into five families by primary failure mode, and a structure is view-invariant when its support crosses at least two families with no undefeated incompatible result, contested when applied views yield incompatible results, and a candidate otherwise. See View Families and the Determination Chain for the families.

ViewOntologyIdentityEvidenceFormalismTarget
InstrumentalPredictive modelPredictive equivalenceForecast, interventionModel theoryBehavior prediction
ContrastiveDifference relative to foilSame contrastive patternFoil-varied patchingContrastive explanationWhy PP rather than QQ
PerspectivalMethod projectionCross-method coherenceMulti-method robustnessMeasurement algebraMethod-relative
ObjectConcrete partComponent overlapAblation, patchingDirected graphSpecific behavior
RoleFunctional roleRole equivalenceRole-specific causal testsFunctional decomp.Functional class
SubspaceCausal subspaceSame projectorDAS/IIA (linear), subspace stabilityGrassmannian Gr(k,d)\mathrm{Gr}(k,d)Representational variable
StructuralGauge-invariant structureGauge orbitHolonomy, compositionFiber bundleComputation class
ProcessFormation trajectorySame trajectory typeCheckpoints, formation knockoutsDynamical systemMechanism origin
StratifiedStratum pointStratum + local equivalenceParticipation ratio, localizationRes.-indexed strataResolution-relative

Each view comes with a test of whether it is the right lens for a given body of evidence, and a negative result that says what to do instead. Where cross-view promotion asks what evidence raises a claim to a view, discrimination asks whether the view was the right one to begin with.

ViewTestWhat failure implies
ObjectAblate the identified component set; check that degradation is (a) specific to the behavior the mechanism should implement and (b) stable across prompt distributionsIf (b) fails, the object-level description is distribution-relative, and a role or subspace description fits better
RoleApply the role specification to a model in which the original component has been ablatedIf no backup component takes over and the behavior does not degrade in a role-consistent way, the role label is not independently validated
SubspaceCheck that the subspace recovered on a held-out distribution is close in geodesic distance to the training one; confirm a basis change inside the subspace leaves IIA unchanged while a rotation out of it reduces IIAThe search found a subspace that permits successful interchange rather than the causally relevant one; confirming it requires weight-space evidence
StructuralApply a computation-preserving symmetry (scale by λ\lambda, counter-scale by λ1\lambda^{-1}) to a component’s weights and retest the activation-space evidenceIf the activation result changes, it was not gauge-invariant and is not a structural claim
ProcessTrain the same architecture under different hyperparameters and compare formation dynamicsIf the final circuit is similar but the trajectory differs, the process claim is hyperparameter-sensitive and must be stated at that generality
StratifiedCompare participation ratio and localizability for the recovered causal subspace at different hook points and under different alignment constraintsA stratum change indicates either a higher-stratum mechanism partially recovered at the coarser probe, or an artifact of the finer one
PerspectivalTest whether a patching-supported claim is also supported by weight-space composition scoresIf not, a third method adjudicates
InstrumentalTest whether the mechanism label generates novel predictions beyond the behavior used to identify itIf not, the label only redescribes the behavior used to identify it
ContrastiveVary the foil systematically — “repeat the subject,” “output a random name,” “output nothing”If the identified circuit changes with the foil, the mechanism is foil-relative and should be stated as such

A view that survives its own discriminating experiment has not thereby been confirmed: the test rules out the view being the wrong lens, not the claim being wrong. Each view’s page carries the full entry — identity criterion, admissible evidence, and the characteristic failure mode that fixes its family.

The views are not mutually exclusive at the level of phenomena. A given model may contain mechanisms best described by different views at different strata. The goal is to identify which commitments different descriptions carry, not to select a single correct view.

See View Families and the Determination Chain for how the views order, group and constrain each other, and for how other positions in the literature map onto these nine.