View Families and the Determination Chain
Section titled “View Families and the Determination Chain”This page shows how the nine views relate to each other, and how other positions in the interpretability literature map onto them.
The nine views are ordered by ontological commitment, grouped into five families by shared failure mode, and linked by the determination chain, which runs from ontology to identity to formalism.
Ontological commitment ordering
Section titled “Ontological commitment ordering”View families
Section titled “View families”The nine views group into five families, where membership is fixed by primary failure mode: two views belong to the same family when one methodological error corrupts both. The Identity family (Object, Role; failure: role inflation), the Mathematical family (Subspace, Structural; failure: alignment vacuousness), the Process family (Process; failure: training-detail overfitting), the Pragmatic family (Instrumental; failure: safety-relevant overreach), and the Analyst-choice family (Contrastive, Perspectival, Stratified; failure: analyst-choice dependence — each individuates mechanisms relative to a parameter the analyst sets).
The determination chain
Section titled “The determination chain”What a mechanism is fixes when two count as the same, and that fixes what mathematics can express the claim. The figure runs the chain for all nine views.
The chain is a strong tendency rather than an entailment. DAS is the instructive deviation: its search space is a subspace parameterization over , but it validates by interchange intervention accuracy, which is role equivalence. The hybrid works because role success on a subspace search space is sufficient for the subspace being causally relevant — so a method that departs from the chain is flagged, as DAS is in the methods table, rather than treated as an error.
Other positions in the literature
Section titled “Other positions in the literature”Several coherent positions appear in the interpretability literature that are not listed as separate views here. In each case, we argue the position maps onto one of the nine views, adds a constraint to an existing view, or lacks practical methods for trained models.
Algorithmic (RASP, Tracr). Treats mechanisms as formal programs the network implements. A coherent philosophical position, but the only cases where “this network implements algorithm X” can be verified are toy models or very simple circuits where the algorithm is already obvious. For real models, the computation is too distributed and approximate to extract a clean program. Algorithmic claims about trained models reduce to detailed role descriptions in practice.
Information-theoretic (mutual information bottlenecks, probing-as-compression). Defines mechanisms as information channels between input and output variables. This collapses into the Perspectival view: what you find depends on your choice of MI estimator, partitioning scheme, and layer — the “mechanism” is relative to the measurement procedure. Information theory is listed as a cross-cutting formalism rather than a view for this reason.
Distributed / superposition (Elhage et al. “Toy Models of Superposition”). Claims that mechanisms are overcomplete codes spread across many dimensions simultaneously. This is not a separate view but a negative result within the Object view: it shows that clean component-level decomposition fails in certain regimes. The positive version — “the mechanism is the interference structure itself” — is a Subspace view claim about the geometry of overcomplete representations.
Representational similarity (CKA, RSA, SVCCA). Identifies mechanisms with representational geometries: two layers or models have the “same mechanism” when their representations have high CKA or RSA similarity. This is the Subspace view with a different metric. CKA compares subspaces; RSA compares geometry within subspaces. The ontological commitment is the same — mechanisms are geometric properties of representation spaces.
Compression / MDL (minimum description length, pruning-as-compression). Defines the mechanism as whatever survives optimal compression. This makes no claim about what the mechanism is, only that it is predictively necessary — exactly the Instrumental view.
Concept bottleneck (TCAV, Network Dissection, concept bottleneck models). Treats mechanisms as human-interpretable concepts used as intermediate variables. This is the Role view with an additional constraint: the role must be human-legible. It does not require a new ontology — a concept is a functional role that happens to align with a human category. It can also appear as the Subspace view plus interpretability (the concept is a direction in activation space that aligns with a human-meaningful axis).
Phase transition / modularity (grokking, induction head formation). Claims that mechanisms emerge at specific training loss thresholds as discrete phase transitions. This is the Process view: a phase transition is a feature of the formation trajectory — the point where a qualitative change occurs along the training dynamics. The dynamical systems formalism already covers bifurcations and critical transitions.