Skip to content

Do SAEs recover the “true” features? Is chain-of-thought faithful to the model’s reasoning? Do probes find features the model actually uses? These are usually treated as empirical questions — run the experiment, get the answer. But each contains a hidden assumption: what “true” means, what “faithful” means, what “uses” means. Different views define these differently, and we find that disagreements in the literature often reflect unstated view commitments rather than conflicting experimental results.

We distilled 18 distinct problem types from eight sources, three of which enumerate hundreds of individual open questions. The foundational ones — why methods disagree, what evidence can establish, when findings are the same — are clarified, reframed, or scoped when you name the view. The methods-level ones turn out to require foundational clarity before they can be answered. Each maps onto specific Mechanistic Validity criteria being violated or neglected.

Of the 18 problems: 8 are clarified (the question becomes answerable once the view is specified), 2 are reframed (the framework clarifies what evidence would answer it), and 8 are scoped (the framework identifies which view owns the problem).

PageProblemThe confusionViews involvedStatus
Decomposition identitySAE features change with widthWhich decomposition is “real”?Object vs. SubspaceClarified
SAE “true” featuresDo SAEs recover the “true” features?”True” depends on the viewObject vs. SubspaceClarified
SuperpositionHow do models represent more features than dimensions?Problem, property, or presupposition?Object vs. Subspace vs. PerspectivalClarified
Circuit disagreementsACDC vs. EAP vs. patching disagreeWhich circuit is correct?Object vs. RoleClarified
Unit of analysisAre attention heads the right unit?Unit depends on the claimObject vs. Role vs. SubspaceClarified
Probe featuresDo probes find features the model uses?”Uses” means four different thingsInstrumental vs. Subspace vs. RoleClarified
Cross-model identity”Same features” across modelsSame by what criterion?Object impossible, Subspace possibleClarified
DAS in random modelsDAS finds subspaces in random modelsReal subspace or optimization artifact?Perspectival diagnosisClarified
CoT faithfulnessIs CoT faithful to internal reasoning?”Faithful” means four thingsInstrumental vs. StructuralReframed
World modelsDo LLMs build internal world models?Subspace evidence used for Structural claimsSubspace vs. StructuralReframed
Intervention limitsAblation triggers reconfigurationAblation reveals function or breaks it?Object ceilingScoped
Level of analysisAblation ≠ localizationComponent or coordinate of a mechanism?Object ceilingScoped
Linear assumptionConcepts form circles, not linesIs linear structure real?Subspace ceilingScoped
Safety evidence gaps”Refusal direction” = refusal mechanism?Lever or mechanism?Instrumental floorScoped
Validation methodologyValidation on discovery dataCircular evidence?Perspectival diagnosisScoped
Training dynamicsFine-tuning changes circuitsEnhanced, bypassed, or relocated?Process territoryScoped
Resolution & completeness”Dark matter” — unexplained behaviorHow much do we understand?Stratified territoryScoped
Deceptive alignmentCan MI detect strategic deception?Instrumental evidence is structurally insufficientInstrumental floor, Structural + Process requiredScoped

Each problem has a resolution status indicating how much the framework addresses it:

Clarified — the question becomes answerable once the view is specified. This does not mean the underlying phenomena aren’t real — it means the question as stated conflates different views and therefore has no single answer. Once you specify which view you’re in, the question either has a clear answer or becomes a different (better) question. Examples: circuit disagreements, decomposition identity, superposition, the probe wars.

Reframed — the framework clarifies what the question actually means and what evidence would answer it, but the empirical question remains open. The conceptual confusion is resolved; the experiment hasn’t been run. Example: CoT faithfulness — “faithful” means four different things, and the safety-relevant notion requires evidence no one has collected.

Scoped — the framework identifies which view owns the problem, what the evidence standard is, and why current methods hit a ceiling. The problem is real and structural — it requires moving up the commitment ladder or working in a view the field hasn’t adopted yet. Examples: intervention limits (Object ceiling), training dynamics (Process territory).

Three types of problem appear:

View confusions — the problem is clarified when you name which view you’re in. Circuit disagreements, decomposition identity, cross-model identity, superposition, SAE “true” features, the probe wars, and the unit-of-analysis debate are all cases where researchers using different views get different answers and mistake the disagreement for an empirical conflict. It isn’t. Different views ask different questions and should give different answers.

View ceilings — the problem is real and structural, caused by the limits of the view being used. Intervention limits, the linear assumption, and safety evidence gaps all arise because a view’s methods cannot produce the evidence needed for the conclusions being drawn. The fix is to move up the commitment ladder: collect evidence at a higher-commitment view where the criterion is possible rather than impossible.

View territories — some problems belong naturally to a specific view that the field hasn’t adopted yet. Training dynamics problems belong to the Process view. Resolution problems belong to the Stratified view. These aren’t confusions — they’re under-explored areas of the view space.

The views framework defines what mechanisms are, what evidence means, and when two mechanisms are the same. It is a framework about mechanistic interpretability — it provides evidence standards and ontological commitments. It does not itself do MI research: it doesn’t find circuits, train SAEs, or detect deception.

Several prominent “open problems” in MI are methods-level problems rather than foundational problems. The framework already provides the evidence standards for evaluating solutions to them:

ProblemWhy it’s not a framework gapWhat the framework provides
CoT faithfulness — is chain-of-thought reasoning faithful to internal computation?This is an empirical question about a specific methodThe Perspectival view already predicts that single-method evidence (including CoT) may be unfaithful. Cross-method convergence is the fix
Deceptive alignment — can MI detect strategic deception?This is a target-property problem: can methods find a specific thing?The framework says what evidence level is needed: at minimum Structural (gauge-invariant, stable under adversarial pressure) + Process (survives retraining). Instrumental evidence (steering vectors) is structurally insufficient. Full analysis →
World models — do LLMs form causal world models?This is a specific empirical finding to chase, not a framework questionLinear encodings of state predicates are Subspace claims. Whether they form a causal world model is a Structural question. Mechanistic Validity E4 (counterfactual validity) is the relevant criterion. Full analysis →
Automated MI at scale — how do we scale MI to frontier models?This is engineering, not ontologyThe framework provides the evidence standards that automated systems should satisfy — especially M1 Reliability and I5 Rival mechanism exclusion (does the automated system find real structure or artifacts?)
Non-transformer architectures — does MI transfer to SSMs, MoE, etc.?The framework is architecture-agnostic by designE4 Cross-model generalization is already a criterion. Whether specific methods (SAEs, DAS) transfer is a methods question
Formal verification of circuits — can mechanistic claims be deductively verified?Deductive certainty is complementary to but separate from empirical validityThe framework assesses empirical validity. Formal verification provides deductive certainty for specific algorithmic claims. Both are useful; they answer different questions

The one genuine scope boundary is multi-agent mechanistic interpretability. The views framework assumes single-model analysis — “mechanism” is defined within one model. Mechanisms that emerge from interaction between models (coordination protocols, collusion strategies) have no natural home in the current view hierarchy. This is a deliberate scope choice (single-model analysis is already hard enough), not an oversight, but it is a real boundary.

SourceYearProblemsFocus
Sharkey et al. “Open Problems in MI”202567Comprehensive survey (30 authors)
Nanda “200 Concrete Open Problems”2022200Enumerated research questions
Apollo Research project ideas202445+Safety-oriented MI projects
Schmidt Sciences “Trustworthy AI” agenda202638Evaluation validity, deception, oversight
Steinhardt “The Case for Evaluating Model Behaviors”202637Behavioral propensity measurement
Orgad, Barez et al. “Interpretability Can Be Actionable”ICML 202648Actionability and comparative advantage
MIB: Mechanistic Interpretability Benchmark2025Standardized causal localization evaluation
Williams et al. “MI Needs Philosophy”2025Conceptual foundations and epistemic status