Open Problem Map
Section titled “Open Problem Map”Do SAEs recover the “true” features? Is chain-of-thought faithful to the model’s reasoning? Do probes find features the model actually uses? These are usually treated as empirical questions — run the experiment, get the answer. But each contains a hidden assumption: what “true” means, what “faithful” means, what “uses” means. Different views define these differently, and we find that disagreements in the literature often reflect unstated view commitments rather than conflicting experimental results.
We distilled 18 distinct problem types from eight sources, three of which enumerate hundreds of individual open questions. The foundational ones — why methods disagree, what evidence can establish, when findings are the same — are clarified, reframed, or scoped when you name the view. The methods-level ones turn out to require foundational clarity before they can be answered. Each maps onto specific Mechanistic Validity criteria being violated or neglected.
Of the 18 problems: 8 are clarified (the question becomes answerable once the view is specified), 2 are reframed (the framework clarifies what evidence would answer it), and 8 are scoped (the framework identifies which view owns the problem).
Problem map
Section titled “Problem map”| Page | Problem | The confusion | Views involved | Status |
|---|---|---|---|---|
| Decomposition identity | SAE features change with width | Which decomposition is “real”? | Object vs. Subspace | Clarified |
| SAE “true” features | Do SAEs recover the “true” features? | ”True” depends on the view | Object vs. Subspace | Clarified |
| Superposition | How do models represent more features than dimensions? | Problem, property, or presupposition? | Object vs. Subspace vs. Perspectival | Clarified |
| Circuit disagreements | ACDC vs. EAP vs. patching disagree | Which circuit is correct? | Object vs. Role | Clarified |
| Unit of analysis | Are attention heads the right unit? | Unit depends on the claim | Object vs. Role vs. Subspace | Clarified |
| Probe features | Do probes find features the model uses? | ”Uses” means four different things | Instrumental vs. Subspace vs. Role | Clarified |
| Cross-model identity | ”Same features” across models | Same by what criterion? | Object impossible, Subspace possible | Clarified |
| DAS in random models | DAS finds subspaces in random models | Real subspace or optimization artifact? | Perspectival diagnosis | Clarified |
| CoT faithfulness | Is CoT faithful to internal reasoning? | ”Faithful” means four things | Instrumental vs. Structural | Reframed |
| World models | Do LLMs build internal world models? | Subspace evidence used for Structural claims | Subspace vs. Structural | Reframed |
| Intervention limits | Ablation triggers reconfiguration | Ablation reveals function or breaks it? | Object ceiling | Scoped |
| Level of analysis | Ablation ≠ localization | Component or coordinate of a mechanism? | Object ceiling | Scoped |
| Linear assumption | Concepts form circles, not lines | Is linear structure real? | Subspace ceiling | Scoped |
| Safety evidence gaps | ”Refusal direction” = refusal mechanism? | Lever or mechanism? | Instrumental floor | Scoped |
| Validation methodology | Validation on discovery data | Circular evidence? | Perspectival diagnosis | Scoped |
| Training dynamics | Fine-tuning changes circuits | Enhanced, bypassed, or relocated? | Process territory | Scoped |
| Resolution & completeness | ”Dark matter” — unexplained behavior | How much do we understand? | Stratified territory | Scoped |
| Deceptive alignment | Can MI detect strategic deception? | Instrumental evidence is structurally insufficient | Instrumental floor, Structural + Process required | Scoped |
Resolution status
Section titled “Resolution status”Each problem has a resolution status indicating how much the framework addresses it:
Clarified — the question becomes answerable once the view is specified. This does not mean the underlying phenomena aren’t real — it means the question as stated conflates different views and therefore has no single answer. Once you specify which view you’re in, the question either has a clear answer or becomes a different (better) question. Examples: circuit disagreements, decomposition identity, superposition, the probe wars.
Reframed — the framework clarifies what the question actually means and what evidence would answer it, but the empirical question remains open. The conceptual confusion is resolved; the experiment hasn’t been run. Example: CoT faithfulness — “faithful” means four different things, and the safety-relevant notion requires evidence no one has collected.
Scoped — the framework identifies which view owns the problem, what the evidence standard is, and why current methods hit a ceiling. The problem is real and structural — it requires moving up the commitment ladder or working in a view the field hasn’t adopted yet. Examples: intervention limits (Object ceiling), training dynamics (Process territory).
The pattern
Section titled “The pattern”Three types of problem appear:
View confusions — the problem is clarified when you name which view you’re in. Circuit disagreements, decomposition identity, cross-model identity, superposition, SAE “true” features, the probe wars, and the unit-of-analysis debate are all cases where researchers using different views get different answers and mistake the disagreement for an empirical conflict. It isn’t. Different views ask different questions and should give different answers.
View ceilings — the problem is real and structural, caused by the limits of the view being used. Intervention limits, the linear assumption, and safety evidence gaps all arise because a view’s methods cannot produce the evidence needed for the conclusions being drawn. The fix is to move up the commitment ladder: collect evidence at a higher-commitment view where the criterion is possible rather than impossible.
View territories — some problems belong naturally to a specific view that the field hasn’t adopted yet. Training dynamics problems belong to the Process view. Resolution problems belong to the Stratified view. These aren’t confusions — they’re under-explored areas of the view space.
What the framework does and doesn’t do
Section titled “What the framework does and doesn’t do”The views framework defines what mechanisms are, what evidence means, and when two mechanisms are the same. It is a framework about mechanistic interpretability — it provides evidence standards and ontological commitments. It does not itself do MI research: it doesn’t find circuits, train SAEs, or detect deception.
Several prominent “open problems” in MI are methods-level problems rather than foundational problems. The framework already provides the evidence standards for evaluating solutions to them:
| Problem | Why it’s not a framework gap | What the framework provides |
|---|---|---|
| CoT faithfulness — is chain-of-thought reasoning faithful to internal computation? | This is an empirical question about a specific method | The Perspectival view already predicts that single-method evidence (including CoT) may be unfaithful. Cross-method convergence is the fix |
| Deceptive alignment — can MI detect strategic deception? | This is a target-property problem: can methods find a specific thing? | The framework says what evidence level is needed: at minimum Structural (gauge-invariant, stable under adversarial pressure) + Process (survives retraining). Instrumental evidence (steering vectors) is structurally insufficient. Full analysis → |
| World models — do LLMs form causal world models? | This is a specific empirical finding to chase, not a framework question | Linear encodings of state predicates are Subspace claims. Whether they form a causal world model is a Structural question. Mechanistic Validity E4 (counterfactual validity) is the relevant criterion. Full analysis → |
| Automated MI at scale — how do we scale MI to frontier models? | This is engineering, not ontology | The framework provides the evidence standards that automated systems should satisfy — especially M1 Reliability and I5 Rival mechanism exclusion (does the automated system find real structure or artifacts?) |
| Non-transformer architectures — does MI transfer to SSMs, MoE, etc.? | The framework is architecture-agnostic by design | E4 Cross-model generalization is already a criterion. Whether specific methods (SAEs, DAS) transfer is a methods question |
| Formal verification of circuits — can mechanistic claims be deductively verified? | Deductive certainty is complementary to but separate from empirical validity | The framework assesses empirical validity. Formal verification provides deductive certainty for specific algorithmic claims. Both are useful; they answer different questions |
The one genuine scope boundary is multi-agent mechanistic interpretability. The views framework assumes single-model analysis — “mechanism” is defined within one model. Mechanisms that emerge from interaction between models (coordination protocols, collusion strategies) have no natural home in the current view hierarchy. This is a deliberate scope choice (single-model analysis is already hard enough), not an oversight, but it is a real boundary.
Sources
Section titled “Sources”| Source | Year | Problems | Focus |
|---|---|---|---|
| Sharkey et al. “Open Problems in MI” | 2025 | 67 | Comprehensive survey (30 authors) |
| Nanda “200 Concrete Open Problems” | 2022 | 200 | Enumerated research questions |
| Apollo Research project ideas | 2024 | 45+ | Safety-oriented MI projects |
| Schmidt Sciences “Trustworthy AI” agenda | 2026 | 38 | Evaluation validity, deception, oversight |
| Steinhardt “The Case for Evaluating Model Behaviors” | 2026 | 37 | Behavioral propensity measurement |
| Orgad, Barez et al. “Interpretability Can Be Actionable” | ICML 2026 | 48 | Actionability and comparative advantage |
| MIB: Mechanistic Interpretability Benchmark | 2025 | — | Standardized causal localization evaluation |
| Williams et al. “MI Needs Philosophy” | 2025 | — | Conceptual foundations and epistemic status |