Skip to content

Is the Component the Mechanism, or a Coordinate of One?

Section titled “Is the Component the Mechanism, or a Coordinate of One?”

The question. Ablation establishes that a component matters. It does not establish that the component is the mechanism rather than a coordinate of one. If a concept occupies a low-dimensional manifold spanned by many units, ablating any contributing unit perturbs the manifold and registers as necessity — so single-unit causal evidence returns the same verdict under both readings.

This is an Object view ceiling that further ablations cannot raise. No amount of additional single-unit intervention distinguishes “this unit is the mechanism” from “this unit is one coordinate of a mechanism spanning many units,” because both predict the same ablation result.

Gallego et al. (2017) observe that “many M1 neurons do not relate to any single movement covariate” and conclude that “the search for representations at the single-neuron level might actually divert us from understanding the neural control of movement.” The repair was to change the unit of analysis to population modes, under which the activity of a single neuron is “simply a reflection of the latent variables.”

The dispute is now running in mechanistic interpretability, and the two sides are separated by unit rather than by evidence quality.

Nikankin et al. (2025) classify 91% of the 3,200 top neurons for each arithmetic operator in Llama3-8B as narrow heuristics, report similar results for GPT-J and Pythia-6.9B, and conclude that these models compute without an algorithm.

Kantamneni et al. (2025) analyze two of the same models — with Llama3.1-8B in place of Llama3-8B — at the level of representational geometry, and recover a helix of 2k+12k+1 basis functions, k4k \leq 4, carrying a clean addition algorithm.

Both validate causally. The described objects differ by more than two orders of magnitude in size, and the disagreement is about what kind of thing a mechanism is — not about whether either measurement is correct.

The framework identifies which view owns the problem. Under the Object view the ceiling is real and further ablation cannot raise it. Under the Subspace view the same evidence is about a subspace’s dimensionality rather than a component’s necessity, and the two results above are compatible descriptions at different resolutions — which is the Stratified view’s reading. A claim must declare which it makes.

  • Gallego et al. (2017): Neural manifolds for the control of movement (Neuron)
  • Nikankin et al. (2025): Arithmetic without algorithms — language models solve math with a bag of heuristics (ICLR 2025)
  • Kantamneni et al. (2025): Language models use trigonometry to do addition (arXiv:2502.00873)