Skip to content

The evidence base is one sentence: a classifier recovers a variable zz from a representation. Three readings follow from it. One holds. Controls have been run against the other two and both failed.

A mechanism is a predictive model; identity is predictive equivalence. The claim is that zz is extractable by a classifier of that family from that representation. Probe accuracy establishes this directly, and nothing further is owed.

This is a claim about how the activation space is organized, and a control disconfirms it rather than leaving it open. Control tasks — the same probe trained to predict randomly assigned labels — attribute most of a nonlinear probe’s accuracy to the classifier rather than to the representation (Hewitt and Liang, 2019). A probe powerful enough to fit noise cannot distinguish structure it found from structure it built.

This is a claim about function, and a second control disconfirms it. A control dataset holding zz non-discriminative for the original task leaves probe accuracy intact (Ravichander et al., 2021): the representation carries zz whether or not the task needs it, so extractability is not evidence of use.

The atlas classifies linear probing under role ontology and role-equivalence identity, and marks both cells as judgments rather than as facts about the method. The live alternative is the instrumental reading — a probe is a predictive model, identity is predictive equivalence — and it is consequential: taking it moves probing from the identity family to the pragmatic family. That changes verdicts. A claim supported by probing plus one identity-family method sits inside a single family under the reading taken in the paper and crosses a family boundary under the alternative, so some structures recorded as candidate would become view-invariant.

The reframings probing carries are themselves analyst-choice evidence: they are disagreements about what the measure quantifies, which is information about the instrument rather than about the model.

ViewThe claimThe verdict on this evidence
Instrumentalzz is extractable by a classifier of that familyHolds
SubspaceThe model represents zzDisconfirmed by control tasks
RoleThe model uses zzDisconfirmed by control datasets

Hewitt, J., Liang, P. “Designing and Interpreting Probes with Control Tasks.” EMNLP-IJCNLP 2019. doi:10.18653/v1/D19-1275.

Ravichander, A., Belinkov, Y., Hovy, E. “Probing the Probing Paradigm: Does Probing Accuracy Entail Task Relevance?” EACL 2021, 3363–3377. doi:10.18653/v1/2021.eacl-main.295.

See also Refusal Direction and the Probe features open problem.