Paper proposes tensor-product representations as a common framework for interpreting language models
An arXiv preprint argues that four widely used language-model interpretability methods can be understood as different operations on a shared filler-role structure.