Study finds language models can represent when they are being evaluated
An arXiv study reports that six language models contained linearly decodable signals associated with being evaluated, while those internal signals only partly matched what the models said. The authors say steering along probe-derived directions changed verbalization scores, raising questions about how reliably…