Study finds AI deception signals transfer unevenly between test settings
An arXiv study of Llama-3.1-70B-Instruct reports that spontaneous and instructed deception share some internal direction, but detection and steering methods transfer asymmetrically between them.