Study reports targeted steering can reduce sycophancy and hallucination in medical AI answers
An arXiv study describes an inference-time method that selectively steers language models when they appear likely to hallucinate or change a correct medical answer under user pressure. In 600 pressure trajectories involving a 4-billion-parameter model, the authors say the unsteered model abandoned its answer 570…