Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion
Researchers introduce a new framework to evaluate the robustness of Large Language Models against persuasion attacks and achieve a 96% success rate with simple attack strategies.