Style-Debiased DPO: Updating LLM Knowledge with Factuality-Aware Synthetic Preference Data
Researchers propose style-debiased direct preference optimization (SD-DPO) to improve the accuracy of large language models (LLMs) in retrieving stored knowledge.