Data-DPO paper proposes target-model-aware selection of fine-tuning data
An arXiv paper introduces Data-DPO, which selects supervised fine-tuning examples using a target model’s responses during brief training probes. Authors report better results than existing selection methods and full-data training on Vision-Flan and LLaVA-CoT, but the record lacks effect sizes or independent validation.