Preprint reports SAFE-G gains for evidence-grounded visual question answering
A new arXiv preprint describes SAFE-G, a multimodal question-answering framework that combines graph-based evidence retrieval with reinforcement learning to keep answers tied to supporting information. The authors report accuracy gains on two benchmarks, though the supplied abstract does not provide absolute scores…