Paper proposes velocity matching to scale reward fine-tuning for diffusion models
Researchers propose reward-based velocity matching, a trajectory-free method that directly updates diffusion models’ velocity fields and, they report, achieves comparable or better results than likelihood-based methods at lower training cost.