뉴스로 돌아가기
혁신AI Understanding 브리핑

논문에서는 확산 모델에 대한 보상 미세 조정 규모를 조정하기 위해 속도 일치를 제안합니다.

연구원들은 확산 모델의 속도 필드를 직접 업데이트하는 궤적 없는 방법인 보상 기반 속도 매칭을 제안하고, 더 낮은 교육 비용으로 우도 기반 방법과 비슷하거나 더 나은 결과를 달성한다고 보고합니다.

5 min readRead the primary source
Primary-source image accompanying Paper proposes velocity matching to scale reward fine-tuning for diffusion models
기본 소스 문서녹음된 소스
출판사
arxiv.org
소스 링크
arxiv.orghttps://arxiv.org/abs/2608.23664
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

미세 조정
사전 훈련된 모델을 특정 작업에 맞게 조정하기 위해 도메인별 데이터에 대한 지속적인 훈련입니다.
강화 학습
에이전트가 장기적인 수익을 극대화하는 행동을 학습하는 보상 신호를 통한 교육입니다.
확산 모델
이미지, 오디오 또는 기타 콘텐츠를 합성하기 위해 노이즈를 역전시키는 방법을 학습하는 생성 아키텍처입니다.
자신을 테스트해 보세요AI 모델 설명 퀴즈

무슨 일이 일어났나요?

A team of seven researchers submitted an arXiv paper proposing reward-based velocity matching (RVM), a method for diffusion models toward human preferences or task-specific goals. RVM updates the model’s velocity field directly instead of reconstructing likelihoods from denoising trajectories or approximating endpoint likelihoods.

The source is an arXiv submission dated Aug. 24, 2026, titled “Scaling for Diffusion Models via Velocity Matching.” The authors present reward as a way to adapt diffusion models to human preferences and task-specific objectives. Their central claim is that existing approaches borrow policy-gradient machinery designed for autoregressive models, creating added complexity because diffusion models do not provide tractable likelihoods for generated samples.

The proposed method, reward-based velocity matching, is described as trajectory-free. Rather than constructing likelihoods from stochastic denoising transitions or approximating endpoint likelihoods with an evidence lower bound, RVM acts directly on the velocity field. According to the paper’s abstract, the update reinforces directions associated with high-reward generations and suppresses directions associated with low-reward generations.

RVM includes an optional anchor term intended to control drift from a reference velocity. The authors also say the framework can recover recent methods, including RAM and DiffusionNFT, as special cases. That makes the proposal a unifying formulation as well as a new update rule, although the supplied source does not explain the exact derivations or the conditions under which those equivalences hold.

The researchers report tests across various large-scale diffusion-model reward- tasks. They say RVM was competitive with or outperformed trajectory-based policy-gradient methods while requiring substantially less training cost. For video generation, they report that standard preference rewards can produce visually clean but nearly static outputs; they introduce a dynamic-tracking reward and say it improved motion while also improving overall VBench performance. No numerical scores, model names, training budgets or comparison tables are included in the supplied source.

소스 세부정보: arxiv.org ↗

왜 중요한가요?

The paper argues that direct velocity updates could simplify and reduce the cost of reward for large-scale image and video diffusion systems. Its video experiments also identify a trade-off in standard preference rewards: they may favor visually clean outputs with little motion.

If the paper’s claims generalize, directly optimizing a ’s native velocity representation could make reward easier to scale. The practical importance is not merely a new loss function: the method targets the computational and algorithmic overhead associated with likelihood-based policy optimization. Lower training cost could make preference or task-specific adaptation more accessible to teams working with large image and video generators.

The paper also focuses attention on the design of the reward itself. In the authors’ video experiments, a reward emphasizing visual cleanliness reportedly favored outputs with little movement. Their dynamic-tracking reward is presented as a way to measure motion more effectively, while still improving the paper’s overall VBench result. This illustrates that improving a generative system depends not only on the update rule but also on what the optimization process is asked to value.

The optional anchor term matters because reward optimization can move a model away from a reference behavior. The source says the anchor controls drift from a reference velocity, but it does not state how that control affects quality, diversity, stability or the risk of overfitting to a reward. Those details will determine whether RVM is useful for controlled adaptation rather than only for benchmark optimization.

The result remains a research claim from one arXiv paper, not evidence that diffusion-model training has broadly changed. The supplied abstract does not identify the tested models, datasets, reward magnitudes, compute savings, statistical variation or failure cases. It also does not establish whether the method works equally well for images, video and other diffusion applications, or whether its advantages depend on particular reward designs.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
대화형 개념 확인+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

다음에 무엇을 볼 것인가

The main questions are whether RVM’s reported cost and quality advantages hold across independently reproduced experiments, reward functions, model sizes and generation tasks, and whether the method can avoid optimizing for narrow or misleading rewards. The supplied source does not provide numerical results, implementation details, compute budgets or evidence of external validation.

Reproduction should focus first on the paper’s comparison with trajectory-based policy-gradient methods. Useful checks would include the same model and task under matched compute budgets, because “substantially reduced training cost” is meaningful only when the accounting includes all relevant training and evaluation work. The supplied source gives no numerical cost ratio, so the size of the claimed advantage is unknown.

The reward design deserves separate evaluation from the velocity update. The authors say that, once the velocity update is simplified, the particular loss variant matters less than reward and anchor design. That claim suggests that improvements could depend heavily on how high- and low-reward generations are labeled or scored. Independent tests should therefore vary the reward functions and measure whether gains persist beyond the paper’s selected objectives.

For video generation, observers should check whether dynamic-tracking rewards improve meaningful motion or merely increase visible change. The source reports better motion and an improved overall VBench result, but it does not provide the underlying metrics or describe the failure modes. Evaluation should examine visual quality, temporal consistency and diversity together, including cases in which motion is undesirable or conflicts with the intended content.

The paper’s broader significance will depend on implementation and validation details not present in the supplied source. Important unknowns include the exact velocity parameterization, anchor settings, model and dataset coverage, training duration, reproducibility of the reported comparisons, and whether RVM remains stable as reward objectives change. Follow-up papers, released code and independent replications would help establish whether the proposal is a general scaling method or a result tied to particular experiments.

관련 가이드 및 퀴즈

AI 모델 설명AI 트레이닝트랜스포머AI의 미래알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.
이것이 유용하다고 생각하시나요?