뉴스로 돌아가기
혁신AI Understanding 브리핑

NC-GRPO 논문은 잠재 공간 롤아웃 다양성을 통해 더욱 강력한 비전 언어 추론을 보고합니다.

arXiv 사전 인쇄에서는 입력 이미지가 아닌 비전 언어 모델의 숨겨진 표현을 교란시키는 강화 학습 방법인 NC-GRPO를 소개합니다. 저자는 Qwen2.5-VL-7B에서 향상된 도메인 외부 수학적 추론과 환각 견고성을 보고하는 동시에…

5 min readRead the primary source
Primary-source image accompanying NC-GRPO paper reports stronger vision-language reasoning from latent-space rollout diversity
기본 소스 문서녹음된 소스
출판사
arxiv.org
소스 링크
arxiv.orghttps://arxiv.org/abs/2608.21595
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

비전-언어 모델(VLM)
시각적 정보와 텍스트 정보를 공동으로 처리하는 다중 모드 모델입니다.
강화 학습
에이전트가 장기적인 수익을 극대화하는 행동을 학습하는 보상 신호를 통한 교육입니다.
메모리(에이전트 메모리)
AI 에이전트는 연속성을 향상하기 위해 여러 단계 또는 세션에서 사용하는 저장된 컨텍스트입니다.
자신을 테스트해 보세요AI 모델 설명 퀴즈

무슨 일이 일어났나요?

An arXiv paper introduces Noise-Contrastive GRPO, or NC-GRPO, a method for training vision-language models with and verifiable rewards. It adds calibrated Gaussian noise to the model’s last hidden layer during prompt encoding for half of each rollout group, creating alternate reasoning paths without changing the image, reward, objective or inference protocol. Tested on Qwen2.5-VL-7B trained on Geometry3K, the authors report stronger in-domain performance, improved out-of-domain mathematical reasoning across five held-out benchmarks, and better hallucination robustness than vanilla GRPO.

The paper studies with verifiable rewards for vision-language models, where candidate outputs can be checked against a known answer or other objective signal. Its central proposal is Noise-Contrastive GRPO, or NC-GRPO. Instead of increasing diversity by changing decoding temperature or altering the input image in pixel space, the method injects scale-calibrated Gaussian noise into the last hidden layer produced while the prompt is being encoded. Half of the rollouts in each optimization group receive the perturbation, while the remaining branches provide an unperturbed comparison.

The authors describe the perturbation as a way to create branches from a displaced internal departure state. A branch that still reaches the correct answer after that displacement receives reinforcement relative to a branch that is derailed. In the paper’s framing, the model’s sensitivity at that internal branch point becomes a policy-gradient signal. The abstract says this leaves the training objective, reward definition and inference protocol unchanged, and that the method can be integrated into a standard reinforcement-learning pipeline through an approximately 50-line inference-engine change. These are claims made by the authors, not an independently verified implementation assessment.

The reported experiment uses Qwen2.5-VL-7B trained on Geometry3K. Against vanilla GRPO, the paper says NC-GRPO significantly improves out-of-domain mathematical reasoning on five held-out benchmarks, with a pooled McNemar test of p less than or equal to 0.001. It also reports improvements on the in-domain task and on hallucination robustness. The abstract says image-space noise produced a larger average out-of-domain result on perception-heavy benchmarks but regressed on hallucination robustness. Mechanism ablations attribute the effect to independent stochastic diversity rather than simply the amount or direction of noise, while a noise-scale study identifies a tradeoff between reasoning specialization and general capability.

소스 세부정보: arxiv.org ↗

왜 중요한가요?

The work addresses a practical problem in for multimodal models: how to produce sufficiently diverse candidate reasoning paths without distorting the visual input. If the reported results generalize, latent-space diversification could offer a relatively small engineering change for improving reasoning transfer while preserving the original task and reward setup. The findings also suggest that training for mathematical specialization and robustness may involve a controllable tradeoff rather than a single universally optimal noise level.

The research is relevant because rollout diversity is part of how searches for better behavior. If every candidate in an optimization group follows nearly the same internal path, the training signal may offer limited information about which decisions are robust. NC-GRPO tests whether diversity can be introduced inside the model’s representation while leaving the observed image and the task’s external reward untouched. That distinction could matter for vision-language systems, where changing pixels may create artifacts or alter the problem itself rather than testing resilience in reasoning.

The reported contrast with image-space perturbation is potentially useful. The authors say pixel-level noise performed better on average for perception-heavy out-of-domain benchmarks, but harmed hallucination robustness, whereas NC-GRPO improved that robustness in their experiment. This points to different perturbation locations producing different capabilities: input changes may test visual tolerance, while latent changes may encourage reasoning paths that remain stable after internal variation. The source does not show enough detail to determine whether this interpretation is causal beyond the authors’ ablations.

The method’s practical appeal comes from its claimed compatibility with an existing RLVR setup. The abstract says the reward, objective and inference procedure are untouched, and that implementation requires roughly 50 lines of inference-engine changes. If those claims are borne out, researchers could evaluate the technique without redesigning the full training stack. That does not mean the approach is cheap overall: the source gives no compute budget, training duration, memory impact or throughput measurement. Its public significance therefore depends on whether the reported quality gains justify any additional rollout or optimization cost.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
대화형 개념 확인+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

다음에 무엇을 볼 것인가

The main open questions are whether the gains hold across more model sizes, visual tasks and training datasets, and whether they survive independent replication. The source does not provide exact accuracy changes, dataset sizes, per-benchmark results, compute costs or full implementation details. It also does not establish production availability or real-world reliability. Follow-up work should test whether latent perturbations improve tasks beyond mathematics and whether the reported hallucination benefits remain under unfamiliar images, prompts and deployment conditions.

The evidence remains limited to one source and one named training setup: Qwen2.5-VL-7B trained on Geometry3K. The abstract does not identify the five held-out benchmarks, report their individual scores, give sample counts or state the size of the hallucination-robustness improvement. It also does not say whether the comparisons used matched compute, the same number of rollouts or identical hyperparameter tuning effort. Those details are necessary to judge whether the result reflects a broadly useful method or a favorable experimental configuration.

Replication should examine model scale, model family, task type and modality. The authors describe NC-GRPO as modality-agnostic, but the source provides no experiment demonstrating that claim outside the reported vision-language setting. Further tests could include perception-heavy tasks, compositional visual reasoning, nonmathematical questions and prompts designed to induce unsupported answers. Researchers should also separate gains from the perturbation itself from gains caused by additional stochastic sampling or other training differences.

The paper’s noise-scale result makes deployment behavior an important unknown. The abstract describes a dial between reasoning specialization and general capability, but does not specify how practitioners should select that setting or how stable the tradeoff is across tasks. There is no claim of a released model, product integration or production evaluation. Readers should treat the findings as an early research result from an arXiv preprint until the exact metrics, code or implementation details and independent replications are available.

관련 가이드 및 퀴즈

AI 모델 설명AI 트레이닝트랜스포머ChatGPT와 LLM알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.
이것이 유용하다고 생각하시나요?