Back to News
SecurityAI Understanding briefing

Study finds benign reinforcement learning can increase language-model leakage of memorized private data

A new arXiv preprint reports that reinforcement learning on factual data containing no personal information made memorized email addresses easier to extract from tested language models.

By 5 min read
AI-generated editorial illustration accompanying Study finds benign reinforcement learning can increase language-model leakage of memorized private data
The short version

A new arXiv preprint reports that reinforcement learning on factual data containing no personal information made memorized email addresses easier to extract from tested language models.

What happened

Researchers Renfei Zhang and Niloofar Mireshghallah report that reinforcement learning with verifiable rewards on benign factual data increased the accessibility of personally identifiable information already memorized by an instruct model. In tests on DeepSeek-V3.1, verbatim recall@k rose from 0.155 to 0.370, a 2.4-fold increase.

The reported pattern also varied with model size. Across three models ranging from 8 billion to 671 billion parameters, the authors say absolute leakage was greatest in the largest model. This comparison describes how the reported result differed across the models examined in the study. It does not replace the central result or add a separate measurement. The models in the comparison are presented as part of the authors' account of the experiment, with the size range providing the stated scale for that comparison. At the same time, the wording preserves the distinction between absolute leakage and the broader question of how reinforcement learning affected access to memorized information.

At the same time, the models retained their reasoning abilities and refusal rates in the study's evaluations. The authors interpret that combination as evidence that reinforcement learning selectively changed access to memorized information rather than broadly changing the models' behavior. In the study's description, the reported increase in accessibility therefore appears alongside retained performance and refusal behavior. The point is not that every aspect of the models changed, but that the reported privacy-related outcome changed while the cited evaluations remained retained. This is the specific combination the authors use when characterizing what the reinforcement-learning process did in the experiments.

The source does not provide enough detail here to independently assess the full training setup, evaluation sample sizes or the exact identities of all three models. That limitation applies to the level of detail available in the source for interpreting the comparison. The reported model-size pattern, the retained reasoning abilities and refusal rates, and the authors' interpretation are all described in the draft. The missing information means the account does not independently establish additional details about the setup or evaluation beyond what is stated. The result can therefore be reported with its stated measurements and interpretation while keeping the source's unresolved details visible.

Read the primary source: arxiv.org

Why it matters

The finding suggests that privacy risk can change during model fine-tuning even when the new training data contains no personal information. A model may retain its reasoning performance and refusal behavior while becoming more likely to surface memorized private data.

The study also complicates the use of refusal rates and general capability tests as privacy indicators. The authors report that refusal rates and reasoning abilities were retained while leakage increased. Read together, those statements describe a mismatch between the behavior measured by broad capability or refusal checks and the accessibility of memorized information. The significance claimed by the study comes from that coexistence: the cited checks did not show a corresponding loss of reasoning or refusal behavior even as the reported leakage measure increased. The concern is therefore about what those checks may leave unmeasured.

That suggests a model can appear stable on broad safety or performance checks while becoming more willing or able to reveal information embedded in its parameters. The statement is framed as a suggestion from the reported finding, not as a claim that every model or deployed system behaves this way. It identifies why the reported result matters for evaluation: stable-looking results in one set of checks may coexist with a different result on access to memorized data. The study's relevance follows from this possible separation between general performance, refusal behavior and the specific leakage outcome described by the authors.

The finding is consequential, but it remains a preprint result from the experiments described by its authors, not a demonstrated failure across all deployed language models. This qualification limits the scope of the claim while preserving its importance. The draft does not present the result as proof that all deployed systems have the same behavior. It presents a preprint finding, reports what the authors observed in their experiments, and identifies the privacy implication that follows from those observations. The distinction between a consequential result and a demonstrated failure across all deployed language models is part of the finding's proper context.

What to watch next

The result needs to be tested across more models, training methods, datasets and privacy attacks. Important unknowns include how broadly the effect generalizes, whether deployed systems are exposed to the same risk, and which safeguards can detect or prevent the change in access to memorized information.

The public impact is still uncertain because the source does not show that a deployed service has leaked private data through this mechanism. That uncertainty concerns the distance between the experiments described in the preprint and real-world service behavior. The reported mechanism involves access to memorized private data after benign reinforcement learning, but the source does not show a deployed-service incident caused by it. The absence of that demonstration does not remove the reported result; it defines what remains unresolved when considering its public impact. The draft therefore keeps the experimental finding separate from a claim about an observed service failure.

It also does not specify how much benign training is required, whether the effect can be reversed, or which controls would reliably block extraction without damaging useful capabilities. These are stated unknowns about the conditions, reversibility and mitigation of the reported effect. They matter because the source does not provide the information needed to determine how the change in accessibility would behave under different amounts of training or under attempted safeguards. The wording also preserves the tradeoff identified in the draft: controls would need to block extraction while avoiding damage to useful capabilities. No additional threshold, reversal method or control is supplied here.

Those unknowns should temper the paper's adversarial implication while keeping the core result in view: according to the authors, training that never touches private data can nevertheless alter how accessible memorized private data becomes. The caution and the core result belong together. The uncertainty about public impact, training requirements, reversibility and controls limits how broadly the implication should be applied. At the same time, the authors' reported claim remains the central point to watch: benign training data containing no personal information may still change access to private information already memorized by a model. The draft does not extend that claim beyond the authors' account.

Related guides & quizzes

AI Models ExplainedAI TrainingAI EthicsFuture of AITest what you know — try a free AI quizLook up an AI term in our glossary
Found this useful?