Back to News
InnovationAI Understanding briefing

Study finds AI preference measurements depend heavily on the testing instrument

A new preprint reports that conclusions about what AI models prefer may change substantially with the prompt format used to measure them.

By 5 min read
AI-generated editorial illustration accompanying Study finds AI preference measurements depend heavily on the testing instrument
The short version

A new preprint reports that conclusions about what AI models prefer may change substantially with the prompt format used to measure them.

What happened

A preprint by Jason Hung tested whether measured AI preferences reflect model behavior or the instrument used to elicit it. The study gave 15 welfare-related outcomes to eight models through five prompt formats, producing 11,400 scored elicitations from 11,528 API calls. The reported results suggest that preferences measured with one instrument provide limited information about what another instrument would find.

The preprint addresses a disagreement among earlier studies that attempted to infer AI model preferences from prompted answers. Its central design keeps the set of outcomes and the set of models fixed while varying only the measurement instrument. The paper describes each instrument as a different prompt format for eliciting a preference. This is intended to isolate whether conflicting findings come from the models themselves or from the way researchers ask the questions.

The study examined 15 outcomes bearing on model welfare, including shutdown, loss of memory between conversations and freedom to leave a distressing interaction. Eight models were tested through five instruments, with each combination run five times. The author reports a corpus of 11,400 scored elicitations drawn from 11,528 API calls. Four outcomes reproduced a published prompt verbatim, while five used the stimulus slot of a published template.

The main reported result is a generalisability coefficient of 0.348 for the ranking a model assigns to the 15 outcomes across instruments. The author estimates that about 38 instruments would be needed to raise that coefficient to 0.80. On four outcomes, the paper reports no variance separating one model from another. The source also says an estimate of 87.6 percent remained after removing any one instrument, any one model or four outcomes whose scales varied probability, delay, duration or count rather than intensity.

The preprint concludes that a preference obtained from one instrument carries little information about what a second instrument would report. Its robustness checks reportedly left the estimate between 0.777 and 0.934 when instruments, models and the four differently scaled outcomes were removed in turn. Every value in that range was above the null distribution’s reported 95th percentile of 0.365. These are claims made by the paper and have not been independently established within the supplied source.

Read the primary source: arxiv.org

Why it matters

The findings challenge the assumption that a single preference test can reliably reveal what an AI model values or how it would respond to welfare-related choices. That matters for research into model behavior, shutdown, memory, distressing interactions and other safety questions, because apparently precise conclusions may depend substantially on wording and test design.

The practical implication is that researchers should be cautious when treating a model’s answer to a preference prompt as a stable measurement of an underlying preference. If the prompt format materially affects the ranking, a result that appears to describe a model may partly describe the instrument. The paper therefore shifts attention from only asking what a model answered to also asking how the answer was elicited.

This is especially relevant to model-welfare research, where the outcomes may involve shutdown, memory, distress and the ability to exit an interaction. Such questions can influence debates about how systems should be evaluated, what safeguards they might need and how much weight to give model-generated statements about their own treatment. The source does not show that models possess welfare interests; it studies the reliability of attempts to measure apparent preferences about welfare-related outcomes.

The reported lack of model-to-model variance on four of the 15 outcomes is another warning sign. A test may fail to distinguish systems even when researchers expect meaningful differences, either because the models respond similarly or because the instrument cannot resolve the difference. The supplied abstract does not identify those four outcomes or explain the mechanism behind the lack of variance, so the public meaning of that result remains limited.

The study also illustrates why repeated questioning alone may not solve measurement problems. Its dataset contains many scored elicitations, yet the reported cross-instrument coefficient remains low. The source’s argument is not that model preference research is impossible, but that confidence in a finding should account for instrument choice. The claim could affect how future evaluations report uncertainty, compare prompt formats and interpret apparent consistency.

What to watch next

The paper is a single preprint, and the source does not establish whether its findings generalize to other models, outcomes, prompt designs or research teams. Further work should test more instruments, independently reproduce the results and clarify why four outcomes showed no model-to-model variance. The paper also reports that substantially more instruments would be needed for a higher reliability coefficient.

The immediate question is whether independent researchers reproduce the reported coefficient and robustness range. The source identifies one author and one preprint, but it does not provide evidence of peer review, external replication or agreement from the authors of the earlier studies whose instruments were compared. Those checks are necessary before treating the numerical estimates as settled.

Future studies should test whether the result holds beyond the eight models, 15 outcomes and five instruments used here. The supplied source does not identify the models, their versions, providers, access conditions or exact prompt wording in the abstract. Without those details, readers cannot determine how broadly the findings apply or whether later model updates would change the outcome.

Researchers will also need to explain the mismatch between the low generalisability coefficient of 0.348 and the robustness range reported for the 87.6 percent estimate. The abstract presents both figures, but does not fully define the estimands or show how they relate. A full reading of the paper and independent analysis would be needed to interpret the numbers without conflating reliability across instruments with robustness of a separate estimate.

The paper says that about 38 instruments would be needed to reach a generalisability coefficient of 0.80. That estimate should be treated as a study-specific projection, not a universal rule. It may depend on the outcomes, models, scoring method and assumptions used. The most useful next step is transparent comparison across independently designed instruments, with uncertainty reported for each outcome and with prompt scales that measure comparable quantities.

The source also leaves open whether better instruments can separate genuine model differences from artifacts of wording. Until that is answered, claims about AI preferences should be framed as conditional findings tied to a stated measurement method rather than as definitive descriptions of what a model wants.

Related guides & quizzes

AI Models ExplainedAI EthicsAI TrainingWhat is AI?Test what you know — try a free AI quizLook up an AI term in our glossary
Found this useful?