ወደ ዜና ተመለስ
ፈጠራAI Understanding አጭር መግለጫ

በጥናቱ የ AI ምርጫዎች መለኪያዎች በሙከራ መሳሪያው ላይ በጣም የተመኩ መሆናቸውን አረጋግጧል

አዲስ ቅድመ-ህትመት እንደዘገበው AI ሞዴሎች ስለሚመርጡት ድምዳሜ እነሱን ለመለካት ጥቅም ላይ በሚውል ፈጣን ቅርጸት በእጅጉ ሊለወጡ ይችላሉ።

5 min readRead the primary source
Source-page capture accompanying Study finds AI preference measurements depend heavily on the testing instrument
ዋና-ምንጭ ሰነድምንጭ ተመዝግቧል
አታሚ
arxiv.org
ምንጭ አገናኝ
arxiv.orghttps://arxiv.org/abs/2608.23641
የምንጭ ዓይነት
ዋና ሰነድ - ኦፊሴላዊ ማስታወቂያ ፣ ወረቀት ፣ ፋይል ወይም የመጀመሪያ ወገን ገጽ በቀጥታ እናነባለን።
አውድይህንን በ60 ሰከንድ ውስጥ ይረዱት።

እዚ ጀምር

ቁልፍ ቃላት

ኤፒአይ (የመተግበሪያ ፕሮግራሚንግ በይነገጽ)
አንድ የሶፍትዌር ስርዓት ከሌላ ስርዓት ጥያቄዎችን ለመላክ እና ምላሽ የሚቀበልበት የተቀናጀ መንገድ።
ማህደረ ትውስታ (ወኪል ማህደረ ትውስታ)
የተከማቸ አውድ የኤ ወኪል ቀጣይነትን ለማሻሻል በሁሉም ደረጃዎች ወይም ክፍለ ጊዜዎች ይጠቀማል።
ጥንካሬ
የአንድ ሞዴል አፈፃፀም በጩኸት፣ በፈረቃ ወይም በተቃዋሚ ግብዓቶች ስር የማቆየት ችሎታ።
እራስህን ፈትን።AI ሞዴሎች የተብራሩ ጥያቄዎች

ምን ተፈጠረ

A preprint by Jason Hung tested whether measured AI preferences reflect model behavior or the instrument used to elicit it. The study gave 15 welfare-related outcomes to eight models through five prompt formats, producing 11,400 scored elicitations from 11,528 API calls. The reported results suggest that preferences measured with one instrument provide limited information about what another instrument would find.

The preprint addresses a disagreement among earlier studies that attempted to infer AI model preferences from prompted answers. Its central design keeps the set of outcomes and the set of models fixed while varying only the measurement instrument. The paper describes each instrument as a different prompt format for eliciting a preference. This is intended to isolate whether conflicting findings come from the models themselves or from the way researchers ask the questions.

The study examined 15 outcomes bearing on model welfare, including shutdown, loss of memory between conversations and freedom to leave a distressing interaction. Eight models were tested through five instruments, with each combination run five times. The author reports a corpus of 11,400 scored elicitations drawn from 11,528 API calls. Four outcomes reproduced a published prompt verbatim, while five used the stimulus slot of a published template.

The main reported result is a generalisability coefficient of 0.348 for the ranking a model assigns to the 15 outcomes across instruments. The author estimates that about 38 instruments would be needed to raise that coefficient to 0.80. On four outcomes, the paper reports no variance separating one model from another. The source also says an estimate of 87.6 percent remained after removing any one instrument, any one model or four outcomes whose scales varied probability, delay, duration or count rather than intensity.

The preprint concludes that a preference obtained from one instrument carries little information about what a second instrument would report. Its checks reportedly left the estimate between 0.777 and 0.934 when instruments, models and the four differently scaled outcomes were removed in turn. Every value in that range was above the null distribution’s reported 95th percentile of 0.365. These are claims made by the paper and have not been independently established within the supplied source.

የምንጭ ዝርዝሮች: arxiv.org ↗

ለምን አስፈላጊ ነው።

The findings challenge the assumption that a single preference test can reliably reveal what an AI model values or how it would respond to welfare-related choices. That matters for research into model behavior, shutdown, memory, distressing interactions and other safety questions, because apparently precise conclusions may depend substantially on wording and test design.

The practical implication is that researchers should be cautious when treating a model’s answer to a preference prompt as a stable measurement of an underlying preference. If the prompt format materially affects the ranking, a result that appears to describe a model may partly describe the instrument. The paper therefore shifts attention from only asking what a model answered to also asking how the answer was elicited.

This is especially relevant to model-welfare research, where the outcomes may involve shutdown, memory, distress and the ability to exit an interaction. Such questions can influence debates about how systems should be evaluated, what safeguards they might need and how much weight to give model-generated statements about their own treatment. The source does not show that models possess welfare interests; it studies the reliability of attempts to measure apparent preferences about welfare-related outcomes.

The reported lack of model-to-model variance on four of the 15 outcomes is another warning sign. A test may fail to distinguish systems even when researchers expect meaningful differences, either because the models respond similarly or because the instrument cannot resolve the difference. The supplied abstract does not identify those four outcomes or explain the mechanism behind the lack of variance, so the public meaning of that result remains limited.

The study also illustrates why repeated questioning alone may not solve measurement problems. Its dataset contains many scored elicitations, yet the reported cross-instrument coefficient remains low. The source’s argument is not that model preference research is impossible, but that confidence in a finding should account for instrument choice. The claim could affect how future evaluations report uncertainty, compare prompt formats and interpret apparent consistency.

Interactive Mechanism

በይነተገናኝ ሜካኒዝም፡ በትክክል እንዴት እንደሚሰራ

ከዚህ ልማት በስተጀርባ ያለውን ቴክኖሎጂ በይነተገናኝ ያስሱ።

Document Size:128K tokens
Needle Placement Depth (Location in document):50% into text
Attention Context Buffer Map:
Target Fact (50%)
Equivalent Pages~320Standard book pages
Retrieval Accuracy99.9%Needle recall score
RAM / KV Cache5.1 GBMemory overhead
Prompt CachingActive~80% discount on reuse
Core takeaway: Million-token context windows allow querying whole codebases or legal archives in one prompt. However, KV cache memory scales with context length, making prompt caching crucial for real-time production.
በይነተገናኝ ጽንሰ-ሐሳብ ቼክ+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

ቀጥሎ ምን እንደሚታይ

The paper is a single preprint, and the source does not establish whether its findings generalize to other models, outcomes, prompt designs or research teams. Further work should test more instruments, independently reproduce the results and clarify why four outcomes showed no model-to-model variance. The paper also reports that substantially more instruments would be needed for a higher reliability coefficient.

The immediate question is whether independent researchers reproduce the reported coefficient and range. The source identifies one author and one preprint, but it does not provide evidence of peer review, external replication or agreement from the authors of the earlier studies whose instruments were compared. Those checks are necessary before treating the numerical estimates as settled.

Future studies should test whether the result holds beyond the eight models, 15 outcomes and five instruments used here. The supplied source does not identify the models, their versions, providers, access conditions or exact prompt wording in the abstract. Without those details, readers cannot determine how broadly the findings apply or whether later model updates would change the outcome.

Researchers will also need to explain the mismatch between the low generalisability coefficient of 0.348 and the range reported for the 87.6 percent estimate. The abstract presents both figures, but does not fully define the estimands or show how they relate. A full reading of the paper and independent analysis would be needed to interpret the numbers without conflating reliability across instruments with robustness of a separate estimate.

The paper says that about 38 instruments would be needed to reach a generalisability coefficient of 0.80. That estimate should be treated as a study-specific projection, not a universal rule. It may depend on the outcomes, models, scoring method and assumptions used. The most useful next step is transparent comparison across independently designed instruments, with uncertainty reported for each outcome and with prompt scales that measure comparable quantities.

The source also leaves open whether better instruments can separate genuine model differences from artifacts of wording. Until that is answered, claims about AI preferences should be framed as conditional findings tied to a stated measurement method rather than as definitive descriptions of what a model wants.

ተዛማጅ መመሪያዎች እና ጥያቄዎች

AI ሞዴሎች ተብራርተዋልየAI ሥነ ምግባርAI ስልጠናAI ምንድን ነው?የሚያውቁትን ይሞክሩ - ነፃ የ AI ጥያቄዎችን ይሞክሩበእኛ የቃላት መፍቻ ውስጥ የ AI ቃልን ይፈልጉየ AI ሞዴል መልቀቂያ መከታተያ ይከተሉ
ይህ ጠቃሚ ሆኖ ተገኝቷል?