Назад към Новини
ИновацияAI Understanding брифинг

The Deontic Gap: Large Language Models and the Modal Language of Obligation

Modal auxiliaries such as must, should, and have to mark necessity and obligation within the contexts of speaker authority and interpersonal stance.

6 min readRead the primary source
Source-page capture accompanying The Deontic Gap: Large Language Models and the Modal Language of Obligation
Документ с първичен източникИзточникът е записан
Издател
arxiv.org
Изходна връзка
arxiv.orghttps://arxiv.org/abs/2608.18144
Тип източник
Първичен документ — официално съобщение, документ, документ или първа страна, която четем директно.
КонтекстРазберете това за 60 секунди

Започнете тук

Ключови термини

Голям езиков модел (LLM)
Езиков модел, обучен върху масивни текстови корпуси за генериране и анализиране на текст.
Бенчмарк
Стандартизиран тест или набор от данни, използвани за измерване и сравняване на ефективността на модела.
Тествайте себе сиКакво е AI? Тест

Какво стана

Researchers examined whether large language models (LLMs) reproduce contemporary human patterns of deontic modal usage. They found that AI-generated text consistently underuses positive deontic modals (must, should, have to, had to) relative to contemporary humans. Historical comparison with the Google Books Ngram corpus (1920-2022) showed that AI modal frequencies fall within the range of formal published English, whereas contemporary human modal rates in informal digital contexts often exceed twentieth-century book baselines.

The researchers analyzed three primary corpora, an external , two controlled replications, and a naturalistic eleven-model replication to examine LLM modal usage. This design keeps the comparison centered on modal usage across the stated sources. It also separates the main analysis from the benchmark and replications, preserving the study's stated structure and keeping the description focused on what was examined.

They found that AI-generated text consistently underuses positive deontic modals (must, should, have to, had to) relative to contemporary humans. That result is stated as a relative pattern, with the comparison directed to contemporary humans. The wording identifies both the direction of the difference and the specific family of modals involved, keeping the finding tied to the study's central question.

Historical comparison with the Google Books Ngram corpus (1920-2022) showed that AI modal frequencies fall within the range of formal published English, whereas contemporary human modal rates in informal digital contexts often exceed twentieth-century book baselines. The comparison therefore places the AI pattern alongside two different reference points: formal published English and informal digital writing. The contrast is part of the reported result, presented here without changing either the period covered or the corpus named by the researchers.

Phrase-level decomposition showed that the AI-human modal gap is concentrated in constructions central to interpersonal stance (should, have to, had to), while AI matches or exceeds humans on need to in instructional and question-answering contexts but not in persuasive student writing. This decomposition narrows the reported gap to the constructions already identified, while retaining the separate context in which need to is described. The sentence keeps the distinction between the instructional and question-answering contexts and persuasive student writing in view.

The findings suggest that LLM modal usage reflects the formal written resources on which these models were trained, while underusing the modal constructions through which contemporary human writers mark immediate, interpersonal obligation. This interpretation connects the observed pattern to the language resources and interpersonal functions named in the study. It describes the result as a matter of what is used more or less often, keeping the emphasis on contemporary obligation and on the formal written resources identified in the findings.

Детайли за източника: arxiv.org

Защо има значение

The findings suggest that LLM modal usage reflects the formal written resources on which these models were trained, while underusing the modal constructions through which contemporary human writers mark immediate, interpersonal obligation.

The study's results have implications for the development and evaluation of LLMs, particularly in terms of their ability to capture and reproduce human-like language use. This significance follows from the reported comparison between AI-generated text and contemporary human patterns. It keeps the importance of the result connected to the specific language behavior under discussion rather than separating it from the study's central observation.

The findings suggest that LLMs may not be able to fully capture the nuances of human language use, particularly in terms of deontic modal usage. The importance of that finding lies in the distinction between producing language and reproducing the modal choices through which obligation is expressed. The wording retains the study's emphasis on nuance and keeps deontic modal usage as the relevant point of comparison.

The study's results have implications for the development of more sophisticated LLMs that can better capture and reproduce human-like language use. In this context, sophistication is linked to the same ability identified throughout the study: capturing and reproducing human-like language use. The implication remains focused on the reported modal pattern, with no change to the result or to the language being evaluated.

The findings suggest that LLMs may need to be trained on a wider range of language data, including informal digital contexts, to better capture human-like language use. This matters because the reported comparison distinguishes formal written resources from the informal digital contexts in which contemporary human modal rates are described. The point remains the relationship between training data and the language patterns that models reproduce.

The study's results have implications for the evaluation of LLMs, particularly in terms of their ability to capture and reproduce human-like language use. Evaluation matters here because the study identifies a particular aspect of language use that can be considered when assessing models. The statement preserves the connection to deontic modal usage and to the human-like language use already named in the findings.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Interactive Concept Check+10 Points
What is AI? Quiz

As use of AI scales up across an organization, what tends to matter most?

Какво да гледате след това

The study's results have implications for the development and evaluation of LLMs, particularly in terms of their ability to capture and reproduce human-like language use.

The study's results have implications for the development and evaluation of LLMs, particularly in terms of their ability to capture and reproduce human-like language use. In practical terms, the central question is whether an LLM can reproduce the same modal choices that appear in the human comparisons described above. The wording keeps that question at the level of language use, linking development and evaluation to the reported pattern without adding a separate result.

The findings suggest that LLMs may not be able to fully capture the nuances of human language use, particularly in terms of deontic modal usage. The concern is therefore about nuance within the usage pattern already reported. This identifies a specific place where the study asks how closely that production aligns with contemporary human usage, while keeping the focus on the deontic modal distinction.

The study's results have implications for the development of more sophisticated LLMs that can better capture and reproduce human-like language use. The proposed direction follows directly from the stated implication: greater sophistication is considered in relation to the ability to capture and reproduce the human-like language use discussed in the study. The focus remains the same modal question and does not introduce a different .

The findings suggest that LLMs may need to be trained on a wider range of language data, including informal digital contexts, to better capture human-like language use. The relevant training question is framed around the kinds of language data represented in the study's comparison. The emphasis on informal digital contexts follows the reported contrast with formal published English, while the purpose remains improved capture of the usage at issue.

The study's results have implications for the evaluation of LLMs, particularly in terms of their ability to capture and reproduce human-like language use. Evaluation is included alongside development because the reported pattern matters both for what models produce and for how that production is assessed. The statement keeps the implication tied to deontic modal usage and to the human-like language use named in the study.

Свързани ръководства и викторини

Какво е AI?ChatGPT и LLMЕтика на ИИAI агентиОбяснени модели на AIТрансформърсБъдещето на ИИAI обучениеPrompt EngineeringТествайте какво знаете — опитайте безплатен тест с изкуствен интелектПотърсете термин за AI в нашия речник
Намирате ли това за полезно?