뉴스로 돌아가기
혁신AI Understanding 브리핑

The Deontic Gap: Large Language Models and the Modal Language of Obligation

Modal auxiliaries such as must, should, and have to mark necessity and obligation within the contexts of speaker authority and interpersonal stance.

6 min readRead the primary source
Source-page capture accompanying The Deontic Gap: Large Language Models and the Modal Language of Obligation
기본 소스 문서녹음된 소스
출판사
arxiv.org
소스 링크
arxiv.orghttps://arxiv.org/abs/2608.18144
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

대형 언어 모델(LLM)
텍스트를 생성하고 분석하기 위해 대규모 텍스트 말뭉치를 학습한 언어 모델입니다.
벤치마크
모델 성능을 측정하고 비교하는 데 사용되는 표준화된 테스트 또는 데이터 세트입니다.
자신을 테스트해 보세요AI란 무엇인가? 퀴즈

무슨 일이 일어났나요?

Researchers examined whether large language models (LLMs) reproduce contemporary human patterns of deontic modal usage. They found that AI-generated text consistently underuses positive deontic modals (must, should, have to, had to) relative to contemporary humans. Historical comparison with the Google Books Ngram corpus (1920-2022) showed that AI modal frequencies fall within the range of formal published English, whereas contemporary human modal rates in informal digital contexts often exceed twentieth-century book baselines.

The researchers analyzed three primary corpora, an external , two controlled replications, and a naturalistic eleven-model replication to examine LLM modal usage. This design keeps the comparison centered on modal usage across the stated sources. It also separates the main analysis from the benchmark and replications, preserving the study's stated structure and keeping the description focused on what was examined.

They found that AI-generated text consistently underuses positive deontic modals (must, should, have to, had to) relative to contemporary humans. That result is stated as a relative pattern, with the comparison directed to contemporary humans. The wording identifies both the direction of the difference and the specific family of modals involved, keeping the finding tied to the study's central question.

Historical comparison with the Google Books Ngram corpus (1920-2022) showed that AI modal frequencies fall within the range of formal published English, whereas contemporary human modal rates in informal digital contexts often exceed twentieth-century book baselines. The comparison therefore places the AI pattern alongside two different reference points: formal published English and informal digital writing. The contrast is part of the reported result, presented here without changing either the period covered or the corpus named by the researchers.

Phrase-level decomposition showed that the AI-human modal gap is concentrated in constructions central to interpersonal stance (should, have to, had to), while AI matches or exceeds humans on need to in instructional and question-answering contexts but not in persuasive student writing. This decomposition narrows the reported gap to the constructions already identified, while retaining the separate context in which need to is described. The sentence keeps the distinction between the instructional and question-answering contexts and persuasive student writing in view.

The findings suggest that LLM modal usage reflects the formal written resources on which these models were trained, while underusing the modal constructions through which contemporary human writers mark immediate, interpersonal obligation. This interpretation connects the observed pattern to the language resources and interpersonal functions named in the study. It describes the result as a matter of what is used more or less often, keeping the emphasis on contemporary obligation and on the formal written resources identified in the findings.

소스 세부정보: arxiv.org

왜 중요한가요?

The findings suggest that LLM modal usage reflects the formal written resources on which these models were trained, while underusing the modal constructions through which contemporary human writers mark immediate, interpersonal obligation.

The study's results have implications for the development and evaluation of LLMs, particularly in terms of their ability to capture and reproduce human-like language use. This significance follows from the reported comparison between AI-generated text and contemporary human patterns. It keeps the importance of the result connected to the specific language behavior under discussion rather than separating it from the study's central observation.

The findings suggest that LLMs may not be able to fully capture the nuances of human language use, particularly in terms of deontic modal usage. The importance of that finding lies in the distinction between producing language and reproducing the modal choices through which obligation is expressed. The wording retains the study's emphasis on nuance and keeps deontic modal usage as the relevant point of comparison.

The study's results have implications for the development of more sophisticated LLMs that can better capture and reproduce human-like language use. In this context, sophistication is linked to the same ability identified throughout the study: capturing and reproducing human-like language use. The implication remains focused on the reported modal pattern, with no change to the result or to the language being evaluated.

The findings suggest that LLMs may need to be trained on a wider range of language data, including informal digital contexts, to better capture human-like language use. This matters because the reported comparison distinguishes formal written resources from the informal digital contexts in which contemporary human modal rates are described. The point remains the relationship between training data and the language patterns that models reproduce.

The study's results have implications for the evaluation of LLMs, particularly in terms of their ability to capture and reproduce human-like language use. Evaluation matters here because the study identifies a particular aspect of language use that can be considered when assessing models. The statement preserves the connection to deontic modal usage and to the human-like language use already named in the findings.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
대화형 개념 확인+10 Points
What is AI? Quiz

As use of AI scales up across an organization, what tends to matter most?

다음에 무엇을 볼 것인가

The study's results have implications for the development and evaluation of LLMs, particularly in terms of their ability to capture and reproduce human-like language use.

The study's results have implications for the development and evaluation of LLMs, particularly in terms of their ability to capture and reproduce human-like language use. In practical terms, the central question is whether an LLM can reproduce the same modal choices that appear in the human comparisons described above. The wording keeps that question at the level of language use, linking development and evaluation to the reported pattern without adding a separate result.

The findings suggest that LLMs may not be able to fully capture the nuances of human language use, particularly in terms of deontic modal usage. The concern is therefore about nuance within the usage pattern already reported. This identifies a specific place where the study asks how closely that production aligns with contemporary human usage, while keeping the focus on the deontic modal distinction.

The study's results have implications for the development of more sophisticated LLMs that can better capture and reproduce human-like language use. The proposed direction follows directly from the stated implication: greater sophistication is considered in relation to the ability to capture and reproduce the human-like language use discussed in the study. The focus remains the same modal question and does not introduce a different .

The findings suggest that LLMs may need to be trained on a wider range of language data, including informal digital contexts, to better capture human-like language use. The relevant training question is framed around the kinds of language data represented in the study's comparison. The emphasis on informal digital contexts follows the reported contrast with formal published English, while the purpose remains improved capture of the usage at issue.

The study's results have implications for the evaluation of LLMs, particularly in terms of their ability to capture and reproduce human-like language use. Evaluation is included alongside development because the reported pattern matters both for what models produce and for how that production is assessed. The statement keeps the implication tied to deontic modal usage and to the human-like language use named in the study.

관련 가이드 및 퀴즈

AI란 무엇인가?ChatGPT와 LLMAI 윤리AI 에이전트AI 모델 설명트랜스포머AI의 미래AI 트레이닝Prompt Engineering알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.
이것이 유용하다고 생각하시나요?