Quay lại Tin tức
Đổi mớiAI Understanding tóm tắt

The Deontic Gap: Large Language Models and the Modal Language of Obligation

Modal auxiliaries such as must, should, and have to mark necessity and obligation within the contexts of speaker authority and interpersonal stance.

6 min readRead the primary source
Source-page capture accompanying The Deontic Gap: Large Language Models and the Modal Language of Obligation
Tài liệu nguồn chínhNguồn đã ghi
Nhà xuất bản
arxiv.org
Liên kết nguồn
arxiv.orghttps://arxiv.org/abs/2608.18144
Loại nguồn
Tài liệu chính - một thông báo chính thức, giấy tờ, hồ sơ hoặc trang của bên thứ nhất mà chúng tôi đọc trực tiếp.
Bối cảnhHiểu điều này trong 60 giây

Bắt đầu ở đây

Thuật ngữ chính

Mô hình ngôn ngữ lớn (LLM)
Một mô hình ngôn ngữ được đào tạo trên kho văn bản lớn để tạo và phân tích văn bản.
Điểm chuẩn
Một bài kiểm tra hoặc tập dữ liệu được tiêu chuẩn hóa dùng để đo lường và so sánh hiệu suất của mô hình.
Tự kiểm traAI là gì? Câu đố

Chuyện gì đã xảy ra

Researchers examined whether large language models (LLMs) reproduce contemporary human patterns of deontic modal usage. They found that AI-generated text consistently underuses positive deontic modals (must, should, have to, had to) relative to contemporary humans. Historical comparison with the Google Books Ngram corpus (1920-2022) showed that AI modal frequencies fall within the range of formal published English, whereas contemporary human modal rates in informal digital contexts often exceed twentieth-century book baselines.

The researchers analyzed three primary corpora, an external , two controlled replications, and a naturalistic eleven-model replication to examine LLM modal usage. This design keeps the comparison centered on modal usage across the stated sources. It also separates the main analysis from the benchmark and replications, preserving the study's stated structure and keeping the description focused on what was examined.

They found that AI-generated text consistently underuses positive deontic modals (must, should, have to, had to) relative to contemporary humans. That result is stated as a relative pattern, with the comparison directed to contemporary humans. The wording identifies both the direction of the difference and the specific family of modals involved, keeping the finding tied to the study's central question.

Historical comparison with the Google Books Ngram corpus (1920-2022) showed that AI modal frequencies fall within the range of formal published English, whereas contemporary human modal rates in informal digital contexts often exceed twentieth-century book baselines. The comparison therefore places the AI pattern alongside two different reference points: formal published English and informal digital writing. The contrast is part of the reported result, presented here without changing either the period covered or the corpus named by the researchers.

Phrase-level decomposition showed that the AI-human modal gap is concentrated in constructions central to interpersonal stance (should, have to, had to), while AI matches or exceeds humans on need to in instructional and question-answering contexts but not in persuasive student writing. This decomposition narrows the reported gap to the constructions already identified, while retaining the separate context in which need to is described. The sentence keeps the distinction between the instructional and question-answering contexts and persuasive student writing in view.

The findings suggest that LLM modal usage reflects the formal written resources on which these models were trained, while underusing the modal constructions through which contemporary human writers mark immediate, interpersonal obligation. This interpretation connects the observed pattern to the language resources and interpersonal functions named in the study. It describes the result as a matter of what is used more or less often, keeping the emphasis on contemporary obligation and on the formal written resources identified in the findings.

Chi tiết nguồn: arxiv.org

Tại sao nó quan trọng

The findings suggest that LLM modal usage reflects the formal written resources on which these models were trained, while underusing the modal constructions through which contemporary human writers mark immediate, interpersonal obligation.

The study's results have implications for the development and evaluation of LLMs, particularly in terms of their ability to capture and reproduce human-like language use. This significance follows from the reported comparison between AI-generated text and contemporary human patterns. It keeps the importance of the result connected to the specific language behavior under discussion rather than separating it from the study's central observation.

The findings suggest that LLMs may not be able to fully capture the nuances of human language use, particularly in terms of deontic modal usage. The importance of that finding lies in the distinction between producing language and reproducing the modal choices through which obligation is expressed. The wording retains the study's emphasis on nuance and keeps deontic modal usage as the relevant point of comparison.

The study's results have implications for the development of more sophisticated LLMs that can better capture and reproduce human-like language use. In this context, sophistication is linked to the same ability identified throughout the study: capturing and reproducing human-like language use. The implication remains focused on the reported modal pattern, with no change to the result or to the language being evaluated.

The findings suggest that LLMs may need to be trained on a wider range of language data, including informal digital contexts, to better capture human-like language use. This matters because the reported comparison distinguishes formal written resources from the informal digital contexts in which contemporary human modal rates are described. The point remains the relationship between training data and the language patterns that models reproduce.

The study's results have implications for the evaluation of LLMs, particularly in terms of their ability to capture and reproduce human-like language use. Evaluation matters here because the study identifies a particular aspect of language use that can be considered when assessing models. The statement preserves the connection to deontic modal usage and to the human-like language use already named in the findings.

Interactive Mechanism

Cơ chế tương tác: Nó thực sự hoạt động như thế nào

Khám phá công nghệ cơ bản đằng sau sự phát triển này một cách tương tác.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Kiểm tra khái niệm tương tác+10 Points
What is AI? Quiz

Which description best fits "narrow AI", the kind of AI in use today?

Xem gì tiếp theo

The study's results have implications for the development and evaluation of LLMs, particularly in terms of their ability to capture and reproduce human-like language use.

The study's results have implications for the development and evaluation of LLMs, particularly in terms of their ability to capture and reproduce human-like language use. In practical terms, the central question is whether an LLM can reproduce the same modal choices that appear in the human comparisons described above. The wording keeps that question at the level of language use, linking development and evaluation to the reported pattern without adding a separate result.

The findings suggest that LLMs may not be able to fully capture the nuances of human language use, particularly in terms of deontic modal usage. The concern is therefore about nuance within the usage pattern already reported. This identifies a specific place where the study asks how closely that production aligns with contemporary human usage, while keeping the focus on the deontic modal distinction.

The study's results have implications for the development of more sophisticated LLMs that can better capture and reproduce human-like language use. The proposed direction follows directly from the stated implication: greater sophistication is considered in relation to the ability to capture and reproduce the human-like language use discussed in the study. The focus remains the same modal question and does not introduce a different .

The findings suggest that LLMs may need to be trained on a wider range of language data, including informal digital contexts, to better capture human-like language use. The relevant training question is framed around the kinds of language data represented in the study's comparison. The emphasis on informal digital contexts follows the reported contrast with formal published English, while the purpose remains improved capture of the usage at issue.

The study's results have implications for the evaluation of LLMs, particularly in terms of their ability to capture and reproduce human-like language use. Evaluation is included alongside development because the reported pattern matters both for what models produce and for how that production is assessed. The statement keeps the implication tied to deontic modal usage and to the human-like language use named in the study.

Hướng dẫn và câu hỏi liên quan

AI là gì?ChatGPT & LLMĐạo đức AIĐại lý AIGiải thích về mô hình AIMáy biến ápTương lai của AIĐào tạo AIPrompt EngineeringKiểm tra những gì bạn biết — thử một bài kiểm tra AI miễn phíTra cứu một thuật ngữ AI trong bảng thuật ngữ của chúng tôi
Tìm thấy điều này hữu ích?