返回新聞
創新AI Understanding 簡報

研究發現自動創造力分數可能與人類對法學碩士故事的判斷不同

一份新的 arXiv 預印本報告稱,自動化指標和法學碩士評審往往無法與人類對短篇小說創造力的評估相匹配,評審系統地偏愛人工智慧生成的寫作風格。

5 min readRead the primary source
Source-provided image accompanying Study finds automatic creativity scores can diverge from human judgments of LLM stories
主要來源文件來源記錄
出版商
arxiv.org
來源連結
arxiv.orghttps://arxiv.org/abs/2608.23705
來源類型
主要文件-我們直接閱讀的官方公告、文件、文件或第一方頁面。
背景60 秒內了解這一點

從這裡開始

關鍵術語

大語言模型(LLM)
在海量文本語料庫上訓練來產生和分析文本的語言模型。
精確度
實際正確的預測陽性的比例。
數據集
用於訓練、驗證或測試的結構化或非結構化範例的集合。
測試一下自己AI 模型解釋測驗

發生了什麼事

A preprint submitted to arXiv on Aug. 24 examines whether automatic systems can reliably evaluate creativity in stories generated by large language models. The authors compare human assessments with objective metrics and evaluations produced by LLM-based judges.

The paper, titled “The Limits of Automatic Evaluation of Creativity in Large Language Models,” investigates the gap between computational measures of creativity and human judgment. The authors use short stories from the WritingPrompts , including stories written by humans and stories generated by AI systems, and collect human evaluations across 11 dimensions of creativity. The source does not identify the evaluators or provide the number of stories in its abstract. In that sense, the study’s comparison is about the relationship between different kinds of evaluation, not about replacing one final definition of creativity with another. The abstract frames the work as an investigation of reliability and alignment.

Those human assessments are compared with two broad classes of automated evaluation: objective metrics and LLM-as-a-Judge systems. The paper reports “substantial misalignment” between the automated results and human assessments. It also reports that widely used automatic metrics showed near-zero correlation with human judgments for both human-authored and AI-generated stories. The result is presented at the level of correspondence between scores and assessments. It does not, in the supplied text, identify a single metric as responsible for the mismatch, so the broad category of objective metrics should not be read as a judgment on every metric.

The most specific finding concerns LLM-based judges. According to the abstract, these judges systematically preferred AI-generated stories, favoring their stylistic characteristics over unpredictability and other qualities associated with human-authored texts. The source presents this as the result of the authors’ experiments; it does not establish that every LLM judge behaves this way or that the result applies to all creative writing. That qualification keeps the finding tied to the reported comparison. It also separates a preference observed in the experiment from a general conclusion about the literary value of either source, which the abstract does not make.

來源詳情: arxiv.org ↗

為什麼這很重要

The findings challenge a common assumption that creative quality can be measured reliably by automated scoring alone. If evaluation systems reward stylistic signals associated with AI-generated writing while missing qualities human readers value, they could distort comparisons between people and language models.

Creativity is often treated as a single quality that can be summarized by a score, but the paper describes it as multidimensional and subjective. Its reported near-zero correlations suggest that a metric can produce a number without capturing the aspects of creative work that human readers consider important. That distinction matters when automated evaluations are used to compare models, guide development, or assess generated content. It also means that the apparent of an automated result should be considered separately from whether the result reflects the qualities readers are being asked to assess.

The reported preference for AI-generated stories raises a specific measurement concern. If an LLM judge rewards recognizable stylistic patterns more readily than surprise or unpredictability, a system could appear more creative because it matches the judge’s preferences, not because it produces writing that human readers value more. The source does not describe a deployment or a documented institutional decision, so these are practical implications of the study’s findings rather than reported real-world consequences. The concern is therefore about how an evaluation may shape interpretation, especially when its score is treated as a direct summary of creative quality.

The work is also relevant to claims that language models are approaching or exceeding human performance in creative tasks. A strong score from an automated evaluator is not necessarily evidence of broad creative superiority if the evaluator is poorly aligned with human judgment. At the same time, the source does not say that humans are reliable or unbiased judges, nor does it show that AI-generated stories are less creative overall. It reports a disagreement among evaluation methods. That disagreement leaves room for a more careful reading of comparative results, in which the evaluator’s behavior is treated as part of the evidence rather than as a neutral backdrop.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下來看什麼

The paper is a preprint, and the source text does not provide details about the evaluators, judge models, prompts, statistical methods, or the size and composition of the story sample. Follow-up work should test whether the reported mismatch holds across genres, languages, models, and independent groups of readers.

The main unknown is how the experiments were conducted in detail. The provided source does not state how many human and AI-generated stories were assessed, which language models produced the AI stories, how the 11 creativity dimensions were defined, or how human ratings were aggregated. Those details are important for judging the strength and scope of the reported conclusions. They would also help readers understand how closely the reported comparison reflects the conditions under which the results might later be interpreted or reproduced.

Further scrutiny should examine the objective metrics and LLM judges used in the comparison. The abstract does not name them, describe their prompts, or indicate whether the judges were tested for reliability across repeated evaluations. Results could depend on the choice of metric, judge model, instructions, story length, or writing style. Without those particulars, the broad finding is informative but difficult to connect to a specific evaluation setup. Clarifying the setup would make it easier to distinguish a general limitation from a result tied to particular tools or instructions.

Replication will show whether the reported pattern extends beyond the WritingPrompts and short stories. Useful tests would include other genres, languages, model families, human reader populations, and creative formats. The source also leaves open whether improved evaluation methods can better reflect human judgments, or whether some aspects of creativity will remain resistant to any single automatic score. Those open questions define the next stage of interpretation: determining how durable the mismatch is and what kinds of assessment can address it.

相關指引和測驗

人工智慧模型解釋ChatGPT 與大型語言模型AI 倫理人工智慧培訓測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?