뉴스로 돌아가기
혁신AI Understanding 브리핑

연구에 따르면 자동 창의성 점수는 LLM 스토리에 대한 인간의 판단과 다를 수 있습니다.

새로운 arXiv 사전 인쇄 보고서에 따르면 자동화된 지표와 LLM 심사위원은 단편 소설의 창의성에 대한 인간의 평가와 일치하지 못하는 경우가 많으며 심사위원은 체계적으로 AI 생성 작문 스타일을 선호합니다.

5 min readRead the primary source
Source-provided image accompanying Study finds automatic creativity scores can diverge from human judgments of LLM stories
기본 소스 문서녹음된 소스
출판사
arxiv.org
소스 링크
arxiv.orghttps://arxiv.org/abs/2608.23705
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

대형 언어 모델(LLM)
텍스트를 생성하고 분석하기 위해 대규모 텍스트 말뭉치를 학습한 언어 모델입니다.
정밀도
예측된 긍정 중 실제로 정확한 비율입니다.
데이터세트
학습, 검증 또는 테스트에 사용되는 구조화된 또는 구조화되지 않은 예제 모음입니다.
자신을 테스트해 보세요AI 모델 설명 퀴즈

무슨 일이 일어났나요?

A preprint submitted to arXiv on Aug. 24 examines whether automatic systems can reliably evaluate creativity in stories generated by large language models. The authors compare human assessments with objective metrics and evaluations produced by LLM-based judges.

The paper, titled “The Limits of Automatic Evaluation of Creativity in Large Language Models,” investigates the gap between computational measures of creativity and human judgment. The authors use short stories from the WritingPrompts , including stories written by humans and stories generated by AI systems, and collect human evaluations across 11 dimensions of creativity. The source does not identify the evaluators or provide the number of stories in its abstract. In that sense, the study’s comparison is about the relationship between different kinds of evaluation, not about replacing one final definition of creativity with another. The abstract frames the work as an investigation of reliability and alignment.

Those human assessments are compared with two broad classes of automated evaluation: objective metrics and LLM-as-a-Judge systems. The paper reports “substantial misalignment” between the automated results and human assessments. It also reports that widely used automatic metrics showed near-zero correlation with human judgments for both human-authored and AI-generated stories. The result is presented at the level of correspondence between scores and assessments. It does not, in the supplied text, identify a single metric as responsible for the mismatch, so the broad category of objective metrics should not be read as a judgment on every metric.

The most specific finding concerns LLM-based judges. According to the abstract, these judges systematically preferred AI-generated stories, favoring their stylistic characteristics over unpredictability and other qualities associated with human-authored texts. The source presents this as the result of the authors’ experiments; it does not establish that every LLM judge behaves this way or that the result applies to all creative writing. That qualification keeps the finding tied to the reported comparison. It also separates a preference observed in the experiment from a general conclusion about the literary value of either source, which the abstract does not make.

소스 세부정보: arxiv.org ↗

왜 중요한가요?

The findings challenge a common assumption that creative quality can be measured reliably by automated scoring alone. If evaluation systems reward stylistic signals associated with AI-generated writing while missing qualities human readers value, they could distort comparisons between people and language models.

Creativity is often treated as a single quality that can be summarized by a score, but the paper describes it as multidimensional and subjective. Its reported near-zero correlations suggest that a metric can produce a number without capturing the aspects of creative work that human readers consider important. That distinction matters when automated evaluations are used to compare models, guide development, or assess generated content. It also means that the apparent of an automated result should be considered separately from whether the result reflects the qualities readers are being asked to assess.

The reported preference for AI-generated stories raises a specific measurement concern. If an LLM judge rewards recognizable stylistic patterns more readily than surprise or unpredictability, a system could appear more creative because it matches the judge’s preferences, not because it produces writing that human readers value more. The source does not describe a deployment or a documented institutional decision, so these are practical implications of the study’s findings rather than reported real-world consequences. The concern is therefore about how an evaluation may shape interpretation, especially when its score is treated as a direct summary of creative quality.

The work is also relevant to claims that language models are approaching or exceeding human performance in creative tasks. A strong score from an automated evaluator is not necessarily evidence of broad creative superiority if the evaluator is poorly aligned with human judgment. At the same time, the source does not say that humans are reliable or unbiased judges, nor does it show that AI-generated stories are less creative overall. It reports a disagreement among evaluation methods. That disagreement leaves room for a more careful reading of comparative results, in which the evaluator’s behavior is treated as part of the evidence rather than as a neutral backdrop.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
대화형 개념 확인+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

다음에 무엇을 볼 것인가

The paper is a preprint, and the source text does not provide details about the evaluators, judge models, prompts, statistical methods, or the size and composition of the story sample. Follow-up work should test whether the reported mismatch holds across genres, languages, models, and independent groups of readers.

The main unknown is how the experiments were conducted in detail. The provided source does not state how many human and AI-generated stories were assessed, which language models produced the AI stories, how the 11 creativity dimensions were defined, or how human ratings were aggregated. Those details are important for judging the strength and scope of the reported conclusions. They would also help readers understand how closely the reported comparison reflects the conditions under which the results might later be interpreted or reproduced.

Further scrutiny should examine the objective metrics and LLM judges used in the comparison. The abstract does not name them, describe their prompts, or indicate whether the judges were tested for reliability across repeated evaluations. Results could depend on the choice of metric, judge model, instructions, story length, or writing style. Without those particulars, the broad finding is informative but difficult to connect to a specific evaluation setup. Clarifying the setup would make it easier to distinguish a general limitation from a result tied to particular tools or instructions.

Replication will show whether the reported pattern extends beyond the WritingPrompts and short stories. Useful tests would include other genres, languages, model families, human reader populations, and creative formats. The source also leaves open whether improved evaluation methods can better reflect human judgments, or whether some aspects of creativity will remain resistant to any single automatic score. Those open questions define the next stage of interpretation: determining how durable the mismatch is and what kinds of assessment can address it.

관련 가이드 및 퀴즈

AI 모델 설명ChatGPT와 LLMAI 윤리AI 트레이닝알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.
이것이 유용하다고 생각하시나요?