Voltar às notícias
InovaçãoInstruções AI Understanding

Study finds automatic creativity scores can diverge from human judgments of LLM stories

A new arXiv preprint reports that automated metrics and LLM judges often fail to match human assessments of creativity in short stories, with judges systematically favoring AI-generated writing styles.

Por 5 min read
AI-generated editorial illustration accompanying Study finds automatic creativity scores can diverge from human judgments of LLM stories
A versão curta

A new arXiv preprint reports that automated metrics and LLM judges often fail to match human assessments of creativity in short stories, with judges systematically favoring AI-generated writing styles.

O que aconteceu

A preprint submitted to arXiv on Aug. 24 examines whether automatic systems can reliably evaluate creativity in stories generated by large language models. The authors compare human assessments with objective metrics and evaluations produced by LLM-based judges.

The paper, titled “The Limits of Automatic Evaluation of Creativity in Large Language Models,” investigates the gap between computational measures of creativity and human judgment. The authors use short stories from the WritingPrompts dataset, including stories written by humans and stories generated by AI systems, and collect human evaluations across 11 dimensions of creativity. The source does not identify the evaluators or provide the number of stories in its abstract. In that sense, the study’s comparison is about the relationship between different kinds of evaluation, not about replacing one final definition of creativity with another. The abstract frames the work as an investigation of reliability and alignment.

Those human assessments are compared with two broad classes of automated evaluation: objective metrics and LLM-as-a-Judge systems. The paper reports “substantial misalignment” between the automated results and human assessments. It also reports that widely used automatic metrics showed near-zero correlation with human judgments for both human-authored and AI-generated stories. The result is presented at the level of correspondence between scores and assessments. It does not, in the supplied text, identify a single metric as responsible for the mismatch, so the broad category of objective metrics should not be read as a judgment on every metric.

The most specific finding concerns LLM-based judges. According to the abstract, these judges systematically preferred AI-generated stories, favoring their stylistic characteristics over unpredictability and other qualities associated with human-authored texts. The source presents this as the result of the authors’ experiments; it does not establish that every LLM judge behaves this way or that the result applies to all creative writing. That qualification keeps the finding tied to the reported comparison. It also separates a preference observed in the experiment from a general conclusion about the literary value of either source, which the abstract does not make.

Leia a fonte primária: arxiv.org

Por que isso importa

The findings challenge a common assumption that creative quality can be measured reliably by automated scoring alone. If evaluation systems reward stylistic signals associated with AI-generated writing while missing qualities human readers value, they could distort comparisons between people and language models.

Creativity is often treated as a single quality that can be summarized by a score, but the paper describes it as multidimensional and subjective. Its reported near-zero correlations suggest that a metric can produce a number without capturing the aspects of creative work that human readers consider important. That distinction matters when automated evaluations are used to compare models, guide development, or assess generated content. It also means that the apparent precision of an automated result should be considered separately from whether the result reflects the qualities readers are being asked to assess.

The reported preference for AI-generated stories raises a specific measurement concern. If an LLM judge rewards recognizable stylistic patterns more readily than surprise or unpredictability, a system could appear more creative because it matches the judge’s preferences, not because it produces writing that human readers value more. The source does not describe a deployment or a documented institutional decision, so these are practical implications of the study’s findings rather than reported real-world consequences. The concern is therefore about how an evaluation may shape interpretation, especially when its score is treated as a direct summary of creative quality.

The work is also relevant to claims that language models are approaching or exceeding human performance in creative tasks. A strong score from an automated evaluator is not necessarily evidence of broad creative superiority if the evaluator is poorly aligned with human judgment. At the same time, the source does not say that humans are reliable or unbiased judges, nor does it show that AI-generated stories are less creative overall. It reports a disagreement among evaluation methods. That disagreement leaves room for a more careful reading of comparative results, in which the evaluator’s behavior is treated as part of the evidence rather than as a neutral backdrop.

O que assistir a seguir

The paper is a preprint, and the source text does not provide details about the evaluators, judge models, prompts, statistical methods, or the size and composition of the story sample. Follow-up work should test whether the reported mismatch holds across genres, languages, models, and independent groups of readers.

The main unknown is how the experiments were conducted in detail. The provided source does not state how many human and AI-generated stories were assessed, which language models produced the AI stories, how the 11 creativity dimensions were defined, or how human ratings were aggregated. Those details are important for judging the strength and scope of the reported conclusions. They would also help readers understand how closely the reported comparison reflects the conditions under which the results might later be interpreted or reproduced.

Further scrutiny should examine the objective metrics and LLM judges used in the comparison. The abstract does not name them, describe their prompts, or indicate whether the judges were tested for reliability across repeated evaluations. Results could depend on the choice of metric, judge model, instructions, story length, or writing style. Without those particulars, the broad finding is informative but difficult to connect to a specific evaluation setup. Clarifying the setup would make it easier to distinguish a general limitation from a result tied to particular tools or instructions.

Replication will show whether the reported pattern extends beyond the WritingPrompts dataset and short stories. Useful tests would include other genres, languages, model families, human reader populations, and creative formats. The source also leaves open whether improved evaluation methods can better reflect human judgments, or whether some aspects of creativity will remain resistant to any single automatic score. Those open questions define the next stage of interpretation: determining how durable the mismatch is and what kinds of assessment can address it.

Guias e questionários relacionados

Modelos de IA explicadosChatGPT e LLMÉtica da IATreinamento de IATeste o que você sabe – experimente um teste gratuito de IAProcure um termo de IA em nosso glossário
Achou isso útil?