Επιστροφή στις Ειδήσεις
ΚαινοτομίαAI Understanding ενημέρωση

Η μελέτη διαπιστώνει ότι οι αυτόματες βαθμολογίες δημιουργικότητας μπορεί να αποκλίνουν από τις ανθρώπινες κρίσεις για τις ιστορίες LLM

Μια νέα προεκτύπωση του arXiv αναφέρει ότι οι αυτοματοποιημένες μετρήσεις και οι κριτές LLM συχνά αποτυγχάνουν να ταιριάζουν με τις ανθρώπινες αξιολογήσεις της δημιουργικότητας στα διηγήματα, με τους κριτές να προτιμούν συστηματικά τα στυλ γραφής που δημιουργούνται από την τεχνητή νοημοσύνη.

5 min readRead the primary source
Source-provided image accompanying Study finds automatic creativity scores can diverge from human judgments of LLM stories
Έγγραφο κύριας πηγήςΗ πηγή καταγράφηκε
Εκδότης
arxiv.org
Σύνδεσμος πηγής
arxiv.orghttps://arxiv.org/abs/2608.23705
Τύπος πηγής
Κύριο έγγραφο — μια επίσημη ανακοίνωση, χαρτί, αρχειοθέτηση ή σελίδα πρώτου μέρους που διαβάζουμε απευθείας.
ΠλαίσιοΚαταλάβετε αυτό σε 60 δευτερόλεπτα

Ξεκινήστε εδώ

Βασικοί όροι

Μεγάλο μοντέλο γλώσσας (LLM)
Ένα μοντέλο γλώσσας εκπαιδευμένο σε τεράστια σώματα κειμένου για τη δημιουργία και ανάλυση κειμένου.
Ακρίβεια
Το ποσοστό των προβλεπόμενων θετικών που είναι πραγματικά σωστά.
Σύνολο δεδομένων
Μια συλλογή δομημένων ή μη παραδειγμάτων που χρησιμοποιούνται για εκπαίδευση, επικύρωση ή δοκιμή.
Δοκιμάστε τον εαυτό σαςΕξηγημένο Κουίζ Μοντέλων AI

Τι έγινε

A preprint submitted to arXiv on Aug. 24 examines whether automatic systems can reliably evaluate creativity in stories generated by large language models. The authors compare human assessments with objective metrics and evaluations produced by LLM-based judges.

The paper, titled “The Limits of Automatic Evaluation of Creativity in Large Language Models,” investigates the gap between computational measures of creativity and human judgment. The authors use short stories from the WritingPrompts , including stories written by humans and stories generated by AI systems, and collect human evaluations across 11 dimensions of creativity. The source does not identify the evaluators or provide the number of stories in its abstract. In that sense, the study’s comparison is about the relationship between different kinds of evaluation, not about replacing one final definition of creativity with another. The abstract frames the work as an investigation of reliability and alignment.

Those human assessments are compared with two broad classes of automated evaluation: objective metrics and LLM-as-a-Judge systems. The paper reports “substantial misalignment” between the automated results and human assessments. It also reports that widely used automatic metrics showed near-zero correlation with human judgments for both human-authored and AI-generated stories. The result is presented at the level of correspondence between scores and assessments. It does not, in the supplied text, identify a single metric as responsible for the mismatch, so the broad category of objective metrics should not be read as a judgment on every metric.

The most specific finding concerns LLM-based judges. According to the abstract, these judges systematically preferred AI-generated stories, favoring their stylistic characteristics over unpredictability and other qualities associated with human-authored texts. The source presents this as the result of the authors’ experiments; it does not establish that every LLM judge behaves this way or that the result applies to all creative writing. That qualification keeps the finding tied to the reported comparison. It also separates a preference observed in the experiment from a general conclusion about the literary value of either source, which the abstract does not make.

Στοιχεία πηγής: arxiv.org ↗

Γιατί έχει σημασία

The findings challenge a common assumption that creative quality can be measured reliably by automated scoring alone. If evaluation systems reward stylistic signals associated with AI-generated writing while missing qualities human readers value, they could distort comparisons between people and language models.

Creativity is often treated as a single quality that can be summarized by a score, but the paper describes it as multidimensional and subjective. Its reported near-zero correlations suggest that a metric can produce a number without capturing the aspects of creative work that human readers consider important. That distinction matters when automated evaluations are used to compare models, guide development, or assess generated content. It also means that the apparent of an automated result should be considered separately from whether the result reflects the qualities readers are being asked to assess.

The reported preference for AI-generated stories raises a specific measurement concern. If an LLM judge rewards recognizable stylistic patterns more readily than surprise or unpredictability, a system could appear more creative because it matches the judge’s preferences, not because it produces writing that human readers value more. The source does not describe a deployment or a documented institutional decision, so these are practical implications of the study’s findings rather than reported real-world consequences. The concern is therefore about how an evaluation may shape interpretation, especially when its score is treated as a direct summary of creative quality.

The work is also relevant to claims that language models are approaching or exceeding human performance in creative tasks. A strong score from an automated evaluator is not necessarily evidence of broad creative superiority if the evaluator is poorly aligned with human judgment. At the same time, the source does not say that humans are reliable or unbiased judges, nor does it show that AI-generated stories are less creative overall. It reports a disagreement among evaluation methods. That disagreement leaves room for a more careful reading of comparative results, in which the evaluator’s behavior is treated as part of the evidence rather than as a neutral backdrop.

Interactive Mechanism

Διαδραστικός Μηχανισμός: Πώς λειτουργεί στην πραγματικότητα

Εξερευνήστε την υποκείμενη τεχνολογία πίσω από αυτήν την εξέλιξη διαδραστικά.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Διαδραστικός Έλεγχος Έννοιας+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Τι να παρακολουθήσετε στη συνέχεια

The paper is a preprint, and the source text does not provide details about the evaluators, judge models, prompts, statistical methods, or the size and composition of the story sample. Follow-up work should test whether the reported mismatch holds across genres, languages, models, and independent groups of readers.

The main unknown is how the experiments were conducted in detail. The provided source does not state how many human and AI-generated stories were assessed, which language models produced the AI stories, how the 11 creativity dimensions were defined, or how human ratings were aggregated. Those details are important for judging the strength and scope of the reported conclusions. They would also help readers understand how closely the reported comparison reflects the conditions under which the results might later be interpreted or reproduced.

Further scrutiny should examine the objective metrics and LLM judges used in the comparison. The abstract does not name them, describe their prompts, or indicate whether the judges were tested for reliability across repeated evaluations. Results could depend on the choice of metric, judge model, instructions, story length, or writing style. Without those particulars, the broad finding is informative but difficult to connect to a specific evaluation setup. Clarifying the setup would make it easier to distinguish a general limitation from a result tied to particular tools or instructions.

Replication will show whether the reported pattern extends beyond the WritingPrompts and short stories. Useful tests would include other genres, languages, model families, human reader populations, and creative formats. The source also leaves open whether improved evaluation methods can better reflect human judgments, or whether some aspects of creativity will remain resistant to any single automatic score. Those open questions define the next stage of interpretation: determining how durable the mismatch is and what kinds of assessment can address it.

Σχετικοί οδηγοί και κουίζ

Επεξήγηση μοντέλων AIChatGPT και LLMΗθική του AIΕκπαίδευση AIΔοκιμάστε τι γνωρίζετε — δοκιμάστε ένα δωρεάν κουίζ AIΑναζητήστε έναν όρο AI στο γλωσσάρι μαςΑκολουθήστε τον ιχνηλάτη έκδοσης μοντέλου AI
Βρήκατε αυτό χρήσιμο;