Study finds automatic creativity scores can diverge from human judgments of LLM stories
A new arXiv preprint reports that automated metrics and LLM judges often fail to match human assessments of creativity in short stories, with judges systematically favoring AI-generated writing styles.