Back to News
InnovationAI Understanding briefing

Preprint reports declining diversity in LLM creative outputs over three years

A preliminary arXiv study finds that responses from language models have become less diverse across open-ended creativity tasks, raising questions about homogenization in human-AI creative work.

By 5 min read
Primary-source image accompanying Preprint reports declining diversity in LLM creative outputs over three years
The short version

A preliminary arXiv study finds that responses from language models have become less diverse across open-ended creativity tasks, raising questions about homogenization in human-AI creative work.

What happened

A preliminary arXiv study examined language-model responses across three years of model releases and reported a statistically significant decline in output diversity. The analysis used open-ended prompts from Infinity-Chat100 and the Alternate Uses Task, measuring response similarity with sentence embeddings.

The headline result is a statistically significant decrease in model output diversity over the three-year period. This is the central reported result described by the source, and it concerns the diversity of language-model outputs during the period covered by the analysis. The wording identifies a measured trend in the analyzed material rather than a conclusion about every possible model response, every creative task or every use of language models. The reported decrease is therefore the specific outcome that the preprint puts forward for consideration.

The authors interpret that pattern as evidence that outputs may be converging in creative substance across models. In that interpretation, the result is connected to the possibility that responses are becoming more alike in the creative substance they provide. This remains an interpretation of the reported pattern, rather than a complete claim about all forms of creativity or every model. The distinction matters because a decrease in measured output diversity and a claim about the underlying nature of creativity are not identical statements. The source supports reporting the convergence as the authors' interpretation of the observed pattern.

The source gives no effect size, confidence interval, model-by-model result or comparison with human responses. Those omissions limit how precisely the reported decrease can be characterized and how directly it can be compared with other kinds of responses. It therefore establishes a reported trend in the analyzed outputs, not a complete account of how language-model creativity works or changes in every setting. The available description supports the existence of the reported result within the analysis while leaving the scale, distribution and wider meaning of that result unspecified.

Read the primary source: arxiv.org

Why it matters

If the pattern holds across models and measurement methods, AI systems could offer increasingly similar creative suggestions rather than a broad range of alternatives. The authors argue that this could affect human agency in co-creative work, although the source does not establish that people are already producing less original work because of these systems.

The evidence should be read within its stated boundaries. It concerns language-model responses to two prompt collections and uses an embedding-based comparison. That scope matters because the reported evidence is tied to the responses and collections examined in the study, not presented as a universal measurement of creative output. The comparison provides the basis for discussing similarity in the analyzed responses, while the source's own description leaves the broader reach of the finding open. The practical interpretation should therefore stay with the reported pattern and its stated limits.

It does not identify a causal mechanism, establish why outputs may be converging, or show that model similarity necessarily reduces human creativity. The reported association between output similarity and the broader concern about creative work is therefore not itself a demonstrated chain of cause and effect. The source also does not establish that a more similar set of model responses must produce a particular outcome for people using those responses. These boundaries keep the finding focused on the analyzed model outputs and on the questions they raise.

The practical significance will depend on whether the pattern survives independent replication and whether people judge the affected responses as less original, less useful or less varied. Until those questions are addressed, the implications remain conditional rather than settled. If the pattern holds across models and measurement methods, AI systems could offer increasingly similar creative suggestions rather than a broad range of alternatives. The authors argue that this could affect human agency in co-creative work, although the source does not establish that people are already producing less original work because of these systems.

What to watch next

The key questions are whether the result replicates across models, prompts and diversity measures, and whether embedding-based convergence corresponds to lower human-rated originality or usefulness. The supplied source does not identify the models, sample sizes, effect sizes, embedding model or underlying cause of the reported trend.

The source offers no explanation for the reported trend, so claims about its cause would be premature. The absence of an explanation means that the reported decrease should be followed as an empirical result whose underlying reason remains unresolved. Future discussion should preserve that distinction and avoid treating a possible cause as an established one. The central unresolved issue is not whether a cause can be imagined, but whether competing explanations can be tested against the reported pattern in the analyzed outputs.

Future research will need to test competing explanations and determine whether the pattern appears only in the two studied tasks or extends to other forms of creative work. The key questions are whether the result replicates across models, prompts and diversity measures, and whether embedding-based convergence corresponds to lower human-rated originality or usefulness. This would clarify whether the reported pattern is tied to the particular prompt collections and comparison method or is also visible under other approaches. The supplied source does not identify the models, sample sizes, effect sizes, embedding model or underlying cause of the reported trend.

It is also unknown whether users can counteract convergence through prompt design, model choice or other workflow decisions. Until those questions are answered, the strongest conclusion is that the preprint reports a measurable and potentially consequential trend that remains incomplete. The result should therefore be watched through replication, broader task coverage and closer comparison between embedding-based convergence and human judgments. Those checks would help determine the practical meaning of the reported trend without claiming more than the supplied source establishes.

Related guides & quizzes

Found this useful?
The Weekly Briefing

Get the AI stories that actually matter.

One useful email a week — what changed in AI, why it matters, plus tools, guides, opportunities, and practical ways to take action.

Free · No spam · Unsubscribe in one click