Back to News
SecurityAI Understanding briefing

Study Finds People Heed GPT-4 More Than Peers on AI-Written News

A preregistered University of Osaka experiment found students shifted their judgments more toward GPT-4 than peer advice, but the result used GPT-2-era text and does not validate modern AI detectors.

By 4 min read
A student compares ambiguous paper fragments while three distinct sources offer competing evidence through abstract light patterns.
The short version

A preregistered University of Osaka experiment found students shifted their judgments more toward GPT-4 than peer advice, but the result used GPT-2-era text and does not validate modern AI detectors.

What happened

A University of Osaka research paper posted on August 2 reports that students gave more weight to GPT-4 advice than to advice from fellow students when estimating how much of a Japanese news item was human-written.

Yuhao Fu and Nobuyuki Hanaki ran a laboratory study with 220 native Japanese-speaking University of Osaka students across a main experiment and three follow-up sessions. The main experiment was approved by the university institute's review board and preregistered; two of the three added conditions were separately preregistered.

Each participant judged 30 Japanese-language items: 10 human-written Wikinews articles, 10 items generated with a Japanese GPT-2 model, and 10 that began with human-written text and ended with generated text. Participants estimated the human-written share, saw a numerical estimate attributed to GPT-4, another student, or—in later sessions—a linguistics expert, and then submitted a revised estimate.

In the 87-person main experiment, the researchers' weight-of-advice measure averaged 0.592 for GPT-4 advice and 0.326 for peer advice, a statistically significant difference. Final estimates were also more accurate in the GPT-4 condition, but the paper's regression analysis attributes most of that gain to the quality of the advice received rather than to the AI label by itself.

Read the primary source: Fu and Hanaki research paper on arXiv

Why it matters

The study separates trust from usefulness: a source can strongly influence a decision, but that influence helps only when its advice is accurate enough for the task.

That distinction matters for schools, newsrooms, platforms, and other organizations considering automated content-detection tools. A familiar AI label or a user's confidence in it is not evidence that the detector works on the content, language, and model generation they actually face.

The study also complicates a simple AI-versus-human story. In the 2025 follow-up sessions, participants placed more weight on advice attributed to linguistics experts than on GPT-4 advice when the two conditions used the same procedure and calendar period. The earlier GPT-4 group showed higher reliance than the expert group, but the authors caution that the comparison spans different years and procedures.

A safer operational lesson is to validate advice quality and preserve human review. The experiment measured how far people moved toward an estimate; it did not show that an AI detector should make final authorship, misconduct, or misinformation decisions.

What to watch next

Watch for independent replications with current generators, dedicated detectors, broader populations, and real-world decisions where false accusations carry consequences.

The paper is a new preprint, not a completed peer-review verdict. Its participants came from one university, and the controlled lab design cannot establish how journalists, teachers, moderators, or the general public would behave outside the experiment.

The technology gap is especially important: the synthetic text came from GPT-2, while the advice came from prompted GPT-4. Today's generators and specialized detection systems may produce different accuracy and reliance patterns. The authors also did not directly observe which linguistic, contextual, or factual strategies participants used.

Future evidence should report false-positive rates, performance across languages and model families, and whether showing uncertainty or provenance changes appropriate reliance. Until then, this study is evidence about advice-taking under one controlled setup—not a recommendation to use ChatGPT as an AI-writing detector.

Related guides & quizzes

Found this useful?
The Monthly Briefing

Get the AI stories that actually matter.

One short email a month — what changed in AI, why it matters, plus the tools and guides worth your time.

Free · No spam · Unsubscribe in one click