What happened
Researchers Paul Sebo and Ting Wang published a study in the Journal of General Internal Medicine testing whether AI models exhibit in peer review. They evaluated 50 medical abstracts using ChatGPT and Claude, assigning fictional author identities to test for gender and geographic bias. The models showed no significant bias and high consistency in scoring.
Paul Sebo of the University of Geneva and Ting Wang of Emporia State University conducted an experimental study to assess the and reproducibility of large language models in scientific peer review. The study, published in the Journal of General Internal Medicine, focused on two leading AI models: OpenAI’s ChatGPT and Anthropic’s Claude.
The researchers selected 50 original research abstracts from ten general internal medicine journals published between 2023 and 2026. Each abstract was evaluated under four fictional author identities: an American woman, an American man, a woman from Côte d’Ivoire, and a man from Côte d’Ivoire. This design allowed the team to isolate identity cues while keeping the scientific content constant.
Each abstract-identity combination was scored twice by each model in fresh chat sessions to prevent memory contamination. The models were prompted to rate the abstracts on scientific quality, novelty, and likelihood of acceptance on a 0-10 scale. The study generated 800 total evaluations, collected between April 15 and April 30, 2026.
The results showed that both models exhibited no significant gender or geographic . For ChatGPT, median scores were identical across identities, and for Claude, all three scores were identical. Multivariable regression analysis confirmed no overall association between author identity and scores, with only minor, non-significant variations noted.
Source details: bioengineer.org ↗
Why it matters
This research provides early evidence that AI could serve as a neutral tool for initial manuscript screening, potentially mitigating human biases in scientific publishing. However, the study's limited scope and lack of human comparison mean it does not yet prove AI can replace human reviewers.
The study highlights the potential for AI to address the well-documented biases in human peer review, such as gender and geographic disparities. By demonstrating that AI models can score identical content consistently regardless of author identity, the research suggests a path toward more equitable initial screening processes in scientific publishing.
However, the authors emphasize that consistency does not equate to validity. The study did not include a human reference standard to calibrate the AI scores, and the task was simplified to abstract scoring rather than full manuscript review. Therefore, the findings do not prove that AI judgments are accurate or useful for final editorial decisions.
The research also revealed that AI models are sensitive to journal impact factors, with abstracts from higher-impact journals receiving higher scores. This suggests that while AI may not exhibit explicit identity , it may still reflect implicit biases related to institutional prestige or writing quality associated with selective venues.
The implications for the scientific community are significant but cautious. AI tools could be useful for low-stakes applications like detecting reporting deficiencies or triaging submissions, but they are not yet ready to replace human reviewers. The study underscores the need for further research to validate AI performance in more complex, real-world peer review scenarios.
Interactive Mechanism: How It Actually Works
Explore the underlying technology behind this development interactively.
Impossibility results in algorithmic fairness (e.g. Kleinberg et al., Chouldechova) show what?
What to watch next
Future studies comparing AI scores against human reviewer decisions, and the adoption of AI triage tools by major journals.
Future studies that compare AI-generated reviews with human reviewer reports and editorial decisions will be crucial for validating the practical utility of AI in peer review. Such head-to-head comparisons will help determine if AI can match or exceed human performance in identifying flaws and assessing novelty.
The adoption of AI triage tools by major scientific journals is a key development to monitor. If journals begin using AI for initial screening, it could significantly reduce the burden on human reviewers and potentially speed up the publication process, but it will also raise questions about transparency and accountability.
The evolution of AI models and their ability to handle more complex tasks, such as full manuscript review and narrative critique, will be important. Current findings are limited to abstract scoring, and it remains unclear if AI can maintain its consistency and lack of in more nuanced evaluative tasks.
The study's limitations, including the small sample size and narrow focus on internal medicine, suggest that broader research is needed to generalize the findings to other scientific disciplines. Future studies should include a wider range of fields, models, and identity variables to provide a more comprehensive understanding of AI in peer review.