Back to News
InnovationAI Understanding briefing

Study finds AI reviewers show no gender or geographic bias in abstract scoring

A new study in the Journal of General Internal Medicine found that ChatGPT and Claude scored scientific abstracts identically regardless of author gender or country, demonstrating high reproducibility but limited scope.

4 min readRead the linked source
Source-provided image accompanying Study finds AI reviewers show no gender or geographic bias in abstract scoring
Source referenceSource recorded
Publisher
bioengineer.org
Source link
bioengineer.orghttps://bioengineer.org/ai-reviewers-show-no-gender-or-geographic-bias-in-abstract-scoring-study-finds/
Source type
Linked source — primary-source status has not been established.
ContextUnderstand this in 60 seconds

Start here

Key terms

Bias
A consistent pattern of error or unfairness in data or model behavior.
Memory (Agent Memory)
Stored context an AI agent uses across steps or sessions to improve continuity.
Test yourselfAI Ethics Quiz

What happened

Researchers Paul Sebo and Ting Wang published a study in the Journal of General Internal Medicine testing whether AI models exhibit in peer review. They evaluated 50 medical abstracts using ChatGPT and Claude, assigning fictional author identities to test for gender and geographic bias. The models showed no significant bias and high consistency in scoring.

Paul Sebo of the University of Geneva and Ting Wang of Emporia State University conducted an experimental study to assess the and reproducibility of large language models in scientific peer review. The study, published in the Journal of General Internal Medicine, focused on two leading AI models: OpenAI’s ChatGPT and Anthropic’s Claude.

The researchers selected 50 original research abstracts from ten general internal medicine journals published between 2023 and 2026. Each abstract was evaluated under four fictional author identities: an American woman, an American man, a woman from Côte d’Ivoire, and a man from Côte d’Ivoire. This design allowed the team to isolate identity cues while keeping the scientific content constant.

Each abstract-identity combination was scored twice by each model in fresh chat sessions to prevent memory contamination. The models were prompted to rate the abstracts on scientific quality, novelty, and likelihood of acceptance on a 0-10 scale. The study generated 800 total evaluations, collected between April 15 and April 30, 2026.

The results showed that both models exhibited no significant gender or geographic . For ChatGPT, median scores were identical across identities, and for Claude, all three scores were identical. Multivariable regression analysis confirmed no overall association between author identity and scores, with only minor, non-significant variations noted.

Source details: bioengineer.org ↗

Why it matters

This research provides early evidence that AI could serve as a neutral tool for initial manuscript screening, potentially mitigating human biases in scientific publishing. However, the study's limited scope and lack of human comparison mean it does not yet prove AI can replace human reviewers.

The study highlights the potential for AI to address the well-documented biases in human peer review, such as gender and geographic disparities. By demonstrating that AI models can score identical content consistently regardless of author identity, the research suggests a path toward more equitable initial screening processes in scientific publishing.

However, the authors emphasize that consistency does not equate to validity. The study did not include a human reference standard to calibrate the AI scores, and the task was simplified to abstract scoring rather than full manuscript review. Therefore, the findings do not prove that AI judgments are accurate or useful for final editorial decisions.

The research also revealed that AI models are sensitive to journal impact factors, with abstracts from higher-impact journals receiving higher scores. This suggests that while AI may not exhibit explicit identity , it may still reflect implicit biases related to institutional prestige or writing quality associated with selective venues.

The implications for the scientific community are significant but cautious. AI tools could be useful for low-stakes applications like detecting reporting deficiencies or triaging submissions, but they are not yet ready to replace human reviewers. The study underscores the need for further research to validate AI performance in more complex, real-world peer review scenarios.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Document Size:128K tokens
Needle Placement Depth (Location in document):50% into text
Attention Context Buffer Map:
Target Fact (50%)
Equivalent Pages~320Standard book pages
Retrieval Accuracy99.9%Needle recall score
RAM / KV Cache5.1 GBMemory overhead
Prompt CachingActive~80% discount on reuse
Core takeaway: Million-token context windows allow querying whole codebases or legal archives in one prompt. However, KV cache memory scales with context length, making prompt caching crucial for real-time production.
Interactive Concept Check+10 Points
AI Ethics Quiz

Impossibility results in algorithmic fairness (e.g. Kleinberg et al., Chouldechova) show what?

What to watch next

Future studies comparing AI scores against human reviewer decisions, and the adoption of AI triage tools by major journals.

Future studies that compare AI-generated reviews with human reviewer reports and editorial decisions will be crucial for validating the practical utility of AI in peer review. Such head-to-head comparisons will help determine if AI can match or exceed human performance in identifying flaws and assessing novelty.

The adoption of AI triage tools by major scientific journals is a key development to monitor. If journals begin using AI for initial screening, it could significantly reduce the burden on human reviewers and potentially speed up the publication process, but it will also raise questions about transparency and accountability.

The evolution of AI models and their ability to handle more complex tasks, such as full manuscript review and narrative critique, will be important. Current findings are limited to abstract scoring, and it remains unclear if AI can maintain its consistency and lack of in more nuanced evaluative tasks.

The study's limitations, including the small sample size and narrow focus on internal medicine, suggest that broader research is needed to generalize the findings to other scientific disciplines. Future studies should include a wider range of fields, models, and identity variables to provide a more comprehensive understanding of AI in peer review.

Related guides & quizzes

AI EthicsAI Models ExplainedFuture of AITest what you know — try a free AI quizLook up an AI term in our glossaryFollow the AI model release tracker
Found this useful?