Back to News
InnovationAI Understanding briefing

Tech Xplore reports benchmark finding bias in expert recommendations across 22 LLMs

Tech Xplore reports that a benchmark of 22 language models found trade-offs between factual accuracy and social representation when models recommend experts. Retrieval improved factuality, while prompting improved representation, but no tested intervention improved both.

By 6 min read
AI-generated editorial illustration accompanying Tech Xplore reports benchmark finding bias in expert recommendations across 22 LLMs
The short version

Tech Xplore reports that a benchmark of 22 language models found trade-offs between factual accuracy and social representation when models recommend experts. Retrieval improved factuality, while prompting improved representation, but no tested intervention improved both.

What happened

Tech Xplore reported on LLMScholarBench, a benchmark developed by researchers associated with the Complexity Science Hub to evaluate how large language models recommend experts. The report says the benchmark tested 22 models across technical-quality and social-representation metrics, finding that recommendations often favored highly cited, senior, male, U.S.-based and white scholars. Retrieval-augmented generation improved factual accuracy, while prompting could steer representation, but the two goals remained in tension.

Tech Xplore reported on LLMScholarBench, developed by researchers associated with the Complexity Science Hub to test expert recommendations from 22 models: 20 open-weight and two proprietary models from families including Gemini, GPT, Llama, Qwen, DeepSeek, Grok, Mistral and Gemma. It measured five technical metrics—factual accuracy, consistency, validity, refusals and duplicate recommendations—and four social metrics: connectedness, bibliometric similarity, diversity and parity.

The researchers first evaluated six open-weight models on five physics recommendation tasks, including requests for the top five and top 100 influential experts. The reference database contained more than 450,000 scientists who published in American Physical Society journals between 1893 and 2020. Tech Xplore says models named scientists in that database in roughly 80% of recommendations, while field and seniority mismatches were harder; physics-subfield mismatches averaged about 40% in an initial test.

The report describes demographic and geographic skews. Women represented 14% to 32% of researchers in the reference record, depending on subfield and period, but models often recommended fewer women or none. Asian scholars were the largest demographic group, yet white scholars were frequently overrepresented while Black and Latino researchers were often absent. Recommendations concentrated on highly published and cited scholars, and smaller models clustered recommendations within fewer countries.

Tech Xplore reports factuality scores of 0.63 to 0.82, diversity scores of 0.44 to 0.69, and parity scores of 54% to 60%. DeepSeek and Gemini were among the strongest on factuality, DeepSeek led on diversity, and Gemma led on parity. Four interventions across the 22 models showed retrieval-augmented generation improving technical quality, especially accuracy, while prompt engineering improved representation; combining them still did not improve every measure at once.

A further study examined whether specified role, language or geographic location changed recommendations across six academic disciplines. Location affected results, while language and role did not. Tests of newer proprietary Gemini 2.5 Pro and Flash models with web retrieval reportedly increased factual accuracy but reduced diversity and parity. The source presents LLMScholarBench as an ongoing auditing tool rather than a definitive ranking, noting that some models may have been updated since the research.

Read the primary source: techxplore.com

Why it matters

AI systems are increasingly used to find people for professional opportunities, including researchers, doctors and lawyers. If recommendations repeatedly favor people who are already highly visible, the systems can narrow access to opportunities while presenting their outputs as neutral. The findings also show why accuracy alone is an incomplete measure for recommendation systems: a list can contain real experts and still reproduce or intensify demographic and geographic imbalances.

Expert recommendation can be a gatekeeping step: conference organizers may find keynote speakers, employers may assemble candidate pools, and patients may seek doctors through an LLM. Tech Xplore reports that researchers see these risks beyond academia. Repeatedly naming highly visible people can direct attention and opportunities toward the same group while leaving less-visible qualified people undiscovered.

The findings separate factuality from fairness. A recommendation may be factually valid because the person exists and works in the field, yet socially unbalanced if it excludes groups present in the relevant professional population. The report’s distinction between technical quality and social representation offers a more useful framework than checking only whether names are real or whether citations are available.

The intervention results complicate the assumption that web search solves recommendation bias. Tech Xplore says retrieval improved factual accuracy but, in tests of newer proprietary models, lowered diversity and parity. Researchers attribute that risk to underrepresentation on the web; the source presents this as their interpretation, not an independently established causal finding. Grounding a system externally therefore does not automatically make its information representative.

The geographic effect matters for international users. If location changes which scholars are recommended, a system may encode assumptions about local relevance or visibility rather than expertise alone. The source says language and role did not change results in the reported tests, but that does not show they never matter in other disciplines, languages or models. The result applies to the study’s tested conditions.

The source does not establish that any model caused someone to lose a job, invitation or opportunity. It also lacks enough methodological detail to determine demographic assignments, the number of runs behind each score, or uncertainty calculations. Scores can vary with prompts, model versions, retrieval sources and sampling. The work’s public value is therefore exposing a measurable risk and proposing a repeatable audit structure, not proving every deployment behaves identically.

What to watch next

The source cites a 2026 ACM SIGKDD publication and two arXiv studies, but those primary materials were not independently reviewed here. Important open questions include how results vary with prompts, sampling, model updates and demographic-classification methods, and whether the benchmark predicts outcomes in real hiring, admissions or conference-selection systems. Future evaluations should test whether broader reference data and structured human review can improve factuality and representation together.

Independent examination of the cited primary research is the first priority. Tech Xplore identifies a 2026 ACM SIGKDD paper on benchmarking and intervention-based auditing, plus arXiv papers on the initial audit and persona-prompting effects. Those materials should clarify prompts, model versions, trial counts, statistical uncertainty, demographic-labeling procedures, reference populations, refusals and duplicates. These details should not be inferred from the secondary report alone.

Researchers and deployers should test whether LLMScholarBench findings extend beyond physics and the six disciplines described. APS publication records may not represent fields with different publication patterns, geographic structures or demographic data, and may miss expertise documented outside academic publishing. Testing medicine, law, engineering, public service and skilled trades would show whether the risk generalizes to settings where people already use AI recommendations.

Model updates and retrieval sources require continuous monitoring. The report says some evaluated models have changed and describes newer Gemini tests with web retrieval that produced a different balance of results. A single run can become outdated. Follow-up should compare releases over time, document retrieved web sources, and measure whether changes improve accuracy and representation together rather than shifting performance between dimensions.

Organizations using AI to recommend people should require transparent criteria, human review and a way for qualified individuals to be considered when omitted. They should record the prompt, model version, retrieval setting and output, and evaluate more than whether names are real. The source reports neither a regulatory requirement nor a confirmed industry response, so these are practical safeguards suggested by the risks, not measures already adopted because of the benchmark.

It remains unresolved whether more representative data can overcome the reported trade-off. Tech Xplore says the team plans to build a broader scholarly knowledge base, but provides no evidence that it is complete or improves benchmark performance. Future results should show reproducible gains in factuality, diversity and parity, plus persistence in realistic recommendation tasks without merely optimizing the benchmark.

Related guides & quizzes

ChatGPT & LLMsAI Models ExplainedAI EthicsTransformersTest what you know — try a free AI quizLook up an AI term in our glossary
Found this useful?