Torna alle notizie
InnovazioneAI Understanding briefing

Tech Xplore segnala errori di ricerca dei benchmark nelle raccomandazioni degli esperti in 22 LLM

Tech Xplore riferisce che un benchmark di 22 modelli linguistici ha trovato compromessi tra accuratezza fattuale e rappresentazione sociale quando i modelli raccomandano esperti. Il recupero ha migliorato la fattualità, promuovendo al contempo una migliore rappresentazione, ma nessun intervento testato ha migliorato entrambi.

6 min readRead the linked source
Source-provided image accompanying Tech Xplore reports benchmark finding bias in expert recommendations across 22 LLMs
Riferimento alla fonteFonte registrata
Editore
techxplore.com
Collegamento alla fonte
techxplore.comhttps://techxplore.com/news/2026-08-ai-expert-benchmark-bias-llms.html
Tipo di fonte
Fonte collegata: lo stato di fonte primaria non è stato stabilito.
ContestoComprendilo in 60 secondi

Inizia qui

Termini chiave

Punto di riferimento
Un test o un set di dati standardizzato utilizzato per misurare e confrontare le prestazioni del modello.
Pregiudizio
Un modello coerente di errore o ingiustizia nei dati o nel comportamento del modello.
RAG (generazione aumentata di recupero)
Un metodo che recupera la conoscenza esterna e la alimenta nella generazione al momento dell'inferenza.
Mettiti alla provaChatGPT e quiz LLM

Cosa è successo

Tech Xplore reported on LLMScholarBench, a developed by researchers associated with the Complexity Science Hub to evaluate how large language models recommend experts. The report says the benchmark tested 22 models across technical-quality and social-representation metrics, finding that recommendations often favored highly cited, senior, male, U.S.-based and white scholars. Retrieval-augmented generation improved factual accuracy, while prompting could steer representation, but the two goals remained in tension.

Tech Xplore reported on LLMScholarBench, developed by researchers associated with the Complexity Science Hub to test expert recommendations from 22 models: 20 open-weight and two proprietary models from families including Gemini, GPT, Llama, Qwen, DeepSeek, Grok, Mistral and Gemma. It measured five technical metrics—factual accuracy, consistency, validity, refusals and duplicate recommendations—and four social metrics: connectedness, bibliometric similarity, diversity and parity.

The researchers first evaluated six open-weight models on five physics recommendation tasks, including requests for the top five and top 100 influential experts. The reference database contained more than 450,000 scientists who published in American Physical Society journals between 1893 and 2020. Tech Xplore says models named scientists in that database in roughly 80% of recommendations, while field and seniority mismatches were harder; physics-subfield mismatches averaged about 40% in an initial test.

The report describes demographic and geographic skews. Women represented 14% to 32% of researchers in the reference record, depending on subfield and period, but models often recommended fewer women or none. Asian scholars were the largest demographic group, yet white scholars were frequently overrepresented while Black and Latino researchers were often absent. Recommendations concentrated on highly published and cited scholars, and smaller models clustered recommendations within fewer countries.

Tech Xplore reports factuality scores of 0.63 to 0.82, diversity scores of 0.44 to 0.69, and parity scores of 54% to 60%. DeepSeek and Gemini were among the strongest on factuality, DeepSeek led on diversity, and Gemma led on parity. Four interventions across the 22 models showed retrieval-augmented generation improving technical quality, especially accuracy, while prompt engineering improved representation; combining them still did not improve every measure at once.

A further study examined whether specified role, language or geographic location changed recommendations across six academic disciplines. Location affected results, while language and role did not. Tests of newer proprietary Gemini 2.5 Pro and Flash models with web retrieval reportedly increased factual accuracy but reduced diversity and parity. The source presents LLMScholarBench as an ongoing auditing tool rather than a definitive ranking, noting that some models may have been updated since the research.

Dettagli della fonte: techxplore.com ↗

Perché è importante

AI systems are increasingly used to find people for professional opportunities, including researchers, doctors and lawyers. If recommendations repeatedly favor people who are already highly visible, the systems can narrow access to opportunities while presenting their outputs as neutral. The findings also show why accuracy alone is an incomplete measure for recommendation systems: a list can contain real experts and still reproduce or intensify demographic and geographic imbalances.

Expert recommendation can be a gatekeeping step: conference organizers may find keynote speakers, employers may assemble candidate pools, and patients may seek doctors through an LLM. Tech Xplore reports that researchers see these risks beyond academia. Repeatedly naming highly visible people can direct attention and opportunities toward the same group while leaving less-visible qualified people undiscovered.

The findings separate factuality from fairness. A recommendation may be factually valid because the person exists and works in the field, yet socially unbalanced if it excludes groups present in the relevant professional population. The report’s distinction between technical quality and social representation offers a more useful framework than checking only whether names are real or whether citations are available.

The intervention results complicate the assumption that web search solves recommendation . Tech Xplore says retrieval improved factual accuracy but, in tests of newer proprietary models, lowered diversity and parity. Researchers attribute that risk to underrepresentation on the web; the source presents this as their interpretation, not an independently established causal finding. Grounding a system externally therefore does not automatically make its information representative.

The geographic effect matters for international users. If location changes which scholars are recommended, a system may encode assumptions about local relevance or visibility rather than expertise alone. The source says language and role did not change results in the reported tests, but that does not show they never matter in other disciplines, languages or models. The result applies to the study’s tested conditions.

The source does not establish that any model caused someone to lose a job, invitation or opportunity. It also lacks enough methodological detail to determine demographic assignments, the number of runs behind each score, or uncertainty calculations. Scores can vary with prompts, model versions, retrieval sources and sampling. The work’s public value is therefore exposing a measurable risk and proposing a repeatable audit structure, not proving every deployment behaves identically.

Interactive Mechanism

Meccanismo interattivo: come funziona realmente

Esplora la tecnologia alla base di questo sviluppo in modo interattivo.

Document Size:128K tokens
Needle Placement Depth (Location in document):50% into text
Attention Context Buffer Map:
Target Fact (50%)
Equivalent Pages~320Standard book pages
Retrieval Accuracy99.9%Needle recall score
RAM / KV Cache5.1 GBMemory overhead
Prompt CachingActive~80% discount on reuse
Core takeaway: Million-token context windows allow querying whole codebases or legal archives in one prompt. However, KV cache memory scales with context length, making prompt caching crucial for real-time production.
Verifica concettuale interattiva+10 Points
ChatGPT & LLMs Quiz

What is a common training objective for an autoregressive language model?

Cosa guardare dopo

The source cites a 2026 ACM SIGKDD publication and two arXiv studies, but those primary materials were not independently reviewed here. Important open questions include how results vary with prompts, sampling, model updates and demographic-classification methods, and whether the predicts outcomes in real hiring, admissions or conference-selection systems. Future evaluations should test whether broader reference data and structured human review can improve factuality and representation together.

Independent examination of the cited primary research is the first priority. Tech Xplore identifies a 2026 ACM SIGKDD paper on benchmarking and intervention-based auditing, plus arXiv papers on the initial audit and persona-prompting effects. Those materials should clarify prompts, model versions, trial counts, statistical uncertainty, demographic-labeling procedures, reference populations, refusals and duplicates. These details should not be inferred from the secondary report alone.

Researchers and deployers should test whether LLMScholarBench findings extend beyond physics and the six disciplines described. APS publication records may not represent fields with different publication patterns, geographic structures or demographic data, and may miss expertise documented outside academic publishing. Testing medicine, law, engineering, public service and skilled trades would show whether the risk generalizes to settings where people already use AI recommendations.

Model updates and retrieval sources require continuous monitoring. The report says some evaluated models have changed and describes newer Gemini tests with web retrieval that produced a different balance of results. A single run can become outdated. Follow-up should compare releases over time, document retrieved web sources, and measure whether changes improve accuracy and representation together rather than shifting performance between dimensions.

Organizations using AI to recommend people should require transparent criteria, human review and a way for qualified individuals to be considered when omitted. They should record the prompt, model version, retrieval setting and output, and evaluate more than whether names are real. The source reports neither a regulatory requirement nor a confirmed industry response, so these are practical safeguards suggested by the risks, not measures already adopted because of the .

It remains unresolved whether more representative data can overcome the reported trade-off. Tech Xplore says the team plans to build a broader scholarly knowledge base, but provides no evidence that it is complete or improves performance. Future results should show reproducible gains in factuality, diversity and parity, plus persistence in realistic recommendation tasks without merely optimizing the benchmark.

Guide e quiz correlati

ChatGPT e LLMSpiegazione dei modelli di intelligenza artificialeEtica dell'IATrasformatoriMetti alla prova ciò che sai: prova un quiz gratuito sull'intelligenza artificialeCerca un termine AI nel nostro glossarioSegui il tracker del rilascio del modello AI
Lo hai trovato utile?