Quay lại Tin tức
Đổi mớiAI Understanding tóm tắt

Tech Xplore báo cáo xu hướng tìm kiếm điểm chuẩn trong các đề xuất của chuyên gia trên 22 LLM

Tech Xplore báo cáo rằng điểm chuẩn của 22 mô hình ngôn ngữ đã tìm thấy sự cân bằng giữa độ chính xác thực tế và tính đại diện xã hội khi các mô hình giới thiệu các chuyên gia. Việc truy xuất đã cải thiện tính xác thực, đồng thời thúc đẩy việc trình bày được cải thiện, nhưng không có sự can thiệp nào được thử nghiệm đã cải thiện cả hai.

6 min readRead the linked source
Source-provided image accompanying Tech Xplore reports benchmark finding bias in expert recommendations across 22 LLMs
Nguồn tham khảoNguồn đã ghi
Nhà xuất bản
techxplore.com
Liên kết nguồn
techxplore.comhttps://techxplore.com/news/2026-08-ai-expert-benchmark-bias-llms.html
Loại nguồn
Nguồn được liên kết - trạng thái nguồn chính chưa được thiết lập.
Bối cảnhHiểu điều này trong 60 giây

Bắt đầu ở đây

Thuật ngữ chính

Điểm chuẩn
Một bài kiểm tra hoặc tập dữ liệu được tiêu chuẩn hóa dùng để đo lường và so sánh hiệu suất của mô hình.
thiên vị
Một dạng lỗi hoặc sự không công bằng nhất quán trong dữ liệu hoặc hành vi của mô hình.
RAG (Thế hệ tăng cường truy xuất)
Một phương pháp truy xuất kiến thức bên ngoài và đưa nó vào thế hệ tại thời điểm suy luận.
Tự kiểm traChatGPT & Câu đố LLM

Chuyện gì đã xảy ra

Tech Xplore reported on LLMScholarBench, a developed by researchers associated with the Complexity Science Hub to evaluate how large language models recommend experts. The report says the benchmark tested 22 models across technical-quality and social-representation metrics, finding that recommendations often favored highly cited, senior, male, U.S.-based and white scholars. Retrieval-augmented generation improved factual accuracy, while prompting could steer representation, but the two goals remained in tension.

Tech Xplore reported on LLMScholarBench, developed by researchers associated with the Complexity Science Hub to test expert recommendations from 22 models: 20 open-weight and two proprietary models from families including Gemini, GPT, Llama, Qwen, DeepSeek, Grok, Mistral and Gemma. It measured five technical metrics—factual accuracy, consistency, validity, refusals and duplicate recommendations—and four social metrics: connectedness, bibliometric similarity, diversity and parity.

The researchers first evaluated six open-weight models on five physics recommendation tasks, including requests for the top five and top 100 influential experts. The reference database contained more than 450,000 scientists who published in American Physical Society journals between 1893 and 2020. Tech Xplore says models named scientists in that database in roughly 80% of recommendations, while field and seniority mismatches were harder; physics-subfield mismatches averaged about 40% in an initial test.

The report describes demographic and geographic skews. Women represented 14% to 32% of researchers in the reference record, depending on subfield and period, but models often recommended fewer women or none. Asian scholars were the largest demographic group, yet white scholars were frequently overrepresented while Black and Latino researchers were often absent. Recommendations concentrated on highly published and cited scholars, and smaller models clustered recommendations within fewer countries.

Tech Xplore reports factuality scores of 0.63 to 0.82, diversity scores of 0.44 to 0.69, and parity scores of 54% to 60%. DeepSeek and Gemini were among the strongest on factuality, DeepSeek led on diversity, and Gemma led on parity. Four interventions across the 22 models showed retrieval-augmented generation improving technical quality, especially accuracy, while prompt engineering improved representation; combining them still did not improve every measure at once.

A further study examined whether specified role, language or geographic location changed recommendations across six academic disciplines. Location affected results, while language and role did not. Tests of newer proprietary Gemini 2.5 Pro and Flash models with web retrieval reportedly increased factual accuracy but reduced diversity and parity. The source presents LLMScholarBench as an ongoing auditing tool rather than a definitive ranking, noting that some models may have been updated since the research.

Chi tiết nguồn: techxplore.com ↗

Tại sao nó quan trọng

AI systems are increasingly used to find people for professional opportunities, including researchers, doctors and lawyers. If recommendations repeatedly favor people who are already highly visible, the systems can narrow access to opportunities while presenting their outputs as neutral. The findings also show why accuracy alone is an incomplete measure for recommendation systems: a list can contain real experts and still reproduce or intensify demographic and geographic imbalances.

Expert recommendation can be a gatekeeping step: conference organizers may find keynote speakers, employers may assemble candidate pools, and patients may seek doctors through an LLM. Tech Xplore reports that researchers see these risks beyond academia. Repeatedly naming highly visible people can direct attention and opportunities toward the same group while leaving less-visible qualified people undiscovered.

The findings separate factuality from fairness. A recommendation may be factually valid because the person exists and works in the field, yet socially unbalanced if it excludes groups present in the relevant professional population. The report’s distinction between technical quality and social representation offers a more useful framework than checking only whether names are real or whether citations are available.

The intervention results complicate the assumption that web search solves recommendation . Tech Xplore says retrieval improved factual accuracy but, in tests of newer proprietary models, lowered diversity and parity. Researchers attribute that risk to underrepresentation on the web; the source presents this as their interpretation, not an independently established causal finding. Grounding a system externally therefore does not automatically make its information representative.

The geographic effect matters for international users. If location changes which scholars are recommended, a system may encode assumptions about local relevance or visibility rather than expertise alone. The source says language and role did not change results in the reported tests, but that does not show they never matter in other disciplines, languages or models. The result applies to the study’s tested conditions.

The source does not establish that any model caused someone to lose a job, invitation or opportunity. It also lacks enough methodological detail to determine demographic assignments, the number of runs behind each score, or uncertainty calculations. Scores can vary with prompts, model versions, retrieval sources and sampling. The work’s public value is therefore exposing a measurable risk and proposing a repeatable audit structure, not proving every deployment behaves identically.

Interactive Mechanism

Cơ chế tương tác: Nó thực sự hoạt động như thế nào

Khám phá công nghệ cơ bản đằng sau sự phát triển này một cách tương tác.

Document Size:128K tokens
Needle Placement Depth (Location in document):50% into text
Attention Context Buffer Map:
Target Fact (50%)
Equivalent Pages~320Standard book pages
Retrieval Accuracy99.9%Needle recall score
RAM / KV Cache5.1 GBMemory overhead
Prompt CachingActive~80% discount on reuse
Core takeaway: Million-token context windows allow querying whole codebases or legal archives in one prompt. However, KV cache memory scales with context length, making prompt caching crucial for real-time production.
Kiểm tra khái niệm tương tác+10 Points
ChatGPT & LLMs Quiz

What is a common training objective for an autoregressive language model?

Xem gì tiếp theo

The source cites a 2026 ACM SIGKDD publication and two arXiv studies, but those primary materials were not independently reviewed here. Important open questions include how results vary with prompts, sampling, model updates and demographic-classification methods, and whether the predicts outcomes in real hiring, admissions or conference-selection systems. Future evaluations should test whether broader reference data and structured human review can improve factuality and representation together.

Independent examination of the cited primary research is the first priority. Tech Xplore identifies a 2026 ACM SIGKDD paper on benchmarking and intervention-based auditing, plus arXiv papers on the initial audit and persona-prompting effects. Those materials should clarify prompts, model versions, trial counts, statistical uncertainty, demographic-labeling procedures, reference populations, refusals and duplicates. These details should not be inferred from the secondary report alone.

Researchers and deployers should test whether LLMScholarBench findings extend beyond physics and the six disciplines described. APS publication records may not represent fields with different publication patterns, geographic structures or demographic data, and may miss expertise documented outside academic publishing. Testing medicine, law, engineering, public service and skilled trades would show whether the risk generalizes to settings where people already use AI recommendations.

Model updates and retrieval sources require continuous monitoring. The report says some evaluated models have changed and describes newer Gemini tests with web retrieval that produced a different balance of results. A single run can become outdated. Follow-up should compare releases over time, document retrieved web sources, and measure whether changes improve accuracy and representation together rather than shifting performance between dimensions.

Organizations using AI to recommend people should require transparent criteria, human review and a way for qualified individuals to be considered when omitted. They should record the prompt, model version, retrieval setting and output, and evaluate more than whether names are real. The source reports neither a regulatory requirement nor a confirmed industry response, so these are practical safeguards suggested by the risks, not measures already adopted because of the .

It remains unresolved whether more representative data can overcome the reported trade-off. Tech Xplore says the team plans to build a broader scholarly knowledge base, but provides no evidence that it is complete or improves performance. Future results should show reproducible gains in factuality, diversity and parity, plus persistence in realistic recommendation tasks without merely optimizing the benchmark.

Hướng dẫn và câu hỏi liên quan

ChatGPT & LLMGiải thích về mô hình AIĐạo đức AIMáy biến ápKiểm tra những gì bạn biết — thử một bài kiểm tra AI miễn phíTra cứu một thuật ngữ AI trong bảng thuật ngữ của chúng tôiTheo dõi trình theo dõi phát hành mô hình AI
Tìm thấy điều này hữu ích?