ወደ ዜና ተመለስ
ፖሊሲAI Understanding አጭር መግለጫ

ብዙ ገምጋሚዎች በICML ኮንፈረንስ ላይ የ AI እገዳዎችን ችላ ሲሉ በጥናት ተረጋግጧል

በ2026 አለም አቀፍ የማሽን መማሪያ ኮንፈረንስ ላይ የተደረገ በዘፈቀደ የተደረገ ሙከራ እንደሚያሳየው በግምት አንድ አራተኛ የሚሆኑ ገምጋሚዎች ትልቅ የቋንቋ ሞዴሎችን መጠቀማቸውን ፖሊሲ ቢከለክልም ተቀባይነት ባላቸው መጠኖች ወይም የግምገማ ውጤቶች ላይ ብዙም ተጽዕኖ አልነበራቸውም።

4 min readRead the original reporting
Source-provided image accompanying Study finds many reviewers ignore AI bans at ICML conference
ሪፖርት ተደርጓልምንጭ ተመዝግቧል
አታሚ
newscientist.com
ምንጭ አገናኝ
newscientist.comhttps://www.newscientist.com/article/2590949-scientists-cant-stop-using-ai-even-when-forbidden-from-doing-so/
የምንጭ ዓይነት
በዜና ማሰራጫ ሪፖርት ማድረግ - የአንደኛ ወገን ሰነድ አይደለም።

በግል ማረጋገጥ ያልቻልነው ነገር: ይህ የይገባኛል ጥያቄ በተሰየመው መውጫ ምክንያት ነው። በአንደኛ ወገን ሰነድ ላይ አላረጋገጥነውም። (newscientist.com)

አውድይህንን በ60 ሰከንድ ውስጥ ይረዱት።

እዚ ጀምር

ቁልፍ ቃላት

ትልቅ የቋንቋ ሞዴል (LLM)
ጽሑፍን ለማፍለቅ እና ለመተንተን በትልቅ ጽሑፍ ኮርፖራ ላይ የሰለጠነ የቋንቋ ሞዴል።
የማሽን መማር (ML)
ስርዓቶች ከውሂብ ንድፎችን እንዲማሩ እና በጊዜ ሂደት እንዲሻሻሉ የሚያስችሉ ዘዴዎች።
እራስህን ፈትን።AI የስነምግባር ጥያቄዎች

ምን ተፈጠረ

A team of Microsoft Research scientists conducted a large‑scale randomized trial during the 2026 International Conference on Machine Learning (ICML) in Seoul. Reviewers were split between a strict policy that banned the use of large language models (LLMs) for any part of the review process and a permissive policy that allowed LLMs for background tasks but not for judging paper merit. The study covered 24,661 submitted papers and 17,886 reviewers. Acceptance rates were virtually identical (27 % vs. 26.5 %) and average review scores differed by only 0.01 points. An anonymous post‑conference survey of 1,486 reviewers revealed that 22.5 % of those instructed not to use AI admitted they did so, mainly to manage workload and generate draft text. Text‑detection analysis showed that only 52.2 % of reviews under the strict policy appeared fully human‑written, compared with 37.0 % under the permissive policy.

The experiment was embedded in the ICML 2026 review workflow. Reviewers could opt into either a "conservative" track that prohibited any LLM assistance or a "permissive" track that allowed LLMs for literature searches, summarisation, and polishing of review language, but not for evaluating the scientific contribution.

Statistical analysis showed no significant difference in acceptance rates (27 % vs. 26.5 %) or average scores (3.31 vs. 3.32 out of 6). Review length increased by 5.5‑7 % under the permissive policy, and expert raters judged those reviews slightly higher in quality, though a reviewer‑by‑reviewer comparison found no meaningful advantage.

The post‑conference survey, with 1,486 respondents, revealed that 22.5 % of reviewers in the strict track used LLMs anyway, citing heavy workloads and unclear rules as primary motivators. The researchers used the Pangram AI‑text detector, which classified only about half of the strict‑policy reviews as fully human‑written, compared with 37 % under the permissive policy, underscoring detector imperfections.

የምንጭ ዝርዝሮች: newscientist.com ↗

ለምን አስፈላጊ ነው።

The findings suggest that outright bans on AI assistance in academic peer review are difficult to enforce, raising concerns for the integrity of scholarly evaluation across computer‑science conferences and potentially other disciplines. If reviewers routinely rely on LLMs to summarize papers or draft feedback, subtle biases or errors introduced by the models could affect acceptance decisions, even if overall acceptance rates appear unchanged. The study also highlights the limitations of current AI‑text detectors, which misclassify a substantial share of reviews. These insights are relevant for conference organizers, journal editors, and funding agencies that are considering how to regulate AI use in peer review to preserve transparency and accountability.

Enforcement challenges imply that simple policy bans may be insufficient to prevent AI‑assisted reviewing, potentially compromising the perceived fairness of the peer‑review process.

The modest impact on acceptance metrics suggests that AI assistance does not dramatically alter outcomes, but the lack of transparency about AI use could mask subtle influences on reviewer judgments.

Detector limitations mean that reliance on automated tools to flag AI‑generated content may produce false positives or miss many AI‑assisted reviews, complicating compliance monitoring.

Interactive Mechanism

በይነተገናኝ ሜካኒዝም፡ በትክክል እንዴት እንደሚሰራ

ከዚህ ልማት በስተጀርባ ያለውን ቴክኖሎጂ በይነተገናኝ ያስሱ።

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
በይነተገናኝ ጽንሰ-ሐሳብ ቼክ+10 Points
AI Ethics Quiz

Why can ethical evaluation not be reduced to one model score?

ቀጥሎ ምን እንደሚታይ

Future policy experiments at major conferences, the development of more reliable AI‑text detection tools, and any formal guidelines issued by professional societies on AI assistance in peer review. Monitoring whether conferences adopt hybrid models—such as mandatory disclosure of AI use or limited‑scope AI tools—will indicate how the community balances efficiency gains against the risk of hidden automation.

Whether major conferences adopt disclosure requirements for AI‑assisted reviews or develop standardized AI‑use guidelines.

Advances in AI‑text detection accuracy that could enable more reliable enforcement of review policies.

Potential policy statements from societies such as the Association for Computing Machinery (ACM) or the Institute of Electrical and Electronics Engineers (IEEE) addressing AI use in scholarly evaluation.

ተዛማጅ መመሪያዎች እና ጥያቄዎች

የAI ሥነ ምግባርAI ሞዴሎች ተብራርተዋልየAI መጪው ጊዜAI ስልጠናየሚያውቁትን ይሞክሩ - ነፃ የ AI ጥያቄዎችን ይሞክሩበእኛ የቃላት መፍቻ ውስጥ የ AI ቃልን ይፈልጉየ AI ደንብ መከታተያ ይከተሉ
ይህ ጠቃሚ ሆኖ ተገኝቷል?