العودة إلى الأخبار
وسائل الإعلامAI Understanding إحاطة

وجدت دراسة أن المحادثات القصيرة باستخدام الذكاء الاصطناعي تقلل من الإيمان بالمؤامرات الجديدة

توصلت تجربتان أمريكيتان إلى أن المحادثات المخصصة Gemini قللت من اعتقاد المشاركين في نظريات المؤامرة التي تتشكل حديثًا أكثر من الدردشات غير ذات الصلة وعادة ما تكون أكثر من صحائف الحقائق الثابتة.

5 min readRead the primary source
وثيقة المصدر الأساسيتم تسجيل المصدر
الناشر
Costello and colleagues' research paper on arXiv
رابط المصدر
arxiv.orghttps://arxiv.org/abs/2608.06151
نوع المصدر
المستند الأساسي - إعلان رسمي أو ورقة أو ملف أو صفحة الطرف الأول التي نقرأها مباشرة.
السياقافهم هذا في 60 ثانية

ابدأ هنا

المصطلحات الرئيسية

التصنيف
مهمة حيث يقوم النموذج بتعيين مدخلات لواحدة أو أكثر من الفئات المحددة مسبقًا.
اختبر نفسكمسابقة أخلاقيات الذكاء الاصطناعي

ماذا حدث

A Carnegie Mellon, MIT, and Cornell research team posted two randomized case studies on August 6 testing whether short AI conversations could counter conspiracy beliefs while facts about major events were still emerging.

The researchers analyzed 472 US adults who expressed conspiratorial views after the July 2024 attempt to assassinate Donald Trump and 1,035 after the September 2025 assassination of Charlie Kirk. Participants were randomly assigned to a tailored debunking dialogue, an unrelated AI conversation, or a static list of contemporaneous facts.

The dialogue groups completed at least five exchanges with Gemini 1.5 Pro in the first experiment and Gemini 2.5 Pro in the second. Both models received a curated fact base that separated confirmed information, debunked claims, and unresolved questions; the second model could also search the web only to verify factual information.

On a 0-to-100 measure of belief in each participant's own stated theory, the tailored dialogues produced reductions of 6.95 points relative to the unrelated-chat control in the first experiment and 7.56 points in the second. Differences from the static fact sheet were 5.60 and 6.56 points. The paper reports standardized effects of roughly 0.32 to 0.38 for those comparisons.

The two case studies were designed around moments when the factual record was still developing. The first followed the July 2024 attempt to assassinate Donald Trump; the second followed the September 2025 assassination of Charlie Kirk. Participants were screened for conspiratorial or uncertain interpretations, then assigned to a conversation that addressed their own stated theory, an unrelated AI chat, or a static contemporaneous fact sheet. The second study was preregistered, while the first was not, so the replication is informative but not equivalent to two independent confirmatory trials.

تفاصيل المصدر: Costello and colleagues' research paper on arXiv

لماذا يهم

The result suggests that a conversational system can do more than repeat a correction: it can adapt facts, questions, and uncertainty to the specific reason a person gives for a belief.

That distinction is useful during fast-moving events, when verified evidence is incomplete and a generic fact sheet may not address the claim a person actually finds persuasive. The paper's strategy analysis found that the model leaned more on credible sources, open questions, and caution when little was known, then used more direct evidence when the second event had a larger factual record.

The authors also report limited evidence of later spillover. Two months after the first experiment, participants assigned to the dialogue were less likely than pooled controls to endorse two conspiratorial interpretations of a subsequent event. In the second follow-up, direct treatment assignment did not significantly reduce beliefs about a later shooting, although a separate preregistered persistence analysis and a broader conspiracy measure suggested smaller carry-over effects.

The design shows why the interaction may outperform a static correction without proving that persuasion is always desirable. Gemini 1.5 Pro in the first study and Gemini 2.5 Pro in the second received researcher-curated fact bases separating confirmed information, debunked claims, and unresolved questions. The model could ask about the participant's specific reasoning and choose whether to answer directly, cite credible evidence, or preserve uncertainty. That adaptability is useful for education, but it also gives the system discretion over which claims to challenge and which sources to foreground.

The study does not justify automatically deploying persuasive bots into public conversations. A system instructed to change beliefs holds unusual power, and the authors note that similar techniques can also increase belief in false claims. Any real service would need transparent authorship, reliable event grounding, safeguards against political targeting, and independent tests of accuracy and unintended effects.

Interactive Mechanism

الآلية التفاعلية: كيف تعمل فعليًا

استكشف التكنولوجيا الأساسية وراء هذا التطور بشكل تفاعلي.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
التحقق من المفهوم التفاعلي+10 Points
AI Ethics Quiz

Impossibility results in algorithmic fairness (e.g. Kleinberg et al., Chouldechova) show what?

ماذا تشاهد بعد ذلك

Watch for peer review, replication beyond two US political crises, and field evidence showing whether people choose to use such a tool without being recruited into a study.

This is a new preprint built around two case studies. The first experiment was not preregistered, the second was, and both measured self-reported beliefs immediately after a short online interaction rather than real-world sharing behavior, voting, or long-term trust.

The analyzed samples included only participants whose open responses were classified as conspiratorial or uncertain by GPT-4o, although the authors report similar patterns under alternate rules. Both studies recruited US adults through CloudResearch, so the findings may not transfer to other countries, languages, events, or populations.

The models were not relying on unaided knowledge: researchers supplied carefully assembled fact bases and explicit persuasion instructions. Future work should test who maintains those facts under deadline pressure, how errors are corrected, whether opposing viewpoints are treated consistently, and when a system should preserve uncertainty instead of trying to persuade.

There is also a measurement distinction between changing a stated belief and improving a person's understanding. The outcomes were self-reported ratings collected after short online conversations, not observed sharing behavior, source checking, voting, or decisions made during a live crisis. The paper's follow-ups provide early evidence about persistence, but the mixed results mean a later study should pre-register long-term outcomes, track attrition, and test whether participants can explain the evidence themselves rather than simply repeating the model's conclusion.

Replication should vary the source of the factual record as well as the population. These experiments supplied researchers' own fact bases and recruited US adults through CloudResearch, which makes the setup unusually controlled but leaves open questions about multilingual events, lower-connectivity settings, and crises in which official information is delayed or disputed. A responsible field trial would make the model's authorship visible, let participants decline the intervention, log which evidence was shown, and use independent adjudicators to check whether the dialogue corrected a misconception without introducing a new one. It should measure trust, emotional response, and willingness to share claims alongside belief scores.

الأدلة والاختبارات ذات الصلة

أخلاقيات الذكاء الاصطناعيسلامة الذكاء الاصطناعيما هو الذكاء الاصطناعي؟اختبر ما تعرفه – جرّب اختبارًا مجانيًا للذكاء الاصطناعيابحث عن مصطلح الذكاء الاصطناعي في قاموسنا
وجدت هذا مفيدا؟