Zpět na Novinky
InovaceInstruktáž AI Understanding

Studie ukazuje, že jemně vyladěné LLM zlepšují detekci cílů nenávistné řeči

Výzkumníci z Zayed University a UAE University publikovali studii v Discover Artificial Intelligence, která prokázala, že supervizované dolaďování velkých jazykových modelů výrazně překonává nabádání a tradiční základní linie při klasifikaci nenávistných projevů konkrétními cílovými skupinami.

5 min readRead the linked source
Source-provided image accompanying Study shows fine-tuned LLMs improve hate speech target detection
Odkaz na zdrojZdroj zaznamenán
Vydavatel
bioengineer.org
Odkaz na zdroj
bioengineer.orghttps://bioengineer.org/fine-tuned-ai-models-learn-to-name-the-target-of-online-hate/
Typ zdroje
Propojený zdroj — stav primárního zdroje nebyl stanoven.
KontextPochopte to za 60 sekund

Začněte zde

Klíčové pojmy

umělá inteligence (AI)
Široká oblast budování systémů, které provádějí úkoly vyžadující rozpoznávání vzorů, uvažování, jazyk nebo rozhodování.
strojové učení (ML)
Metody, které umožňují systémům učit se vzory z dat a zlepšovat se v průběhu času.
Klasifikace
Úloha, kde model přiřadí vstup jedné nebo více předdefinovaným kategoriím.
Otestujte seEtický kvíz AI

Co se stalo

A new study published in Discover Artificial Intelligence evaluates the ability of large language models to classify hate speech not just as a binary category, but by the specific group targeted, including gender, race, religion, and immigration status. Researchers Sanaa Kaddoura and Sumaia Al-Kohlani tested five models, including DeepSeek-R1-Distill-Qwen-32B, Mistral-7B-Instruct, and GPT-4o, using a dataset of 12,957 English tweets. The study found that while zero-shot and few-shot prompting yielded inconsistent and often poor results, supervised fine-tuning using QLoRA transformed these models into the strongest classifiers tested. The fine-tuned DeepSeek model achieved the highest performance, with a weighted F1 score of 76.71% and an external generalization F1 of 88.31% on the ETHOS benchmark, significantly outperforming traditional machine learning baselines like BERT and logistic regression.

Researchers Sanaa Kaddoura of Zayed University and Sumaia Al-Kohlani of United Arab Emirates University published a study in Discover Artificial Intelligence evaluating how well large language models can classify hateful tweets by the specific group being targeted. The study moved beyond the traditional binary approach of labeling content as simply 'hateful' or 'not hateful,' instead using a six-class taxonomy that includes gender, race and ethnicity, immigration and xenophobia, religion, and general abuse.

The team assembled a dataset of 12,957 English tweets by re-annotating the hate speech subset of the TweetEval benchmark. Three independent annotators labeled each instance, achieving a moderate inter-annotator agreement of 0.6003 on Krippendorff’s alpha. The dataset was imbalanced, reflecting real-world platform data, with non-hate content comprising nearly 58 percent of the samples and religious hate speech making up only 1.3 percent.

Five models were evaluated: DeepSeek-R1-Distill-Qwen-32B, Mistral-7B-Instruct, WizardLM-13B, Phi-4-mini, and OpenAI’s GPT-4o, alongside a fine-tuned BERT-base-uncased baseline. The models were tested using zero-shot prompting, few-shot prompting, and supervised fine-tuning. The open-source models were adapted using QLoRA, a parameter-efficient technique that allows training on a single NVIDIA A100 GPU.

Zero-shot and few-shot prompting produced inconsistent results, with some models refusing to stay within the label space or producing unclassified outputs. In contrast, supervised fine-tuning dramatically improved performance. The fine-tuned DeepSeek model achieved the best multiclass performance with a weighted F1 score of 76.71 percent and eliminated unclassified outputs entirely. It also demonstrated strong generalization, scoring an F1 of 88.31 percent on the external ETHOS benchmark.

The study highlights that while fine-tuned LLMs outperform traditional baselines like logistic regression and BERT, they come with higher computational costs. The DeepSeek model required 4.71 hours of training and 32.5 minutes to classify a 1,620-tweet test set, whereas logistic regression completed the same task in seconds. The authors recommend using LLMs where fine-grained discrimination is necessary and simpler classifiers for high-throughput settings.

Podrobnosti o zdroji: bioengineer.org ↗

Proč na tom záleží

This research addresses a critical limitation in automated content moderation: the failure to distinguish between different targets of hate speech, which often leads to disproportionate misclassification of attacks against minority groups. By demonstrating that fine-tuned LLMs can accurately identify the specific target of abuse, the study provides a technical pathway for more equitable and precise moderation systems. This is practically significant because it allows platforms to apply context-aware policies rather than blunt binary filters, potentially reducing the suppression of counter-speech or satire while better protecting vulnerable communities. The public release of the fine-tuned model and dataset also lowers the barrier for other researchers and organizations to implement or audit these more granular detection capabilities.

Traditional hate speech detectors often collapse all forms of hate into a single category, forcing models to rely on surface-level lexical cues. This leads to uneven performance, where hate directed at minority or rarely referenced groups is misclassified far more often than hate against well-represented targets. By explicitly modeling the target of abuse, this research addresses a structural weakness in earlier approaches.

The ability to distinguish between different targets of hate speech is crucial for equitable content moderation. False positives can suppress counter-speech, satire, or educational discussion, potentially silencing the very communities these systems claim to protect. More precise allows for context-aware policies that can better balance the need to remove harmful content with the need to preserve legitimate discourse.

The study provides a practical benchmark for the research community, as the fine-tuned model and dataset have been released publicly. This allows other researchers and organizations to build on a framework that asks not just whether a post is hateful, but whom it harms. The public availability of these resources lowers the barrier to entry for implementing more granular detection capabilities.

The findings suggest that supervised fine-tuning is a more reliable method than prompting for complex tasks involving nuanced social categories. This has implications for the broader field of AI safety and alignment, where the ability to accurately interpret and categorize human behavior is essential for developing effective moderation tools.

Interactive Mechanism

Interaktivní mechanismus: Jak to vlastně funguje

Interaktivně prozkoumejte základní technologii tohoto vývoje.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Interaktivní kontrola konceptu+10 Points
AI Ethics Quiz

Why can ethical evaluation not be reduced to one model score?

Na co se dále dívat

The adoption of target-specific metrics in industry content moderation standards. Additionally, watch for the integration of these fine-tuned models into open-source moderation toolkits, as well as further research into multilingual extensions and multi-label formulations that the authors have outlined for future work.

Industry adoption of target-specific metrics in content moderation systems. As platforms seek to improve the fairness and accuracy of their moderation tools, they may begin to incorporate these more granular evaluation methods into their standard practices.

Further research into multilingual extensions of this framework. The current study focuses on English tweets, but hate speech is a global phenomenon. Future work will likely explore how these models perform across different languages and cultural contexts.

The development of multi-label formulations for posts that target multiple groups simultaneously. The current study uses a single-label approach, but real-world hate speech often involves multiple targets. Future research will need to address this complexity.

The integration of these fine-tuned models into open-source moderation toolkits. As the model and dataset are publicly available, we may see them adopted by non-profit organizations and smaller platforms that lack the resources to develop their own sophisticated moderation systems.

Související průvodci a kvízy

Etika AIVysvětlení modelů AIŠkolení AIOtestujte si, co víte – vyzkoušejte bezplatný kvíz AIVyhledejte si termín AI v našem slovníkuPostupujte podle sledování vydání modelu AI
Považujete to za užitečné?