Вернуться к новостям
БезопасностьAI Understanding брифинг

Исследование показало, что проверка разжигания ненависти в LLM менее последовательна в сценариях на урду

В препринте arXiv сообщается, что пять крупных языковых моделей изменили классификацию вредоносного контента в сценариях урду и английских переводах, что выявило измеримый пробел в безопасности «пропущено на урду».

5 min readRead the primary source
Source-page capture accompanying Study finds LLM hate-speech screening is less consistent in Urdu scripts
ПервоисточникИсточник записан
Издатель
arxiv.org
Ссылка на источник
arxiv.orghttps://arxiv.org/abs/2608.24191
Тип источника
Первичный документ — официальное объявление, документ, файл или собственная страница, которую мы читаем напрямую.
КонтекстПоймите это за 60 секунд

Начните здесь

Ключевые термины

Модель большого языка (LLM)
Языковая модель, обученная на массивных текстовых корпусах для генерации и анализа текста.
API (интерфейс прикладного программирования)
Структурированный способ отправки одной программной системой запросов и получения ответов от другой системы.
Набор данных
Коллекция структурированных или неструктурированных примеров, используемых для обучения, проверки или тестирования.
Проверьте себяВикторина по этике ИИ

Что случилось

An arXiv preprint tested GPT-4o, Claude Sonnet 4.5, Gemini 2.5 Flash, Qwen-2.5 and Llama-3.1 on six datasets covering Nastaliq Urdu, Roman Urdu, English and code-switched Urdu-English. The authors report that classifications changed between original Urdu-script content and English translations, with some harmful content passed as normal in Urdu. The paper also reports finding no dedicated Urdu papers across 205 papers in nine editions of the ALW/WOAH proceedings.

The paper, submitted to arXiv on Aug. 25, 2026, studies whether large language models classify hate speech consistently across Urdu script forms. Its experiment covers five models—GPT-4o, Claude Sonnet 4.5, Gemini 2.5 Flash, Qwen-2.5 and Llama-3.1—and six datasets spanning Nastaliq Urdu, Roman Urdu, English and code-switched Urdu-English. The source presents the work as an investigation of cross-script safety inconsistency, rather than as a product launch or a report of a specific platform incident.

The authors compare classifications of content in its original script with classifications of English translations. They report label instability across the five Urdu-script datasets ranging from 15.9% for Gemini 2.5 Flash to 31.6% for Qwen-2.5. In the paper’s terminology, a “Missed-in-Urdu” case is content that is flagged as harmful in English translation but classified as normal in its original Urdu script.

The reported Missed-in-Urdu rates range from 2.4% to 9.9%, with a median of 4.3% across the models. The abstract says smaller open-weight models showed substantially higher instability and missed-harm rates than the frontier closed models included in the test. These are results reported by the authors; the supplied source does not provide the underlying examples, per- breakdowns, confidence intervals or statistical tests.

The paper also examines the research record around Urdu safety evaluation. According to the abstract, the authors used the ACL Anthology API to enumerate 205 papers across nine ALW/WOAH editions and found zero dedicated Urdu papers during that period. The source does not identify every inclusion criterion in that enumeration or explain whether relevant work may have appeared under other venues, labels or research categories.

Подробности об источнике: arxiv.org ↗

Почему это важно

The study suggests that a content-moderation system can appear safer when tested through English translations than when it handles Urdu directly. That creates a practical evaluation problem for platforms and services serving multilingual users: aggregate safety scores may conceal failures concentrated in particular scripts or language varieties. The findings are claims from a single arXiv preprint and do not establish how widely the measured rates generalize beyond the datasets and models tested.

The central implication is that safety evaluation can depend on how language is represented, not only on what the text means. If a moderation model is tested mainly on English or on translated material, its measured performance may not reveal failures that occur when users write in Urdu script. A system could therefore receive reassuring aggregate results while handling materially different inputs unevenly.

That matters because content moderation is often used to decide whether material is blocked, escalated or allowed. The reported Missed-in-Urdu pattern describes a directionally important failure: text judged harmful after translation was sometimes passed as normal in the original script. The study does not show how often such cases occur in real-world platforms, what types of speech were involved, or what consequences followed, so its practical impact should be described as a documented evaluation concern rather than a quantified estimate of harm.

The work also gives safety teams a concrete measurement vocabulary. Overall accuracy or agreement can hide disagreements between scripts, while a cross-script comparison and a Missed-in-Urdu rate can expose one specific class of missed harmful content. Such measures could help organizations audit language coverage, compare model versions and identify where translation-based testing is insufficient. The source does not claim that the proposed score is a complete safety metric.

The reported gap between open-weight and closed models could influence how organizations choose or deploy multilingual moderation systems, but the comparison requires care. The models may differ in training data, tuning, access conditions and default behavior, and the abstract does not provide enough methodological detail to isolate the cause of the difference. The paper therefore supports scrutiny of model and script coverage, not a blanket conclusion that one model category is universally safer.

Interactive Mechanism

Интерактивный механизм: как он на самом деле работает

Изучите технологию, лежащую в основе этой разработки, в интерактивном режиме.

Document Size:128K tokens
Needle Placement Depth (Location in document):50% into text
Attention Context Buffer Map:
Target Fact (50%)
Equivalent Pages~320Standard book pages
Retrieval Accuracy99.9%Needle recall score
RAM / KV Cache5.1 GBMemory overhead
Prompt CachingActive~80% discount on reuse
Core takeaway: Million-token context windows allow querying whole codebases or legal archives in one prompt. However, KV cache memory scales with context length, making prompt caching crucial for real-time production.
Интерактивная проверка концепции+10 Points
AI Ethics Quiz

Why can ethical evaluation not be reduced to one model score?

Что посмотреть дальше

The important next questions are whether the results survive peer review and replication, how the datasets were assembled and labeled, and whether the same pattern appears in current production systems. Researchers and deployers should also examine Roman Urdu, code-switching, regional variation, translation quality, and model updates separately. The source does not establish which prompts, moderation thresholds, sampling procedures or mitigation methods would reduce the reported inconsistency.

The first uncertainty is methodological. The supplied source is an arXiv v1 preprint, not evidence that the work has completed peer review. Readers need the full paper’s sizes, label definitions, annotator procedures, translation process, prompts, decoding settings and statistical analysis to judge how robust the reported percentages are.

Replication will be especially important. The study tests five named models and a particular set of datasets, but the abstract does not establish whether the rates extend to newer model versions, other moderation systems, additional Urdu varieties or related languages and scripts. It also does not show whether errors are concentrated in particular topics, forms of hate speech, spelling patterns or levels of code-switching.

Deployers should watch for evaluations that test original-script inputs directly rather than treating English translation as a sufficient proxy. Useful follow-up work would compare script-specific results over time, report false positives alongside missed harms, and test whether mitigation changes improve one language variety without degrading another. The source does not identify a validated intervention.

The research-coverage finding also warrants verification. The authors report zero dedicated Urdu papers in 205 papers across nine ALW/WOAH editions, but the abstract does not provide the complete search record or explain how “dedicated Urdu” was defined. Future audits should make those criteria reproducible and examine work outside the proceedings. Until then, the paper’s strongest supported conclusion is that Urdu script coverage deserves explicit measurement in LLM safety testing.

Сопутствующие руководства и викторины

Этика ИИОбъяснение моделей искусственного интеллектаОбучение искусственному интеллектуПроверьте свои знания — пройдите бесплатную викторину по искусственному интеллектуНайдите термин ИИ в нашем глоссарии.Следите за трекером регулирования ИИ
Нашли это полезным?