Volver a Noticias
SeguridadAI Understanding sesión informativa

Study finds LLM hate-speech screening is less consistent in Urdu scripts

An arXiv preprint reports that five large language models changed harmful-content classifications across Urdu scripts and English translations, revealing a measurable “Missed-in-Urdu” safety gap.

Por 5 min read
AI-generated editorial illustration accompanying Study finds LLM hate-speech screening is less consistent in Urdu scripts
La versión corta

An arXiv preprint reports that five large language models changed harmful-content classifications across Urdu scripts and English translations, revealing a measurable “Missed-in-Urdu” safety gap.

que paso

An arXiv preprint tested GPT-4o, Claude Sonnet 4.5, Gemini 2.5 Flash, Qwen-2.5 and Llama-3.1 on six datasets covering Nastaliq Urdu, Roman Urdu, English and code-switched Urdu-English. The authors report that classifications changed between original Urdu-script content and English translations, with some harmful content passed as normal in Urdu. The paper also reports finding no dedicated Urdu papers across 205 papers in nine editions of the ALW/WOAH proceedings.

The paper, submitted to arXiv on Aug. 25, 2026, studies whether large language models classify hate speech consistently across Urdu script forms. Its experiment covers five models—GPT-4o, Claude Sonnet 4.5, Gemini 2.5 Flash, Qwen-2.5 and Llama-3.1—and six datasets spanning Nastaliq Urdu, Roman Urdu, English and code-switched Urdu-English. The source presents the work as an investigation of cross-script safety inconsistency, rather than as a product launch or a report of a specific platform incident.

The authors compare classifications of content in its original script with classifications of English translations. They report label instability across the five Urdu-script datasets ranging from 15.9% for Gemini 2.5 Flash to 31.6% for Qwen-2.5. In the paper’s terminology, a “Missed-in-Urdu” case is content that is flagged as harmful in English translation but classified as normal in its original Urdu script.

The reported Missed-in-Urdu rates range from 2.4% to 9.9%, with a median of 4.3% across the models. The abstract says smaller open-weight models showed substantially higher instability and missed-harm rates than the frontier closed models included in the test. These are results reported by the authors; the supplied source does not provide the underlying examples, per-dataset breakdowns, confidence intervals or statistical tests.

The paper also examines the research record around Urdu safety evaluation. According to the abstract, the authors used the ACL Anthology API to enumerate 205 papers across nine ALW/WOAH editions and found zero dedicated Urdu papers during that period. The source does not identify every inclusion criterion in that enumeration or explain whether relevant work may have appeared under other venues, labels or research categories.

Lea la fuente principal: arxiv.org

Por qué es importante

The study suggests that a content-moderation system can appear safer when tested through English translations than when it handles Urdu directly. That creates a practical evaluation problem for platforms and services serving multilingual users: aggregate safety scores may conceal failures concentrated in particular scripts or language varieties. The findings are claims from a single arXiv preprint and do not establish how widely the measured rates generalize beyond the datasets and models tested.

The central implication is that safety evaluation can depend on how language is represented, not only on what the text means. If a moderation model is tested mainly on English or on translated material, its measured performance may not reveal failures that occur when users write in Urdu script. A system could therefore receive reassuring aggregate results while handling materially different inputs unevenly.

That matters because content moderation is often used to decide whether material is blocked, escalated or allowed. The reported Missed-in-Urdu pattern describes a directionally important failure: text judged harmful after translation was sometimes passed as normal in the original script. The study does not show how often such cases occur in real-world platforms, what types of speech were involved, or what consequences followed, so its practical impact should be described as a documented evaluation concern rather than a quantified estimate of harm.

The work also gives safety teams a concrete measurement vocabulary. Overall accuracy or agreement can hide disagreements between scripts, while a cross-script comparison and a Missed-in-Urdu rate can expose one specific class of missed harmful content. Such measures could help organizations audit language coverage, compare model versions and identify where translation-based testing is insufficient. The source does not claim that the proposed score is a complete safety metric.

The reported gap between open-weight and closed models could influence how organizations choose or deploy multilingual moderation systems, but the comparison requires care. The models may differ in training data, tuning, access conditions and default behavior, and the abstract does not provide enough methodological detail to isolate the cause of the difference. The paper therefore supports scrutiny of model and script coverage, not a blanket conclusion that one model category is universally safer.

Qué ver a continuación

The important next questions are whether the results survive peer review and replication, how the datasets were assembled and labeled, and whether the same pattern appears in current production systems. Researchers and deployers should also examine Roman Urdu, code-switching, regional variation, translation quality, and model updates separately. The source does not establish which prompts, moderation thresholds, sampling procedures or mitigation methods would reduce the reported inconsistency.

The first uncertainty is methodological. The supplied source is an arXiv v1 preprint, not evidence that the work has completed peer review. Readers need the full paper’s dataset sizes, label definitions, annotator procedures, translation process, prompts, decoding settings and statistical analysis to judge how robust the reported percentages are.

Replication will be especially important. The study tests five named models and a particular set of datasets, but the abstract does not establish whether the rates extend to newer model versions, other moderation systems, additional Urdu varieties or related languages and scripts. It also does not show whether errors are concentrated in particular topics, forms of hate speech, spelling patterns or levels of code-switching.

Deployers should watch for evaluations that test original-script inputs directly rather than treating English translation as a sufficient proxy. Useful follow-up work would compare script-specific results over time, report false positives alongside missed harms, and test whether mitigation changes improve one language variety without degrading another. The source does not identify a validated intervention.

The research-coverage finding also warrants verification. The authors report zero dedicated Urdu papers in 205 papers across nine ALW/WOAH editions, but the abstract does not provide the complete search record or explain how “dedicated Urdu” was defined. Future audits should make those criteria reproducible and examine work outside the proceedings. Until then, the paper’s strongest supported conclusion is that Urdu script coverage deserves explicit measurement in LLM safety testing.

Guías y cuestionarios relacionados

Ética de la IAModelos de IA explicadosEntrenamiento de IAPon a prueba lo que sabes: prueba un cuestionario gratuito sobre IABusque un término de IA en nuestro glosario
¿Encontró esto útil?