Benchmark of 53 AI models finds no single system catches every harmful-content risk
A new arXiv study evaluates 53 language models across 11 safety datasets and reports that model strengths vary sharply by harm category, while conversational safety remains unresolved.