LongGuard study finds safety guardrails lose more than half their unsafe-input recall on long context
An arXiv paper reports that safety guardrails’ ability to detect unsafe content falls sharply as input length grows, and proposes training-free mitigations that improve average results in its tests.