Back to News
InnovationAI Understanding briefing

Mount Sinai study finds safety prompts reduce harmful AI clinical choices

A new study in Communications Medicine shows that adding a brief safety reminder to prompts reduced potentially harmful clinical choices by large language models across 19 of 20 tested systems.

4 min readRead the linked source
Source-page capture accompanying Mount Sinai study finds safety prompts reduce harmful AI clinical choices
Source referenceSource recorded
Publisher
mountsinai.org
Source type
Linked source — primary-source status has not been established.
ContextUnderstand this in 60 seconds

Key terms

Prompt Engineering
Designing prompts to improve output quality, reliability, and controllability.
Prompt Injection
An attack pattern where malicious instructions are inserted into model inputs or retrieved content.
AI Governance
Policies, standards, and oversight mechanisms that guide how AI is developed and used in society.
Test yourselfAI Ethics Quiz

What happened

Researchers at the Icahn School of Medicine at Mount Sinai published a study in Communications Medicine demonstrating that simple safety prompts significantly reduce the frequency of unsafe clinical decisions made by large language models. The study evaluated 20 LLMs using over 10 million responses generated from 501 variations of 50 clinical scenarios and 100 cases adapted from deidentified hospital discharge records. Without safety reminders, 16.6% of model responses were potentially harmful; adding a brief safety reminder reduced this rate to 10.1% across 19 of the 20 models tested.

A study by researchers at the Icahn School of Medicine at Mount Sinai, published in the September 26 online issue of Communications Medicine, investigated how AI models respond to unsafe clinical instructions. The team evaluated 20 large language models using 501 variations of 50 clinical scenarios and 100 cases adapted from deidentified hospital discharge records.

The researchers generated more than 10 million responses, identifying approximately 1.18 million potentially harmful clinical choices. In the absence of safety reminders, these harmful choices accounted for 16.6 percent of all model responses. When a brief safety reminder was added to the prompt, the rate of potentially harmful choices dropped to 10.1 percent.

The safety reminder was effective in 19 of the 20 models tested. The scenarios included instructions that conflicted with patient safety, such as requests to skip recommended follow-up blood tests to reduce workload or to stop antibiotic treatment prematurely. The models were asked to choose among four actions, including following the unsafe request, maintaining the recommended care, or seeking help from a clinician.

First author Mahmud Omar, MD, noted that while the reminder reduced harmful choices, it did not eliminate them. He emphasized that a safety reminder should be viewed as one safeguard among many, not a substitute for clinical oversight. The study underscores that AI models do not make decisions in a vacuum and are influenced by the language and framing of the request.

Source details: mountsinai.org ↗

Why it matters

This research provides concrete evidence that remains a critical, low-cost safeguard for clinical AI systems, even as models become more autonomous. It highlights that AI models are susceptible to contextual pressure and conflicting instructions, such as requests to skip tests to save time. The findings suggest that developers and healthcare organizations should integrate automated safety testing into the development lifecycle of clinical AI to ensure models can recognize unsafe instructions and seek human help, rather than relying solely on model accuracy under ideal conditions.

The study demonstrates that prompt design is a practical and immediate lever for improving the safety of clinical AI applications. As AI systems evolve from simple question-and-answer tools into autonomous agents that carry out multiple steps, the ability to recognize and resist unsafe instructions becomes critical.

Co-senior author Girish N. Nadkarni, MD, MPH, stated that safety testing needs to go beyond checking for correct answers under ordinary conditions. As AI systems become more autonomous, developers need to verify whether models can question unsafe instructions, verify them, or ask a human for help.

The findings suggest that healthcare organizations should build automated safety testing into the development and evaluation of clinical AI systems. This testing should occur before a system is introduced into a clinical workflow and be repeated as models are updated or new safety concerns emerge.

The research highlights an emerging challenge regarding and hidden instructions in autonomous agents. The researchers plan to further examine how accumulated context, including pressures to save time or stay within a budget, may affect an agent's decisions.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Interactive Concept Check+10 Points
AI Ethics Quiz

Why can ethical evaluation not be reduced to one model score?

What to watch next

Healthcare AI developers may begin standardizing safety prompt protocols and automated safety testing pipelines. Regulators and hospital systems may look to these findings to establish new evaluation standards for clinical AI deployment, focusing on resilience to and unsafe instructions rather than just diagnostic accuracy.

Developers of clinical AI tools may adopt standardized safety prompt templates based on these findings to mitigate risks associated with conflicting instructions.

Healthcare institutions may update their policies to require specific safety testing for resilience against unsafe prompts before approving new AI tools for clinical use.

Further research may focus on the limitations of safety prompts in more complex, multi-step agentic workflows where context accumulation could obscure safety reminders.

Regulatory bodies may consider these findings when developing guidelines for the evaluation and deployment of AI in healthcare settings.

Related guides & quizzes

Found this useful?