Back to News
InnovationAI Understanding briefing

Systematic review maps the growing use of large language models in mental health

A systematic review surveys how large language models are being studied for mental-health analysis, risk assessment, therapy support and multimodal monitoring, while stressing unresolved ethical and regulatory challenges.

By 5 min read
An empty mental-health clinic consultation room with two chairs, a plain table, and an unbranded closed laptop beside a small audio microphone.
The short version

A systematic review surveys how large language models are being studied for mental-health analysis, risk assessment, therapy support and multimodal monitoring, while stressing unresolved ethical and regulatory challenges.

What happened

The paper reviews existing research on large language models in mental health, covering social-media analysis, clinical conversational agents, therapy-support tools, prompt engineering, multimodal learning and ethics.

The primary source is an arXiv record for a paper by Yisong Chen, Yifan Gao, Sijing Yu, Chuqing Zhao and Yang Lu. The record identifies the work as a systematic review and says it was submitted on May 31, 2026. It also lists a Journal of Industrial Integration and Management reference from 2025, but does not explain the relationship between that publication reference and the later arXiv submission.

The source establishes that this is a synthesis of prior work, not an announcement of a new model, product, clinical service or regulatory decision. According to the abstract, the review brings together interdisciplinary studies using several kinds of information: social-media posts, electronic medical records and multimodal inputs. The authors say this body of work has been used for early detection of depression, suicide-risk assessment, personalized therapy support and generation of psychoeducational content.

Those are claims about the literature surveyed by the authors. The supplied source does not identify the individual studies, participant populations, model versions, clinical settings, evaluation metrics or effect sizes, so it cannot independently establish how well any of these applications work. The review also highlights advances in language models and annotation strategies that the authors say can improve interpretability and clinical relevance.

It identifies prompt engineering as important for adapting systems to a domain, and discusses multimodal fusion involving text, speech and sensor data for mental-health diagnosis and monitoring. The abstract describes these as areas of advancement or emergence, but provides no head-to-head comparison, prospective trial, deployment record or evidence that one approach is safer or more accurate than another.

The final component is governance. The authors describe ethical, sociotechnical and regulatory challenges as ongoing and advocate frameworks for safe, equitable and accountable deployment in real-world mental-health care. The source does not say that such a framework has been adopted by a health system, regulator or model provider. It also does not state that any system reviewed is available to patients, clinicians or the public, or that the paper reports a new safety protocol.

Read the primary source: arxiv.org

Why it matters

The review brings together applications involving depression detection, suicide-risk assessment, personalized therapy support and psychoeducation, but the supplied source does not establish that these uses are clinically validated or ready for routine care.

The significance of the review comes from the range of uses it places in one frame. The abstract connects language models with analysis of mental-health signals, conversations with patients or users, therapy support and educational content. It also includes applications involving depression and suicide risk. If systems are used in these settings, their outputs could influence how concerns are noticed, prioritized or communicated. That makes the difference between a promising research use and a clinically dependable tool important, even though the supplied source does not measure that difference itself.

The review is potentially useful because it links technical choices to clinical and social questions. Prompt design, annotation and multimodal fusion are not presented as isolated engineering techniques; the authors discuss them in relation to interpretation and clinical relevance. A system may produce fluent language while still being difficult to evaluate or explain. The source does not demonstrate that the reviewed methods solve those problems, but it makes clear that technical performance alone is not the complete standard the authors believe should govern mental-health applications.

The paper's focus on ethics and accountability also matters for the public because the source describes uses involving sensitive personal information and high-consequence judgments. The abstract names social-media posts, medical records, speech and sensor data as inputs considered in the literature. It does not provide details about consent, retention, access controls, bias testing, clinician oversight or appeal mechanisms. These omissions are meaningful unknowns, not evidence that the underlying studies failed to address them. They show why the review's call for safe and equitable deployment cannot be treated as proof that those safeguards already exist.

This is consequential research rather than a product announcement. Its immediate public value is to organize a fast-moving field and identify where deployment claims would require scrutiny. The source does not establish improved patient outcomes, reduced suicide deaths, diagnostic accuracy, cost savings or clinical approval. Readers should therefore treat the paper as a map of applications and challenges, not as evidence that large language models are ready to replace mental-health professionals or independently make care decisions.

What to watch next

The key questions are whether the underlying studies demonstrate reliable performance in real-world settings, how they handle safety and accountability, and whether the review's proposed governance frameworks lead to measurable protections.

The first priority is the review's underlying evidence. The abstract does not provide its search strategy, inclusion and exclusion criteria, number of studies, quality assessment or treatment of conflicting results. Those details would show whether the review distinguishes controlled evaluations from demonstrations, retrospective analyses and conceptual proposals. They would also clarify whether the cited evidence is concentrated in a few languages, populations or clinical environments, which would affect how broadly its conclusions can be applied.

Future reporting should examine the specific performance and failure patterns behind the uses named in the source. For depression detection and suicide-risk assessment, that includes whether systems were tested prospectively, how false positives and false negatives were handled, and whether outputs were judged against qualified clinical assessment. For therapy support and psychoeducation, it includes whether responses were evaluated for harmful advice, inappropriate confidence and consistency across cases. None of these results is supplied in the abstract, so they remain open questions.

The multimodal direction warrants particular scrutiny. Combining text, speech and sensor data could add information, but it could also increase the amount and sensitivity of data collected. The source calls multimodal fusion an emerging technique and says it may improve diagnosis and monitoring; it does not report a validated system, explain which sensors were used, or establish that added modalities improve outcomes. Evidence about consent, data security, interpretability and performance across different groups will be necessary before claims of practical benefit can be assessed.

Finally, watch whether the governance frameworks advocated by the authors become concrete requirements. The source does not name a regulator, health provider or company adopting them, and it gives no implementation timetable. The paper's publication status also deserves careful reading because the record lists a 2025 journal reference alongside a 2026 arXiv submission. Until the full methods, cited evidence and any subsequent evaluations are available, the defensible conclusion is that the review identifies important opportunities and risks, while the real-world effectiveness and safety of the applications remain unresolved.

Related guides & quizzes

Found this useful?
The Weekly Briefing

Get the AI stories that actually matter.

One useful email a week — what changed in AI, why it matters, plus tools, guides, opportunities, and practical ways to take action.

Free · No spam · Unsubscribe in one click