Back to News
InnovationAI Understanding briefing

Study finds LLMs can produce brief counseling replies but struggle to know when to use them

A new EMNLP 2026 paper finds that large language models often overproduce long responses in counseling settings, where brief acknowledgments can better support attentive listening and continued disclosure.

By 6 min read
Primary-source image accompanying Study finds LLMs can produce brief counseling replies but struggle to know when to use them
The short version

A new EMNLP 2026 paper finds that large language models often overproduce long responses in counseling settings, where brief acknowledgments can better support attentive listening and continued disclosure.

What happened

A paper by Zhiyang Qi presents a cross-lingual analysis of minimal responses in counseling dialogues, including short backchannel cues and concise empathic statements. It reports that these responses are common in human-collected counseling datasets but substantially underrepresented in LLM-generated dialogues.

The paper studies a specific tension in AI counseling systems: human counselors often use very short responses, while dialogue systems and evaluation frameworks tend to favor replies that contain explicit information. The source describes minimal responses as including backchannel cues and concise empathic statements. Their purpose, according to the paper, is interactional rather than informational: they can signal attentive listening, express empathy and encourage clients to continue speaking.

The study conducts a systematic cross-lingual analysis across multiple counseling dialogue datasets. Its method first filters utterances using length and content, then applies contextual verification with a large language model. The source does not identify the datasets, languages, sample sizes or exact thresholds in the filtering process, so the supplied abstract supports the broad methodological description but not a detailed assessment of the study’s coverage or statistical strength.

The paper reports that minimal responses are common in human-collected datasets but substantially less common in LLM-generated ones. It then evaluates current LLMs in manually curated contexts drawn from exchanges in which human counselors used minimal responses. According to the source, strong commercial LLMs can generate brief responses when explicitly instructed, but have difficulty deciding when that response type is appropriate.

The source also reports weaker performance from counseling-specific models trained on synthetic data. Those systems tended to generate longer, more information-rich replies instead of minimal responses. In addition, the paper finds that LLM-based response-quality evaluations may undervalue minimal responses even when they are interactionally appropriate. The supplied material does not give model-by-model scores, baselines, error examples or details about how appropriateness was judged.

The paper is identified as a camera-ready version accepted to the EMNLP 2026 Main Conference and was submitted to arXiv on Aug. 25, 2026. That makes it timely within the stated news window, but the source remains an abstract and bibliographic page rather than the full paper. Claims about the strength, generality and reproducibility of the results therefore remain limited to what the abstract explicitly reports.

Read the primary source: arxiv.org

Why it matters

The findings suggest that response length and information density are incomplete measures of quality for AI systems used in emotionally sensitive conversations. A system that gives a longer answer may appear more helpful to an automated evaluator while missing the conversational function of listening, acknowledgment and making space for the client.

The central implication is that conversational quality cannot be judged only by how much useful-sounding information an AI system provides. In counseling dialogue, a short acknowledgment can serve a different function from advice, explanation or problem-solving. If an AI system fills every turn with guidance, it may reduce the space available for a person to describe their situation, even if the response reads as detailed and compassionate.

This matters for the design of systems intended to support mental-health conversations or other sensitive exchanges. Developers may need to evaluate whether a system recognizes conversational context, not merely whether it can produce a short sentence on command. The paper’s distinction between generating a minimal response and selecting one at the right moment points to a more specific capability: timing and interactional judgment.

The findings also raise questions about automated evaluation. If evaluators reward explicit content, longer answers or visible problem-solving, they may systematically score some appropriate brief responses as poor. That could create a feedback loop in which models are trained toward verbosity because verbosity is easier to measure, even when human dialogue data shows that shorter turns are common and useful.

The result is relevant beyond counseling because many forms of assistance depend on listening, turn-taking and user disclosure. However, the source directly studies counseling dialogues, and it does not establish that the same patterns hold in customer service, education, crisis response or other settings. Nor does it show that minimal responses by themselves improve clinical outcomes or constitute adequate mental-health support.

There are important safeguards and limitations to keep in view. A brief response can be appropriate in one context and inadequate in another, especially where a user expresses imminent danger, asks for concrete information or needs referral to a qualified professional. The source does not claim that less language is always better; it argues that systems and evaluators should account for cases in which less is interactionally appropriate.

What to watch next

The paper’s results should be tested across more counseling settings, languages, models and evaluation methods. The supplied source does not provide dataset sizes, model names, quantitative results or evidence that the findings translate into safer or more effective real-world counseling, leaving the practical impact uncertain.

A key next step is access to the paper’s full methodological and quantitative details. Useful checks would include the number and identity of datasets, the languages represented, the definition of a minimal response, the reliability of the contextual verification process and the size and composition of the manually curated evaluation set. Without those details, it is difficult to determine how broadly the reported pattern applies.

Further evaluation should compare human judgments with automated response-quality scores. Reviewers would need to assess not only whether a response is empathic or factually useful, but also whether it fits the preceding turn, preserves the speaker’s opportunity to continue and avoids giving inappropriate advice. The source reports a mismatch between these dimensions but does not describe the human rating procedure in the supplied text.

It will also be important to test whether models can learn context-sensitive brevity without becoming evasive, repetitive or emotionally flat. The paper reports that explicit instructions can make strong commercial LLMs generate minimal responses, but instruction-following alone may not show that the model understands when to use them. Future work should distinguish prompted behavior from a stable ability to select an appropriate conversational strategy.

The performance of counseling-specific models trained on synthetic data deserves particular scrutiny. The source says these models tended toward longer responses, but it does not explain how their synthetic training data were produced, what objectives shaped them or whether their behavior changes with different training mixtures. Comparing synthetic and human data could clarify whether the problem comes from data distribution, evaluation incentives, model architecture or another factor.

Finally, the practical question remains open: whether better recognition of minimal responses improves outcomes for people using AI counseling systems. The paper’s abstract does not report clinical validation, user outcomes, deployment evidence or comparisons with qualified human counseling. Until such evidence exists, the result is best understood as a useful warning about conversational evaluation and model behavior, not as evidence that any AI system is suitable for providing mental-health care.

Related guides & quizzes

ChatGPT & LLMsAI Models ExplainedAI EthicsPrompt EngineeringTest what you know — try a free AI quizLook up an AI term in our glossary
Found this useful?