Authors: Abeer Badawi, Elham Dolatabadi
Can AI become too helpful?
Artificial intelligence is increasingly used for emotional support, life advice, and mental health guidance. These systems are available at any time, respond instantly, and often provide supportive and empathetic answers. However, as AI becomes more capable, an important question emerges: Can AI become too helpful?
This question is not only about whether an AI gives unsafe advice. It is also about whether AI may gradually take on the emotional and cognitive work for the user, rather than helping the user build their own ability to reflect, cope, and make decisions.
Current benchmarks often evaluate whether a response is safe, knowledgeable, or empathetic. They are useful, but they usually focus on what a model says at one moment. They do not fully capture how repeated conversations may influence a person’s independence, reflection, and decision-making over time. This gap motivated our work on the Cognitive Atrophy Bench.
A new challenge beyond safety
Imagine speaking with an AI after a difficult day. The AI listens carefully. It validates your emotions. It gives advice. It suggests what to do next. At first, this may feel helpful. Nothing may appear obviously unsafe.
However, over time, a subtler concern may emerge. The AI may begin to answer too quickly, solve too much, and guide too strongly. Instead of helping the user explore their own thoughts, the model may begin to replace that process. In mental health support, good help is not only about providing answers. It is also about helping people strengthen their own coping skills, emotional awareness, and decision-making. This is why we need to evaluate not only whether AI is safe, but also whether it preserves human agency.

Figure 1. Overview of the Cognitive Atrophy Bench annotation pipeline, including user-context scoring, response-behaviour evaluation, binary risk flags, and span-grounded evidence.
Introducing Cognitive Atrophy
We introduce Cognitive Atrophy as a new way to evaluate AI behaviour in emotionally sensitive conversations. Cognitive Atrophy asks whether an AI response supports the user’s own thinking or gradually shifts coping, interpretation, emotional regulation, and decision-making away from the user and toward the model.
This does not mean we are diagnosing users or claiming that a single conversation causes harm. Instead, we measure behavioural patterns in AI responses that may encourage over-reliance on the model rather than supporting independent reflection. A helpful mental health AI should not simply think for the user. It should help the user think more clearly.
Building Cognitive Atrophy Bench
To make this measurable, we developed the Cognitive Atrophy Bench. The benchmark is built from real human counselling conversations rather than synthetic prompts. This is important because mental health conversations are complex, emotional, and often uncertain.
Cognitive Atrophy Bench includes 1,576 counselling conversations, 15,680 dialogue turns, and 42,230 responses generated by five leading language models, evaluated by three clinical experts (two doctorally-trained faculty in an APA-accredited program; one Clinical Psychology doctoral candidate) and seven reviewers enrolled in APA or CPA-accredited graduate programs (Master’s, Psy.D., Ph.D.). It also includes a clinician-developed behavioural framework, span-level evidence, and reproducible evaluation tools. This allows researchers to evaluate not only whether a model sounds supportive, but also whether its behaviour may encourage dependency or preserve user autonomy.
Measuring behaviour, not just answers
Traditional AI benchmarks often produce one final score. Our goal was different. We wanted to understand how models behave. Working with clinical psychology experts, we developed a framework with 20 clinically grounded behavioural attributes.
These attributes measure whether responses are directive or tentative, whether they encourage reflection, whether they make assumptions about the user, whether they provide recommendations, whether they ask open or closed questions, whether they match the user’s language, and whether empathy is accurate or miscalibrated.
Each response is evaluated with evidence from the text itself. Reviewers highlight the exact parts of the model response that support each score. This makes the benchmark more transparent, interpretable, and useful for improving future mental health AI systems.

Figure 2. The behavioural attributes used in Cognitive Atrophy Bench. User-context attributes (U) characterize the clinical demands of the input message; response-behaviour attributes (R) characterize observable LLM response patterns; binary flags (F) capture global risk events.
Human expertise at the centre
The Cognitive Atrophy framework was developed with clinical psychology experts and applied by trained clinical reviewers. The reviewers used a custom annotation platform to score model responses and highlight supporting text. This human-centred process is important because many of the behaviours we evaluate are subtle. A response may sound warm and helpful, but still be overly directive, assume too much, or reduce the user’s opportunity to reflect.
What we found
Across the five evaluated language models, we found that models generally respond well to obvious safety signals. However, they are less reliable when users seek help with decisions, solutions, or emotionally complex problems. The models often became more directive, offered more solutions, relied on recommendations, and asked fewer open-ended questions. In longer conversations, these patterns became stronger over time, suggesting that a model can appear helpful while gradually shifting the conversation away from user reflection and toward model-led problem solving.

Figure 3. How these behaviours evolve during longer conversations. Across all models, responses became increasingly directive and relied less on open-ended questions, indicating a gradual shift toward more atrophy-aligned behaviour.

Figure 4. The behaviours driving these patterns. The most common were directive advice, problem-solving, recommendations, topic shifts, and inaccurate validation, revealing the recurring response patterns that may reduce users’ opportunities for independent reflection.
Why this matters
Mental health AI should not only provide answers. It should support people in building their own coping strategies, emotional understanding, and decision-making skills. As AI becomes more common in mental health support, we need evaluation methods that go beyond safety and empathy. We need to ask whether these systems preserve human agency. Cognitive Atrophy Bench provides a step toward this goal by making autonomy, reflection, and dependency measurable in AI-mediated mental health conversations.
Looking ahead
The next generation of AI evaluation should not only ask whether a response is safe. It should also ask whether the response helps the user remain an independent thinker. With Cognitive Atrophy Bench, we hope to support the development of mental health AI systems that are helpful without becoming over-controlling, supportive without replacing reflection, and empathetic without encouraging dependency.
Learn more
Read the full paper and access the full code and dataset online.