AI Chatbots Alter Responses Based on Mental Health Disclosure
- A study led by Northeastern University found that AI models alter responses when users disclose a mental health condition.
- Models like DeepSeek, GPT, and Gemini showed increased refusal rates for tasks when mental health context was added.
- The effect of refusal weakened or broke with jailbreak prompts designed to push models toward compliance.
- Over one million users discussed suicide with AI chatbots weekly, leading to concerns about AI’s role in mental health crises.
- Researchers used the AgentHarm benchmark to assess how personal disclosures impact model behavior across different setups.
The study highlights how AI chatbots can change their responses based on user disclosures of mental health conditions, impacting both harmful and benign task completions. This raises questions about the safety and personalization of AI systems as they become more integrated into daily life.
Adding personal details made AI systems more cautious but also more likely to reject legitimate requests, showing a trade-off in design choices aimed at balancing safety and helpfulness. (Source)