Funding better evaluations of AI’s impact on wellbeing
AI systems have become central to how people work, learn, and seek emotional support, but the industry still lacks clear standards for how models should behave in sensitive conversations, such as when users seek companionship or navigate mental health crises. Assessing wellbeing is difficult because it requires context: a user in distress may not reveal self-harm immediately, and a response that is benign in one scenario (e.g., diet advice) could be harmful for a user with a history of disordered eating.
To address this, Anthropic is committing $5 million to fund independent research and open-source evaluation development. Grantees will receive direct funding, access to Anthropic's models, and technical support, while working fully independently and publishing their work as open-source projects that any developer can use.
Anthropic also shares guidance from its Safeguards team on what makes a wellbeing evaluation rigorous. Evaluations should clearly state what they measure and why it matters, involve clinical and subject-matter experts in design and validation, test for both overcompliance and overrefusal (precautions and harms), and reflect how users actually use AI, often by constructing realistic scenarios. The program hopes to attract clinicians, psychologists, methodologists, and others to this emerging field.