Anthropic offers $5M to evaluate AI's impact on wellbeing
Anthropic has launched a $5 million grant program to fund independent, open-source research into how AI affects user wellbeing, helping the industry measure and mitigate psychological risks.
Anthropic is launching a $5 million grant initiative to support independent research focused on how artificial intelligence systems affect the wellbeing of their users. The program will offer selected researchers direct financial backing, technical support, and access to Anthropic's proprietary models. Grantees will operate with complete independence, and their final evaluations and benchmarks must be published as open-source projects available to any developer. Applications for the funding are due by September 21, with selected finalists notified to submit full proposals by October 5.
The initiative addresses a critical gap in current AI safety methodologies. While traditional evaluations can easily flag a single inaccurate or inappropriate response, assessing a user's psychological wellbeing requires analyzing long-term, multi-turn conversations. For instance, a user experiencing a mental health crisis or seeking emotional companionship might not immediately disclose their distress. Standard safety filters might also fail to recognize when a helpful response, such as offering diet advice, becomes harmful if the user has a history of disordered eating.
To guide applicants, Anthropic's safeguards team outlined several criteria for rigorous wellbeing evaluations. The company is seeking benchmarks that clearly define what constitutes a pass or fail, involve clinical and psychological experts in their design, and validate automated grading systems against real-world specialists. Crucially, these evaluations must test for both overcompliance and overrefusal to ensure models do not become overly restrictive while trying to prevent harm.
For AI practitioners and developers, this program promises to deliver standardized, open-source tools to measure how conversational agents influence human emotion and mental health. By moving beyond basic content moderation, developers will gain access to sophisticated frameworks capable of tracking nuanced, multi-turn interactions. This collective effort aims to establish industry-wide guardrails, ensuring that conversational models like Claude can safely navigate sensitive topics without triggering unnecessary fallbacks or causing unintended psychological harm.
This is our own summary of reporting by Anthropic



