Anthropic wants to answer an increasingly urgent question: how does artificial intelligence really affect people’s well-being? The company announced a $5 million grant program to fund independent research on this topic.
The project will offer direct funding, access to Anthropic’s models, and technical support. The selected teams will work independently and must publish their evaluations as open-source projects, so any developer can use them.
Why Measuring Well-Being Is So Complex
AI is already part of work, learning, and problem-solving. It has also become a conversational companion for many people and, during difficult moments, may be used as a source of emotional support.
But evaluating whether a response helps or causes harm is not as simple as checking whether a fact is correct. Context can completely change the meaning of a conversation.
For example, someone going through a crisis may not mention thoughts of self-harm in their first messages. That signal may appear later, after an extended conversation, when the risk is more evident. A useful evaluation must be able to detect that evolution.
The same applies to topics such as nutrition and exercise. Advice about balanced diets might seem reasonable for someone who wants to lose weight. However, it could be inappropriate or harmful if that person has a history of an eating disorder.
In conversations about well-being, a response cannot be evaluated in isolation. The history, the user’s intent, and the evolution of the dialogue also matter.
What Kind of Research Is Anthropic Looking For
The program is aimed at specialists from different fields, including clinical professionals, psychologists, methodologists, and experts in AI system evaluation.
Anthropic published a guide outlining the criteria that, from its perspective, can make a well-being evaluation rigorous and useful for the industry. They include:
- Clearly explain what is being measured, what counts as a correct or incorrect outcome, and why that outcome matters.
- Include clinical experts and subject-matter specialists during the study’s design and validation.
- Evaluate both safeguards and potential harms, including over-compliance and over-refusal.
- Represent real-world uses of AI through multi-turn conversations in which the context and level of risk change.
- Verify that the systems used to grade responses reliably align with human expert evaluations.
The Challenge of Avoiding Two Opposite Mistakes
A model can make a mistake by responding too much. For example, it might provide specific recommendations when it should act more cautiously and guide the user toward professional help.
But it can also fail by refusing too much. A response that rejects any conversation related to mental health, nutrition, or emotions does not necessarily protect the user. In some cases, it may leave them without basic guidance or cause them to abandon the conversation.
That is why the goal is not to build an AI that always says no. The aim is to develop systems that can better interpret context and respond proportionally to the level of risk.
Open Evaluations for an Industry with Better Standards
Anthropic says it is working on safeguards to identify sensitive conversations and respond more appropriately in Claude. It also publishes research on the types of conversations people have with the model.
However, the company acknowledges that this is a nuanced field with important consequences. Standards will need to evolve alongside the models and the new ways people use them.
Funding independent evaluations could help reduce a common problem in technology: each company measures its own systems using different criteria that are difficult to compare. If the results are published as open tools, researchers and developers will be able to use them to build more consistent tests.
Applications for the program must be submitted by September 21. The people and teams selected to submit full proposals will be notified on October 5.
The initiative does not solve the challenge of measuring well-being on its own, but it does recognize something essential: safe AI is not defined only by the accuracy of its responses. We must also consider how those responses affect real people, especially when they are going through vulnerable situations.
