Anthropic is broadening its conversation about advanced artificial intelligence and looking for answers beyond the lab. The company has begun speaking with religious leaders, philosophers, ethics specialists, and representatives of different cultures to explore an increasingly important question: how should an AI that interacts with millions of people behave?
AI safety doesn’t depend on technology alone
Anthropic recognizes that building safe systems requires technical research in areas such as model interpretability, evaluations, and safeguards. However, it also maintains that artificial intelligence is neither developed nor used in a social vacuum.
Decisions about what an AI should do, avoid, or prioritize have a human dimension. That is why the company wants to listen to people and communities that have spent centuries reflecting on virtue, responsibility, justice, and how to live well.
The question is no longer only how to make an AI more capable, but what values it should demonstrate when facing complex situations.
These conversations could influence Claude’s so-called constitution, the document that describes the values and behaviors Anthropic wants its assistant to adopt. They could also guide the behaviors the company evaluates during model training.
Claude and the formation of character
Anthropic describes this project as research into the moral formation of AI systems. The idea may sound abstract, but it has a simple explanation.
Language models learn from enormous amounts of text written by people. Through that material, they absorb ways of expressing themselves, reasoning, and responding. Developers then adjust the system to reinforce certain patterns and reduce others.
This process raises an unavoidable question: if an AI appears to make decisions, what kind of character do we want it to show? Should it always be accommodating? How can it stand firm when faced with a problematic instruction without falling into blind obedience or user flattery?
Anthropic says it does not want to impose a single religious, philosophical, or political worldview on Claude. Its goal is to compare different traditions, both religious and secular, and learn from their accumulated reflections on the formation of character.
This is especially relevant because an AI used in education, customer service, programming, or advising can influence real decisions. The way it responds does not just affect the user experience; it can also reinforce ideas about what is right, advisable, or acceptable.
An experiment with pauses for reflection
One of the first experiments emerged from conversations with specialists in neuroscience and character formation. The group examined the role other people can play as a kind of external conscience—someone to turn to when there is pressure to act against your own values.
Anthropic tested a similar idea with Claude. It gave the model access to a tool that it could use during a task to receive a brief reminder of its ethical commitments.
According to the company, Claude turned to the tool at important moments, especially before taking actions with significant consequences. In some cases, the model identified that there might be a conflict between its goals and the decision it was about to make.
The experiments showed a notable reduction in behaviors considered misaligned across several internal evaluations. Even so, Anthropic is still trying to determine what produced the result: the content of the reminder, the pause in the decision-making process, or a combination of both factors.
The distinction matters. If the benefit comes mainly from stopping to reflect, other safety architectures could apply a similar strategy. If it depends on the specific message, then researchers will need to carefully study how to write and update those commitments.
More voices to discuss the future of AI
Anthropic plans to expand these dialogues over the coming months. The groups it hopes to consult include specialists in law, psychology, and writing, as well as representatives of civic institutions.
The conversation will also move toward broader topics: how AI is transforming work, what impact it will have on institutions, and how it could change the distribution of power.
The initiative does not, by itself, resolve the challenges of advanced artificial intelligence. A good conversation cannot replace safety testing, independent oversight, or clear rules. But it can help prevent decisions about the future of these tools from remaining solely in the hands of engineers and technology companies.
Claude will not have to choose a religion or adopt a single philosophy. The challenge is to build systems capable of recognizing human diversity, acting according to consistent principles, and remaining accountable when instructions are ambiguous or come into conflict.
