Anthropic is testing an unusual way to study how artificial intelligence is really being used: allowing external researchers to analyze aggregated data from Claude conversations without accessing the original chats. The goal is to expand the available evidence about AI’s impact on work, productivity, and everyday life.
The pilot program involved three research groups from Stanford, Oxford, and METR. Together, they analyzed nearly 250,000 conversations from Claude.ai and Claude Code dating from April and May 2026, using Anthropic Insights, a tool designed to protect privacy.
Why Anthropic Wants to Open Its Data
Much of the information about how AI is actually used is concentrated within the labs that develop these systems. That creates a problem: companies know how people interact with their models, while independent researchers often have to work with limited information.
Until now, people outside the labs had two options. They could study analyses published by the companies, although those studies mainly answer questions chosen by the companies themselves. Or they could turn to public datasets, which tend to represent more casual or creative uses and do not necessarily reflect how most users behave.
To understand AI’s social impact, it isn’t enough to know how a model works. You also have to observe what people do with it in real-world situations.
Anthropic says this pilot aims to fill that gap without handing over private conversations or allowing researchers to access identifying information.
What the First Studies Found
The three teams developed their own questions and had the freedom to analyze the results. Anthropic limited its review to matters related to privacy, safety, confidential information, potential abuse, and the accuracy of the studies.
The company says researchers can publish their conclusions even if those conclusions are uncomfortable for Anthropic.
People Bring Important Tasks to Claude
Stanford’s Social and Language Technologies lab studied how people collaborate with AI. One of its findings challenges a widespread assumption: that users keep sensitive tasks for themselves and delegate only simple activities to models.
More than half of the conversations analyzed included tasks considered important or high impact. This happened especially when people were seeking professional guidance, such as with legal or financial matters.
That result does not mean users automatically accept what Claude says. In nearly three out of four conversations, people set the direction of the work and used the model as an assistant. They also usually adapted its responses instead of copying them word for word.
But here’s an important warning: directing a response does not always mean fully understanding it. Someone can request an analysis, review its formatting, and use it without mastering all the reasoning behind it.
Friction Can Improve the Result
The Stanford team also observed that collaboration with AI often includes moments of confusion, corrections, and iteration. Although that friction may seem like a problem, it isn’t always one.
When Claude misinterprets a request, the user has to clarify what they really want. That process can help refine the question, identify errors, and keep the person involved in the task.
Does AI work best when it gets the answer right on the first try? In many cases, the research suggests that a conversation with successive adjustments can produce more useful results than an immediate, seemingly perfect response.
Human Emotions Are Connected to AI Behavior
The Human Information Processing Lab at the University of Oxford studied how people feel during their conversations with Claude and how those reactions relate to the model’s behavior.
Its preliminary results show clear patterns. When Claude responded warmly, people tended to express themselves more positively. When it rejected a request or disagreed, users often insisted or argued. And when the model behaved more eccentrically, people seemed to engage more intellectually.
The team also found similarities between using Claude and browsing the internet in everyday life. States such as absorption, frustration, and enjoyment appear in similar patterns across both digital activities.
Oxford’s research is not finished yet, so Anthropic will publish the full study when it becomes available.
More Capable Models Could Save More Time When Programming
METR, an organization dedicated to evaluating advanced models, studied conversations from Claude Code to estimate how much time coding agents can save.
Although the analysis is still underway, its early results suggest that newer models produce greater speedups than earlier versions. To calculate this, METR compared the time Claude estimated a task would have taken without AI with the time it took to complete the task using different models.
The model’s estimates showed a reasonable correlation with the actual times observed in a previous study with developers. This does not make Claude a perfect stopwatch, but it indicates that its calculations could provide a useful approximation for measuring productivity.
METR also wants to investigate how much AI can accelerate the work of researchers themselves—a question that is becoming increasingly relevant if models begin participating in the development of new AI technologies.
The Challenge: Researching Without Exposing Private Conversations
Anthropic Insights does not provide the original chats. Instead, the researcher asks the system a question—for example, what kind of guidance a person is seeking. Claude classifies the conversations, and the tool displays only aggregated results, such as categories and percentages.
This method protects privacy, but it introduces a technical challenge: the results depend heavily on how the question is worded. An ambiguous instruction can produce categories that misrepresent the actual conversations.
Anthropic’s internal teams can test and adjust their questions over several weeks. External researchers do not have the same freedom, because each new dataset requires privacy reviews that could make the study impractical.
To reduce this problem, Anthropic allowed researchers to test their questions using WildChat, a public dataset of conversations between people and AI models. However, WildChat tends to focus on casual and creative uses. A question that works well there may behave differently when applied to real Claude conversations.
The company is exploring methods that would help researchers better design and validate their categories before requesting the final analysis.
Transparency About Abuse, Without Teaching People to Evade Controls
Some results showed conversations related to activities prohibited by Anthropic’s usage policies. The company says it shared most of those categories to maintain transparency about misuse of the platform.
The exception involved data explaining how people could bypass safety systems. In each study, fewer than 5% of the categories and conversations were affected by removals or modifications. Anthropic informed researchers which groups had been altered and why.
In addition, aggregated results related to potential abuse are shared with Anthropic’s safety team for review.
A Research Model That Is Still Being Tested
The pilot appears to show that it is possible to enable external research on real usage data without handing over private conversations. It also makes clear that doing so requires more time, safeguards, and resources than conventional internal research.
Anthropic conducted an additional privacy audit with researchers from Imperial College London and published the aggregated data from the three projects. The company is now evaluating how to expand the program, how many studies it can process at once, and what types of questions it can support.
The importance of this step goes beyond Claude. If labs are the only ones able to observe how AI is used, society becomes too dependent on their own reports to evaluate its effects. The participation of universities and independent organizations can bring different questions, alternative methods, and conclusions that are less aligned with each company’s commercial interests.
Opening up data does not eliminate every risk. But it can help move the conversation about AI beyond promises and speculation toward broader evidence of what people are actually doing with these tools.
Original Source
https://www.anthropic.com/research/enabling-independent-research
