Can artificial intelligence discover something important in a dataset? Yes, but finding an anomaly is not the same as proving a scientific finding. That was one of the main lessons from an academic challenge in which University of Washington students worked with AutoDiscovery, an AI agent created by Ai2 to explore scientific research.
The experience highlighted an increasingly necessary skill: learning to question AI, review its evidence, and decide when a hypothesis deserves deeper investigation.
AutoDiscovery proposes hypotheses, but it does not replace the researcher
AutoDiscovery analyzes datasets, formulates hypotheses, runs computational experiments to test them, and ranks the results according to Bayesian surprise. In simple terms, this metric measures the distance between what the model expected to find and what the data actually show.
A highly unexpected result can be a valuable clue. It can also be a coincidence, a data error, or the consequence of the model not fully understanding the scientific context. How do you distinguish a possibility from a discovery? That is where human work begins.
In the course GenAI for Science: Ai2-UW Materials Challenge, professor Luna Yue Huang invited her students to test this approach. Twenty-five teams proposed datasets and research questions. Ten projects were selected to work with AutoDiscovery for several weeks, with support from instructors and Ai2 researchers.
The teams had to review the generated hypotheses, compare them with the scientific literature, identify reasoning flaws, and assess whether the tool actually reduced the manual work involved in analysis.
AI could accelerate exploration, but it could not decide on its own which result was scientifically sound.
Four examples of AI-assisted research
The possible effect of battery testing
One team analyzed data on the aging of lithium-ion batteries and found a counterintuitive clue: the periodic tests used to evaluate a cell’s health could be accelerating its degradation.
However, the result showed a correlation, not a causal relationship. In other words, the data indicated that both phenomena occurred together, but they did not prove that the tests were responsible for the additional wear.
Confirming the hypothesis would require designing a controlled experiment. AutoDiscovery’s contribution was to point out a possibility that might otherwise go unnoticed, while the students had to determine what the next logical step should be.
An exception in atomic charge
Another group worked with six million simulations of polymers. The team found an apparent exception to a rule documented decades ago concerning two methods for estimating an atom’s charge.
The difference between the two methods is usually greater in atoms with a strong tendency to attract electrons, such as oxygen and nitrogen. The polymer data suggested a different behavior. When the students traced the origin of the pattern through the literature, they found no published explanation for the divergence.
In this case, AI helped turn a statistical signal into an open scientific question. Even so, the fact that an explanation does not appear in the articles consulted does not automatically mean that nobody has solved it or that the phenomenon is real. Bibliographic and experimental verification remain essential.
Proteins that respond to light
A third team studied proteins capable of activating under illumination. To build them, researchers must insert a light-sensitive segment at a precise position in the protein, a process that may require testing numerous designs in the laboratory.
AutoDiscovery analyzed which characteristics distinguished successful insertions from those that failed. Its hypothesis was that functional designs anchored the light-sensitive segment through a few strong points rather than many weak ones.
The students compared that prediction with their own experimental results. The pattern held across all the designs tested and matched an established principle concerning protein binding. Here, the tool did more than produce an interesting idea: it generated a hypothesis compatible with real evidence and prior knowledge.
When synthetic data confuse the model
The experience also revealed AutoDiscovery’s limitations. In a set of 24,000 simulated nuclear reactor configurations, approximately half of the generated hypotheses turned out to be illogical, and none seemed promising enough to pursue.
The students linked the problem to how the system works. AutoDiscovery calculates surprise based on measurements from the real world, but the simulation’s synthetic numbers did not represent that type of observation. The model was evaluating the data against an unsuitable reference point.
When the team used real reactor physics measurements obtained through experiments, performance improved considerably. The tool produced nearly 15 potentially useful leads. The students then had to prioritize them according to their physical plausibility, logical consistency, and connection to previous research.
AI can generate ideas, but not scientific judgment
The project results pointed in the same direction: AutoDiscovery was most useful for proposing lines of research that students might not have considered. However, the more unexpected a result was, the greater the need for a person to examine where it came from.
Interpreting evidence, understanding how the system reached a conclusion, and detecting incorrect assumptions require scientific expertise. It also takes the ability to design experiments, distinguish causation from correlation, and recognize when a result contradicts established physical or biological principles.
This difference is crucial. A generative AI can review millions of combinations in a short time, but it does not possess, by itself, the sense of relevance needed to answer questions such as these:
- Is the signal strong enough to justify more work?
- Is there a simpler alternative explanation?
- Could the result be caused by a flaw in the dataset?
- Can it be reproduced in another experiment?
- What evidence would be needed to turn the hypothesis into a conclusion?
A new way to teach science with AI
For Huang, the challenge was not simply about teaching students how to use a tool. The goal was to prepare them for a different way of conducting research: asking AI to explore, generate hypotheses, and find patterns, while the researcher questions the results and decides which ones deserve attention.
Instead of using AI to complete a perfectly defined task, the students used it as a scientific exploration partner. This shift moves human value toward skills such as critical thinking, experimental design, evidence-based reasoning, and the ability to justify conclusions.
The teams transformed raw datasets into hypotheses supported by literature reviews, quantitative analyses, experimental validation plans, and reasoned reports. AI accelerated part of the process, but the quality of the results depended on the students’ curiosity, persistence, and judgment.
When AI can generate ideas quickly, scientific training does not lose value. It becomes more important, because someone has to know which ideas are worth testing.
The lesson also applies outside the laboratory. In business, data analysis, or product development, a tool can find patterns quickly, but its results need context, verification, and a responsible decision. Automation can broaden the search, but it does not eliminate the responsibility to evaluate what is found.
AutoDiscovery represents an interesting vision for scientific AI: systems capable of exploring vast amounts of information without completely hiding how they reach their results. But the experience at the University of Washington is a reminder that technical transparency is not enough. You also need researchers prepared to question the answers, understand the limitations of the data, and turn an anomaly into reliable knowledge.
