Artificial intelligence can already search scientific literature, analyze data, write code, and propose hypotheses. But research is not only about producing answers: it also involves deciding what is worth exploring, changing direction when new evidence appears, and checking whether an idea makes sense in the real world.
That was the central topic of an event held on August 27 by the Allen Institute for AI, known as Ai2, to mark the expansion of its work with the Paul G. Allen Research Center at the Providence Swedish Cancer Institute. The collaboration on cancer research served as a starting point for discussing a broader question: what is scientific AI still missing?
The challenge is not just generating hypotheses
Current systems can find statistical patterns and suggest relationships that a researcher may not have considered. However, a striking relationship does not always represent an important discovery. It may be trivial, implausible, already known, or impossible to apply in a clinical context.
Bodhisattwa Prasad Majumder, a senior researcher at Ai2, described this human ability as scientific judgment. It is a mix of experience, intuition, and understanding of the problem that makes it possible to distinguish a promising lead from a result that does not deserve more time.
The goal does not have to be transferring this complete judgment to a machine. A more realistic objective is to build systems that allow scientists to incorporate their priorities, experience, and changes of mind during the research process.
The AutoDiscovery project showed why this collaboration is necessary. The system could find statistically surprising hypotheses, but some did not make biological or clinical sense when analyzed with expert knowledge. AI contributed signals; researchers helped interpret them.
AI can expand the space of possibilities, but experts still need to decide which ones deserve rigorous investigation.
Five challenges for scientific AI
1. Maintaining human control
Research rarely follows the initial plan. An experiment may produce an unexpected result, a new paper may change what is known, or a researcher may decide to incorporate another dataset and use a different tool.
The problem is that many AI agents are still difficult to direct during long projects. If the scientist changes the instructions, context, or available tools, the system may lose continuity or require the entire process to start over.
The solution requires agents with working memory, updatable context, and the ability to adapt their behavior without having to retrain the base model every time. In other words, it is not enough for AI to follow an initial assignment: it must support research as that research changes.
2. Deciding which tasks to delegate
Hoifung Poon, chief AI officer at Recursion, distinguished between two types of assistance. The first produces productivity gains: searching records, organizing information, summarizing papers, or structuring data. These tasks are relatively easy to describe and verify.
The second seeks creativity gains. Here, AI could propose an unknown biological mechanism for the team or suggest an unexpected experiment. The potential is enormous, but checking the quality of these ideas often requires additional analysis, experiments, and replication.
Abraham Flaxman shared a revealing case from his work as an editor of the Journal of Privacy and Confidentiality. A researcher used an AI tool to test algorithms included in their own published papers. The system detected a possible error and, after reviewing it, the researcher concluded that the criticism was correct and requested that the paper be withdrawn.
The lesson is not that we should automatically accept what AI says. Its value was in pointing out something that deserved to be examined.
Science is not only about looking for an answer. You have to discover, question, and verify.
3. Preventing AI from amplifying bad science
Analyzing more data at greater speed does not fix a poorly designed study or turn deficient data into reliable evidence. Stephen Salerno described AI as an amplifier, not an equalizer.
When the experimental design is solid, AI can increase its reach. But if there are weak assumptions, selection biases, or methodological errors, it can also cause those problems to be reproduced on a much larger scale.
That is why basic questions remain essential:
- Where did the data come from?
- Why was it collected?
- Who was included or excluded?
- Does the observed relationship imply causation, or only correlation?
- Does the result make sense from a scientific perspective?
Automation does not eliminate the need to think like a researcher. In fact, the more hypotheses and datasets AI can process, the more important quality controls become.
4. Connecting AI with experiments
One of the most ambitious scenarios is to move beyond the idea of AI working in isolation in front of a screen. The next step would be connecting its analyses to the laboratory so that the results of one experiment help decide what comes next.
Kyle Travaglini explained that some neuroscience projects study hundreds of cell types and thousands of genes that change over time. It is difficult for one person to track all the relevant relationships across the scientific literature. Agents could synthesize that evidence and prioritize hypotheses for the next phase.
In the long term, these systems could interact with laboratory instruments. The cycle would become tighter: AI analyzes the results, proposes a test, the laboratory runs the experiment, and the new evidence updates the research.
This does not mean handing the laboratory over to a machine. It means reducing the time between an observation, a hypothesis, and a verifiable test.
5. Building more complete computational models
Poon also raised the possibility of developing richer biological models. One option is to create a virtual tissue that connects individual-cell models with broader simulations of the organism.
That approach could become a bridge toward the concept of a virtual patient, capable of helping predict how a disease will progress or how a person will respond to a treatment. It is still a complex goal, because biology depends on multiple scales, interactions, and changing conditions.
Even so, the direction is clear: scientific AI should not be limited to summarizing what is already known. It should also integrate new evidence, represent biological systems with greater context, and help choose the next experiment.
Scientific AI needs to adapt, not just get bigger
The future of AI-assisted research does not depend solely on models with more parameters or better benchmark results. It also requires systems that can update their knowledge, accept new instructions, retain the context of a project, and show why they are proposing a hypothesis.
For scientists, the ability to direct and verify will be just as important as the ability to generate. A useful tool does not decide on its own which question deserves an answer. It helps explore more possibilities without losing the human judgment that turns an interesting result into reliable science.
The promise is not to replace the researcher, but to shorten the path between a good question, relevant evidence, and an experiment that can genuinely help us learn something new.
