Anthropic presents two experiments showing how Claude can reduce the time and technical expertise needed to advance biomedical research. The model designed proteins capable of binding to biological targets and analyzed experimental chemistry data with results comparable to those from specialized laboratories.
The difference? Claude did not merely answer questions. In one case, it coordinated specialized models, ran optimization cycles, and proposed candidates for laboratory validation. In the other, it interpreted instrument files without relying on the manufacturer’s proprietary software.
Claude Designs Proteins from Scratch
A minibinder is a small protein designed to attach with high precision to another target protein. This type of binding is fundamental to many therapies because it can block a signal, activate a process, or carry a molecule to a specific site.
Designing a binder from scratch, also called de novo design, usually requires weeks or months of work involving simulations, sequence selection, structural prediction, and experimental testing. Machine learning models have sped up this process, but they typically require careful orchestration by specialists.
Anthropic evaluated Claude Opus 4.8 and Mythos Preview in a campaign against 15 targets. The designs were produced computationally and then independently manufactured and tested by Adaptyv Bio and Twist Bioscience.
The main results were:
- Claude designed proteins capable of binding to 14 of the 15 targets.
- 354 confirmed binders were produced from 1,320 designs.
- The overall success rate reached 22.6% for Opus 4.8 and 26.7% for Mythos Preview in multi-target mode.
- Mythos Preview reached 35.1% when it worked with a single target in separate sessions.
- The usual success rate for protein design campaigns is around 10% to 15%.
The success rate, known as the hit rate, indicates what proportion of designs ultimately shows real binding in the laboratory. It does not mean that all of these candidates are drugs, but it does mean they cleared an initial stage that traditionally consumes a great deal of time.
Results Against Difficult Targets
For the RBX1 target, Mythos Preview achieved a 40% success rate in single-target mode, compared with the 3.7% observed among participants in a specialized competition. In addition, its highest-ranked design showed stronger affinity than the winning design from that contest.
Affinity measures how strongly a protein binds to its target. High affinity may allow for lower doses and potentially reduce side effects and manufacturing costs. Even so, it is only one of the properties a molecule must have before it can become a treatment.
Claude Opus 4.8 also successfully designed binders against TNF alpha, a complex target associated with inflammation and the biological basis of medications such as Humira. Some of these designs bound to human, mouse, and macaque versions of the target, a useful feature when planning preclinical studies.
The result also offers an interesting lesson: the model that is generally more capable does not always win on every specific target. Mythos Preview failed against TNF alpha, while Opus 4.8 succeeded. In complex scientific tasks, performance depends on the structure of the problem—not just on a model’s overall ranking.
Claude also produced 15 confirmed binders with at least 20% beta-sheet structure. These structures are more difficult to design than alpha helices because they require amino acid chains to align correctly and can misfold or aggregate.
However, the system did not solve every challenge. Against MBP, it did not confirm any binder among 90 designs, although one showed a weak and reproducible signal. Against BBF-14, it produced three independent binders, but with more modest affinities.
An Autonomous Campaign with Specialized Models
To perform the design, Claude operated public tools for structure design, sequence generation, folding, and protein-protein complex evaluation. It also ran optimization cycles and filtered candidates according to criteria such as diversity, solubility, and predicted expression capacity.
Anthropic allowed the system to work with minimal human intervention. Researchers approved infrastructure access and monitored whether the sessions remained active, but they did not provide additional scientific instructions after the start.
In multi-target mode, the models worked for 48 hours and used up to 12,500 hours of computing on NVIDIA H100 GPUs. In single-target mode, Mythos Preview had 24-hour sessions and up to 2,500 H100 hours per target.
This is not the same as automating the complete development of a drug. A promising binder still has to undergo stability, toxicity, pharmacology, immunogenicity, manufacturing, and efficacy testing. Claude’s contribution occurs mainly at an early stage: finding experimental candidates more quickly.
Designing a binder is the beginning of a therapeutic process, not confirmation that a new drug exists.
Claude Analyzes NMR and LC-MS Data
The second experiment focused on a less flashy but very common task in synthetic chemistry: confirming the identity and purity of a molecule.
Chemists often use nuclear magnetic resonance spectroscopy, or NMR, to verify the structure of a compound. The spectrum contains peaks associated with hydrogen atoms. Their position helps determine the chemical environment of each atom, while their size indicates how many hydrogens they represent.
They also use liquid chromatography coupled with mass spectrometry, known as LC-MS. This technique separates the components of a sample, estimates how much of each one is present, and measures the molecular mass of the detected substances.
The instrument can generate the data in a few minutes, but interpreting the files is often the slow part. In addition, each manufacturer may use proprietary formats that require specialized software.
Anthropic gave Claude Opus 5 the raw files from a contracted laboratory and an instruction of just two sentences. Without using the manufacturer’s software or relying on a human operator, Claude processed the files in parallel:
- It completed the NMR analysis in 23 minutes.
- It processed the LC-MS data in 19 minutes.
- It identified 18 peaks in the NMR spectrum and calculated the number of hydrogens corresponding to each one.
- It estimated a purity of 96.4%, compared with the 96.33% calculated by the laboratory.
- It kept the hydrogen counts within a difference of 0.08 units compared with the professional analysis.
Claude also identified four broad peaks that probably corresponded to hydrogens bonded to nitrogen or oxygen. It then proposed a check using heavy water, a procedure that can cause those signals to diminish or disappear.
In an initial reading, the model stated that all four peaks had disappeared. However, when reviewing its own calculations, it detected that only two had disappeared and corrected its interpretation. The final result matched the laboratory’s independent conclusion.
Reverse Engineering a Proprietary File
The LC-MS analysis presented an additional challenge: the file format was not publicly documented. Claude deduced how the data were encoded and verified its interpretation by reproducing the totals recorded by the instrument itself across 2,664 measurements.
It then generated the results a chemist would typically need: the separation plot, mass and ultraviolet spectra, a purity table, the estimated molecular mass, and reusable code for reading similar files.
It also included warnings about the result’s limitations. For example, it pointed out that this type of instrument determines mass only to an approximate precision of one unit. The ability to report uncertainty is just as important as providing a final figure.
The laboratory took four days to deliver the complete report. Claude produced its report in approximately 25 minutes. This comparison does not mean the model should replace expert review, but it does point to a practical way to automate repetitive and verifiable tasks.
What Changes for Scientific Research
The two cases represent different extremes of laboratory work. In protein design, Claude coordinated specialized tools to create molecules that were later manufactured and tested. In analytical chemistry, it interpreted measurements from existing molecules and produced a reproducible report.
The main advantage is not that AI eliminates the need for scientists. It is that AI can reduce the time spent on searching, processing, and analysis, allowing experts to focus on deciding which experiments matter and how to interpret their results.
For people working in biotechnology, this could mean more candidates evaluated per campaign. For a small laboratory, it could reduce reliance on specialized software. And for researchers who are not programming experts, it opens the possibility of working with complex scientific data through natural-language instructions.
But the risk is real, too. Protein design is a dual-use capability: it can contribute to the development of therapies, but it could also facilitate dangerous biological research if it falls into the wrong hands. Anthropic keeps these functions blocked in its most advanced models with general access and is exploring controlled-access programs for scientists.
Experimental verification will remain essential. A model can propose a structure, process a spectrum, or find a pattern, but science requires controls, replication, independent review, and safety protocols. Speed is an advantage only when quality and traceability are maintained.
Claude points here toward a concrete direction for scientific AI: not merely discussing discoveries, but executing complete workflows, using technical tools, reviewing its own mistakes, and delivering results that can be checked. The next step will be determining how far that autonomy can extend without sacrificing the oversight that makes science trustworthy.
Original Source
https://www.anthropic.com/research/Claude-accelerates-protein-design
