Artificial intelligence is no longer evaluated only by its ability to write text, code, or answer questions. Anthropic is testing something much more sensitive: whether models can help identify people, locate targets, and develop software for drones in simulated military environments. How close are we to making tasks once reserved for analysts and specialists accessible to groups with fewer resources?
A new way to measure AI’s military risk
Anthropic’s Frontier Red Team designed evaluations to analyze two areas: target intelligence and conventional weapons engineering. The work is part of a broader concern about the misuse of advanced models in real-world conflicts.
So far, much of the conversation about AI risks has focused on cybersecurity and biology. However, many conflicts depend on more conventional activities: locating people, interpreting scattered information, guiding drones, and overcoming jamming systems.
These tasks are often described through a sequence known as the kill chain: find, fix, track, target, engage, and assess. If AI improves any of these steps, it can reduce the time, cost, and expertise needed to complete an operation.
The risk does not depend solely on whether a model can act on its own. It also matters whether it makes knowledge that was once scarce and expensive more accessible.
Models capable of linking digital identities
One test examined whether models could link accounts belonging to the same person across platforms such as WhatsApp, Telegram, Instagram, and Facebook. To do this, Anthropic used synthetic content that simulated user activity in two fictional scenarios.
The goal was to determine whether a model could connect different digital identities and classify people according to their potential relevance to an investigation. The metric used was F1, which combines precision and recall to evaluate how well the correct cases are identified without multiplying false positives.
Mythos Preview delivered the best performance in account linking and individual classification. Kimi K3, an open-weights model, came close to frontier models in the simple and medium scenarios, although its performance declined as noise and identity-protection measures increased.
Anthropic cautions that these data do not fully represent real human behavior. Synthetic content may contain artificial phrases and unnatural patterns. For that reason, the results are useful for comparing capabilities across models, but they should not be interpreted as a definitive measure of performance in real operations.
One detail is especially striking: a medium-sized sample contained nearly 37,000 words. A human analyst would take about 2.5 hours to read it and probably much longer to cross-reference all the information. Mythos Preview produced a complete analysis in approximately 11 minutes.
Geolocation is already approaching expert level
Anthropic also tested whether models could estimate where a photograph was taken using only its visual elements, without metadata, reverse searches, or external tools.
Across 6,000 images, Mythos Preview reached a median error of 37 kilometers, while Mythos 5 reached 47.2 kilometers. In addition, both placed about 23 percent of the images within one kilometer of the correct location.
As a reference point, Anthropic compared these results with expert GeoGuessr players. The best human competitors included in the study recorded median errors of between 151 and 174 kilometers. The comparison is not perfect, because GeoGuessr uses Street View images and allows players to explore the scene, while Anthropic’s test used more varied static photographs.
Opus 5 recorded a median error of 181 kilometers, similar to the level of advanced human players. Sonnet 5 reached 384 kilometers. Kimi K3 recorded 385 kilometers, although it placed a higher percentage of images within one kilometer than Sonnet 5.
The conclusion is significant: the most advanced models appear to combine world knowledge and vision effectively enough to find geographic clues that a person might overlook. A sign, an architectural style, a landscape, or a local business can become a location signal.
Inferring a location from posts
In another evaluation, the models tried to estimate the city where anonymous users lived based on their posts. The dataset contained geolocated messages from nearly 1,700 people, but usernames, mentions, and reposts were removed.
At least one of the models located 135 users within one kilometer, or about 8 percent of the sample. In 70 percent of those cases, the people had revealed their location through direct references, such as universities, well-known places, streets, or ZIP codes.
But in another 13 percent, the model used indirect signals: dialect, local expressions, television shows, radio stations, transit lines, events, and sports teams. In other words, you do not have to post your address to leave enough clues about where you live.
Mythos Preview achieved a median error of 20.1 kilometers, followed by Mythos 5 at 20.9 and Opus 5 at 21.7. Sonnet 5 recorded 31.3 kilometers and Kimi K3, 26.4 kilometers.
Anthropic applied safeguards to prevent the models from simply remembering the dataset or searching for exact texts online. Even so, the company acknowledges that a system with broader search access and several rounds of research could achieve more accurate results.
Drone testing in simulated environments
The second part of the study analyzed the models’ ability to write and improve drone-control software in simulators. The tests included navigation, visual tracking of a target, delivery of a simulated payload, and operation with degraded or manipulated positioning signals.
The models received a written report, a programming environment, and data from each attempt, such as trajectories, sensor readings, and camera snapshots. They could then modify their software and test it again within a limited number of launches.
The results showed a clear difference between the models. In the simplest scenarios, Opus 5 delivered the best performance, followed by Mythos Preview and Mythos 5. Kimi K3 outperformed Sonnet 5 in some tests, especially in simulated payload delivery, but fell behind the frontier models in navigation and tracking.
In the more difficult scenarios, with poorly visible targets, obstacles, unpredictable movement, or positioning interference, every model’s performance dropped sharply. None consistently solved all the conditions being evaluated.
Anthropic observed that Opus 5 tended to make small, progressive changes to its code, while other models rewrote larger parts of the system. It also used internal physics models and more sophisticated estimation and control techniques. This behavior allowed it to make better use of the limited number of tests available.
Why simulation does not tell the whole story
The results do not mean that a model can independently build a reliable military system. The tests did not include real hardware, manufacturing, maintenance, full atmospheric conditions, or all the problems that arise on a battlefield.
In fact, the authors themselves describe these evaluations as a floor, not a ceiling. A real-world group could give the model access to more information, better cameras, more tests, human assistance, and specialized tools. All of that could improve performance.
There are also material limits. Having a model capable of writing software is not the same as having the sensors, components, infrastructure, or personnel needed to deploy a functional system. The important question is how much AI can reduce those barriers over time.
The problem does not belong to a single lab
Anthropic’s central conclusion is that both proprietary and open-weights models already show concerning capabilities in intelligence and simulated military systems. Open-weights models lag behind the best commercial models, but not enough to consider them irrelevant from a security standpoint.
This raises difficult decisions for developers and governments. What tasks should be evaluated before a model is released? How can privacy be protected when an apparently innocent post makes it possible to infer someone’s location? What controls should apply when a system can help develop navigation or control software?
The discussion is not limited to drones, either. Anthropic anticipates that similar advances could influence space systems, underwater operations, and other areas where interpreting the environment and reacting quickly is valuable.
AI applied to national security is not a distant possibility. It is already being tested, used, and measured. The challenge is to understand precisely which capabilities are maturing, limit harmful uses, and prevent the falling cost of knowledge from expanding the number of actors capable of causing harm.
Original source
https://www.anthropic.com/research/intelligence-targeting-conventional-weapons-capabilities
