Anthropic's models can geolocate a photograph and write simulated drone guidance. The missing number is human uplift
Anthropic tested frontier and open-weight models on intelligence triage, photo and text geolocation, terminal guidance, payload delivery and navigation under GPS interference. The results show real dual-use capability, but the report does not measure how much more effective a human operator becomes.
By Parminder Kumar Sharma · · 9 min read

The result is wider than one alarming benchmark
Anthropic has published one of the clearest public attempts to measure how modern AI changes military and intelligence work. The report deserves more care than either of the easy headlines. It is neither proof that an autonomous weapon can now be built from a prompt nor reassurance that the work is merely a simulation.
The study asks a more useful question: which pieces of expertise that once required scarce analysts or engineers can a model now perform alone, under controlled conditions?
The evaluations cover intelligence triage, geolocation from photographs, geolocation from social posts, terminal guidance, payload delivery and navigation when GPS is denied or spoofed. Anthropic tested its Claude models, including Opus 5 and the newer Mythos systems, alongside open-weight models including Kimi K3 and GLM 5.2. The work is dual-use in the strict sense. The same vision, search, navigation and control skills can support rescue, inspection and defensive analysis. They can also reduce the expertise needed to identify or reach a person or vehicle.
The central security finding is capability diffusion. A state already employing experienced analysts gains speed and scale. A smaller group may gain access to methods it could not previously staff. The report shows signals for both effects. It does not yet measure either effect directly.
Where the tested capabilities sit in a targeting chain
US Army doctrine describes dynamic targeting as find, fix, track, target, engage and assess. The stages can overlap, but the names prevent one capability from being mistaken for the whole chain.
Anthropic's image and text work mostly addresses find and fix: identifying relevant material and narrowing location. Its simulator work addresses pieces of track and engage: perceiving a designated vehicle, estimating its movement, generating flight-control code and adjusting navigation. The study did not test the legal, command and intelligence processes that decide whether a target is valid, whether force is proportionate or whether a human is authorized to act.
That boundary is essential. A model can be technically capable of following an object while the overall system remains operationally unreliable, legally prohibited or unsafe to deploy. The model evaluation is evidence about one component, not certification of a weapon.
The 37 kilometre number needs its denominator and comparison
For image geolocation, the models received 6,000 permissively licensed Flickr photographs with location metadata hidden. They had no reverse-image search, metadata or external tools. Mythos Preview produced a median error of 37.0 kilometres and placed 23.7% of images within one kilometre. Mythos 5 recorded 47.2 kilometres and 23.1% within one kilometre. Opus 5 recorded 181 kilometres and 18% within one kilometre.
Anthropic describes the Mythos result as approaching superhuman performance, using elite GeoGuessr players as a proxy. That comparison is suggestive and imperfect. The human data came from a different task using navigable Street View scenes, while the models saw static social photographs. Anthropic says it has no human baseline on the same 6,000-image dataset. The report should therefore be read as evidence of strong machine geolocation, not a clean model-versus-human contest.
The one-kilometre rate matters alongside the median. A 37 kilometre median tells us half the guesses are closer and half are farther. It does not say every image can be located usefully. Fewer than one in four Mythos guesses landed within one kilometre. For intelligence triage across thousands of images, that can be significant. For an action requiring confidence about one specific location, the error distribution and verification process matter more than the headline median.
Text created another route. Using one week of anonymised posts from 1,697 users, the strongest systems reached median errors near 20 kilometres when allowed a sandboxed search tool. At least one model placed 135 users within one kilometre. In most successful cases, users had exposed affiliations, venues or explicit locations. In some, dialect, transit, media markets, local events or sports discussion was enough to narrow the area.
Selected results reported by Anthropic. The GeoGuessr figure is a proxy from a different task, not a matched human trial.
| System or proxy | Median image error | Within 1 km | Important condition |
|---|---|---|---|
| Mythos Preview | 37.0 km | 23.7% | Static Flickr image, no tools or metadata |
| Mythos 5 | 47.2 km | 23.1% | Static Flickr image, no tools or metadata |
| Opus 5 | 181 km | 18.0% | Static Flickr image, no tools or metadata |
| Elite GeoGuessr proxy | 151 km | Not comparable | Different dataset with movable Street View |
The drone tests show a steep capability gradient
Anthropic placed models in a physics simulator and asked them to write guidance, navigation and control software. They worked from a written brief, sensor data and a limited set of trials. The exercise began with an airborne drone and a designated vehicle. The models had to track the object and translate estimates into frequent control commands.
Performance fell sharply as conditions became harder. Against a parked, high-contrast vehicle, Opus 5 succeeded in 80% of launches and Mythos Preview in 70%. At road speed, those figures fell to 47% and 20%. With low contrast, only Opus recorded any strikes, at 8%. Camouflage, evasion and decoys were not consistently solved.
Across nine terminal-guidance settings, Opus hit the simulated target in 20% of 540 launches, Mythos Preview in 13%, Mythos 5 in 10%, Kimi K3 in 1.6% and Sonnet 5 in 0.7%. Those totals are operationally poor and technically consequential. They show that a model can generate a functioning control stack, iterate using telemetry and solve easier variants without a specialist writing the code.
The navigation tests make the same point. Frontier models often detected sudden GPS denial and used inertial data to continue, though accuracy deteriorated. Slow GPS drift defeated all models in the hardest setting. The distinction resembles security testing: a system may handle a conspicuous failure and remain vulnerable to a subtle, plausible signal that changes gradually.
Selected simulated terminal-guidance results. These are model-alone simulator trials, not field trials or an assessment of a complete weapon system.
| Condition | Opus 5 | Mythos Preview | What changed |
|---|---|---|---|
| Parked, high contrast vehicle | 80% | 70% | Simple visual separation and no target motion |
| Vehicle at road speed | 47% | 20% | Tracking and prediction become necessary |
| Added roadside clutter | 30% | Not separately highlighted | Background features interfere with tracking |
| Low contrast | 8% | 0% | Visual signal approaches the background |
The missing measurement is uplift
Anthropic is clear that the study does not directly measure uplift. That is the most important limitation.
A model-alone score answers whether the model can complete a bounded task in a sandbox. Threat modelling needs a different comparison: how does a real actor perform without the model, with the model, with ordinary open-source tools, and with a skilled human partner? The change between those conditions is the enabled harm.
For example, a weak actor might fail to build a robust guidance loop alone. A model may produce a fragile version that works only against a stationary object. A knowledgeable operator can then inspect the telemetry, correct the control law and choose a better detector. The combined team could be far better than the model-alone benchmark. The reverse can also happen: the model's confident mistakes may waste scarce hardware and make the actor worse.
The report calls its model-alone results a floor because human collaboration, web access, libraries and real-world testing could repair errors. That is reasonable as a hypothesis. It remains an empirical question, and the answer will differ by task and actor. Security policy should avoid turning a floor into a forecast without measuring the stairs between them.
Controls that follow from the evidence
Controls tied to the observed capability rather than to a general fear of AI.
| Risk | Control | Evidence to retain |
|---|---|---|
| Bulk geolocation of people | Detect repeated dossier building and restrict identity-resolution searches | Queries, target count, location confidence and escalation decision |
| Weapons engineering assistance | Classify high-risk control and delivery workflows across prompts, code and tool use | Classifier result, model response, tools called and reviewer outcome |
| Open-weight diffusion | Control access to sensors, test ranges, flight hardware and deployment infrastructure | Hardware identity, firmware, operator and mission authorization |
| False confidence | Require independent sensor confirmation and calibrated uncertainty | Raw sensor input, model estimate, human check and final coordinate source |
| Capability growth | Repeat matched uplift evaluations with real domain experts | Baseline, assisted result, time, quality, failures and intervention level |
Anthropic says it has introduced classifiers for requests related to weapons development and acknowledges that dual-use engineering makes those classifiers imperfect. That is a necessary layer, not a complete boundary. The model is only one component. Search, maps, sensors, code execution, flight hardware and operator authority each offer a separate place to apply control.
Open-weight models make provider enforcement less comprehensive, but they do not make governance pointless. A model that runs locally still needs data, compute, hardware and access to act. The control strategy should move from one provider refusal to several observable boundaries.
The P.K. view
The report is valuable because it measures components that safety discussions often leave abstract. It also demonstrates how easily a benchmark can be overstated. Thirty-seven kilometres is impressive geolocation at scale and inadequate precision for many individual actions. An 80% result on the easiest simulation is a capability signal and not battlefield reliability.
The near-term concern is not a model independently deciding to conduct a military operation. It is the compression of specialised work. An analyst can triage more material. A small engineering team can reach a functioning prototype faster. A harmful actor can ask better questions before a human specialist becomes involved.
The next evaluation should put domain practitioners into the experiment and publish the uplift. That result would tell defenders and policymakers whether AI saves minutes for experts, weeks for novices, or creates a new class of actor who could not perform the task before. Until then, the honest conclusion is serious but bounded: present models can perform meaningful pieces of intelligence and weapons engineering, and we still do not know how much more capable they make the people using them.
Sources
- PrimaryMeasuring AI capabilities in intelligence targeting and conventional weaponsAnthropicaccessed 2026-09-14
- PrimaryFM 3-60: Army TargetingUS Department of the Armyaccessed 2026-09-14


