P.K. SHARMA

Cyber security intelligence, AI governance, practitioner analysis

Anthropic's models can geolocate a photograph and write simulated drone guidance. The missing number is human uplift

Anthropic tested frontier and open-weight models on intelligence triage, photo and text geolocation, terminal guidance, payload delivery and navigation under GPS interference. The results show real dual-use capability, but the report does not measure how much more effective a human operator becomes.

By Parminder Kumar Sharma · · 9 min read

An intelligence analysis room linking public photographs to broad map probability rings, human review and a simulated drone.

The result is wider than one alarming benchmark

Anthropic has published one of the clearest public attempts to measure how modern AI changes military and intelligence work. The report deserves more care than either of the easy headlines. It is neither proof that an autonomous weapon can now be built from a prompt nor reassurance that the work is merely a simulation.

The study asks a more useful question: which pieces of expertise that once required scarce analysts or engineers can a model now perform alone, under controlled conditions?

The evaluations cover intelligence triage, geolocation from photographs, geolocation from social posts, terminal guidance, payload delivery and navigation when GPS is denied or spoofed. Anthropic tested its Claude models, including Opus 5 and the newer Mythos systems, alongside open-weight models including Kimi K3 and GLM 5.2. The work is dual-use in the strict sense. The same vision, search, navigation and control skills can support rescue, inspection and defensive analysis. They can also reduce the expertise needed to identify or reach a person or vehicle.

The central security finding is capability diffusion. A state already employing experienced analysts gains speed and scale. A smaller group may gain access to methods it could not previously staff. The report shows signals for both effects. It does not yet measure either effect directly.

Where the tested capabilities sit in a targeting chain

US Army doctrine describes dynamic targeting as find, fix, track, target, engage and assess. The stages can overlap, but the names prevent one capability from being mistaken for the whole chain.

Anthropic's image and text work mostly addresses find and fix: identifying relevant material and narrowing location. Its simulator work addresses pieces of track and engage: perceiving a designated vehicle, estimating its movement, generating flight-control code and adjusting navigation. The study did not test the legal, command and intelligence processes that decide whether a target is valid, whether force is proportionate or whether a human is authorized to act.

That boundary is essential. A model can be technically capable of following an object while the overall system remains operationally unreliable, legally prohibited or unsafe to deploy. The model evaluation is evidence about one component, not certification of a weapon.

Six-stage flow from find and fix through track, target, engage and assess, distinguishing tested AI capabilities from human authority.
The evaluation measures components. Validation, legal review, authorization and assessment belong to the wider system.

The 37 kilometre number needs its denominator and comparison

For image geolocation, the models received 6,000 permissively licensed Flickr photographs with location metadata hidden. They had no reverse-image search, metadata or external tools. Mythos Preview produced a median error of 37.0 kilometres and placed 23.7% of images within one kilometre. Mythos 5 recorded 47.2 kilometres and 23.1% within one kilometre. Opus 5 recorded 181 kilometres and 18% within one kilometre.

Anthropic describes the Mythos result as approaching superhuman performance, using elite GeoGuessr players as a proxy. That comparison is suggestive and imperfect. The human data came from a different task using navigable Street View scenes, while the models saw static social photographs. Anthropic says it has no human baseline on the same 6,000-image dataset. The report should therefore be read as evidence of strong machine geolocation, not a clean model-versus-human contest.

The one-kilometre rate matters alongside the median. A 37 kilometre median tells us half the guesses are closer and half are farther. It does not say every image can be located usefully. Fewer than one in four Mythos guesses landed within one kilometre. For intelligence triage across thousands of images, that can be significant. For an action requiring confidence about one specific location, the error distribution and verification process matter more than the headline median.

Text created another route. Using one week of anonymised posts from 1,697 users, the strongest systems reached median errors near 20 kilometres when allowed a sandboxed search tool. At least one model placed 135 users within one kilometre. In most successful cases, users had exposed affiliations, venues or explicit locations. In some, dialect, transit, media markets, local events or sports discussion was enough to narrow the area.

Selected results reported by Anthropic. The GeoGuessr figure is a proxy from a different task, not a matched human trial.

System or proxyMedian image errorWithin 1 kmImportant condition
Mythos Preview37.0 km23.7%Static Flickr image, no tools or metadata
Mythos 547.2 km23.1%Static Flickr image, no tools or metadata
Opus 5181 km18.0%Static Flickr image, no tools or metadata
Elite GeoGuessr proxy151 kmNot comparableDifferent dataset with movable Street View

The drone tests show a steep capability gradient

Anthropic placed models in a physics simulator and asked them to write guidance, navigation and control software. They worked from a written brief, sensor data and a limited set of trials. The exercise began with an airborne drone and a designated vehicle. The models had to track the object and translate estimates into frequent control commands.

Performance fell sharply as conditions became harder. Against a parked, high-contrast vehicle, Opus 5 succeeded in 80% of launches and Mythos Preview in 70%. At road speed, those figures fell to 47% and 20%. With low contrast, only Opus recorded any strikes, at 8%. Camouflage, evasion and decoys were not consistently solved.

Across nine terminal-guidance settings, Opus hit the simulated target in 20% of 540 launches, Mythos Preview in 13%, Mythos 5 in 10%, Kimi K3 in 1.6% and Sonnet 5 in 0.7%. Those totals are operationally poor and technically consequential. They show that a model can generate a functioning control stack, iterate using telemetry and solve easier variants without a specialist writing the code.

The navigation tests make the same point. Frontier models often detected sudden GPS denial and used inertial data to continue, though accuracy deteriorated. Slow GPS drift defeated all models in the hardest setting. The distinction resembles security testing: a system may handle a conspicuous failure and remain vulnerable to a subtle, plausible signal that changes gradually.

Selected simulated terminal-guidance results. These are model-alone simulator trials, not field trials or an assessment of a complete weapon system.

ConditionOpus 5Mythos PreviewWhat changed
Parked, high contrast vehicle80%70%Simple visual separation and no target motion
Vehicle at road speed47%20%Tracking and prediction become necessary
Added roadside clutter30%Not separately highlightedBackground features interfere with tracking
Low contrast8%0%Visual signal approaches the background

The missing measurement is uplift

Anthropic is clear that the study does not directly measure uplift. That is the most important limitation.

A model-alone score answers whether the model can complete a bounded task in a sandbox. Threat modelling needs a different comparison: how does a real actor perform without the model, with the model, with ordinary open-source tools, and with a skilled human partner? The change between those conditions is the enabled harm.

For example, a weak actor might fail to build a robust guidance loop alone. A model may produce a fragile version that works only against a stationary object. A knowledgeable operator can then inspect the telemetry, correct the control law and choose a better detector. The combined team could be far better than the model-alone benchmark. The reverse can also happen: the model's confident mistakes may waste scarce hardware and make the actor worse.

The report calls its model-alone results a floor because human collaboration, web access, libraries and real-world testing could repair errors. That is reasonable as a hypothesis. It remains an empirical question, and the answer will differ by task and actor. Security policy should avoid turning a floor into a forecast without measuring the stairs between them.

Controls that follow from the evidence

Controls tied to the observed capability rather than to a general fear of AI.

RiskControlEvidence to retain
Bulk geolocation of peopleDetect repeated dossier building and restrict identity-resolution searchesQueries, target count, location confidence and escalation decision
Weapons engineering assistanceClassify high-risk control and delivery workflows across prompts, code and tool useClassifier result, model response, tools called and reviewer outcome
Open-weight diffusionControl access to sensors, test ranges, flight hardware and deployment infrastructureHardware identity, firmware, operator and mission authorization
False confidenceRequire independent sensor confirmation and calibrated uncertaintyRaw sensor input, model estimate, human check and final coordinate source
Capability growthRepeat matched uplift evaluations with real domain expertsBaseline, assisted result, time, quality, failures and intervention level

Anthropic says it has introduced classifiers for requests related to weapons development and acknowledges that dual-use engineering makes those classifiers imperfect. That is a necessary layer, not a complete boundary. The model is only one component. Search, maps, sensors, code execution, flight hardware and operator authority each offer a separate place to apply control.

Open-weight models make provider enforcement less comprehensive, but they do not make governance pointless. A model that runs locally still needs data, compute, hardware and access to act. The control strategy should move from one provider refusal to several observable boundaries.

The P.K. view

The report is valuable because it measures components that safety discussions often leave abstract. It also demonstrates how easily a benchmark can be overstated. Thirty-seven kilometres is impressive geolocation at scale and inadequate precision for many individual actions. An 80% result on the easiest simulation is a capability signal and not battlefield reliability.

The near-term concern is not a model independently deciding to conduct a military operation. It is the compression of specialised work. An analyst can triage more material. A small engineering team can reach a functioning prototype faster. A harmful actor can ask better questions before a human specialist becomes involved.

The next evaluation should put domain practitioners into the experiment and publish the uplift. That result would tell defenders and policymakers whether AI saves minutes for experts, weeks for novices, or creates a new class of actor who could not perform the task before. Until then, the honest conclusion is serious but bounded: present models can perform meaningful pieces of intelligence and weapons engineering, and we still do not know how much more capable they make the people using them.

Sources

  1. PrimaryMeasuring AI capabilities in intelligence targeting and conventional weaponsAnthropicaccessed 2026-09-14
  2. PrimaryFM 3-60: Army TargetingUS Department of the Armyaccessed 2026-09-14

Share this briefing

Know someone who owns this problem? Send it to them.

Related briefings

The briefing, in your inbox

Practitioner analysis of cyber and AI security news. No vendor noise.

One email per briefing. Unsubscribe any time.