P.K. SHARMA

Cyber security intelligence, AI governance, practitioner analysis

The same image, a different position: a vision test AI models missed

After motion adaptation, people and macaque inferior temporal cortex shifted their representation of a stationary object. The tested artificial vision networks located objects but did not reproduce the history-dependent bias.

By Parminder Kumar Sharma · · 10 min read

The test image did not move

A stationary object can appear displaced after a person has watched motion in one direction. In a study announced by York University and published in Current Biology, researchers used that motion aftereffect to ask a sharper question than whether humans or AI can locate an object: does recent visual history change the position code even when the test image's pixels are unchanged?

Human observers reported a shift opposite to the adapting motion. Neural population activity in macaque inferior temporal (IT) cortex also carried a shifted position signal. The artificial vision systems the team tested could represent object location but did not spontaneously show the corresponding adaptation-driven displacement. That is a difference in a specific computation, not a claim that all AI vision is incapable of seeing or that biological perception is always more accurate.

The hidden ambiguity in a normal location test

A vision system can appear to know where an object is because its internal features preserve enough of the image's pixel coordinates for a decoder to recover them. That alone says little about whether the system represents where an observer perceives the object. In an ordinary test, physical and perceived positions agree, so the two explanations make the same prediction. The researchers needed a case in which they diverged. Motion adaptation provides it: after watching movement, a stationary target can appear displaced although the target itself has not moved.

This is why the study does not simply compare humans with a model on a static image benchmark. It manipulates the history before the target. If the target image is identical after leftward and rightward adaptation, a position change cannot be attributed to a changed target pixel. The comparison asks whether the preceding motion altered a perceptual or neural position code. The biological and model measurements are still different kinds of evidence: a person's click is a report, IT activity is a neural signal from which researchers decode position, and a network feature is a mathematical representation. Agreement in direction is informative, but none is a direct readout of another system's subjective experience.

One stationary target, different histories

The researchers first established that human observers can report object location and that position is decodable from macaque IT activity. They then used drifting gratings as adaptors, followed by a briefly shown stationary object. Because the object's physical location stayed fixed, a shift in reported or decoded position reflects the preceding motion condition rather than a moved target. The research preprint describes an initial 30-second adaptation and three-second top-ups during the human task; its methods and figures should be read alongside the final journal version for exact replication.

This design connects behaviour and neural signals without pretending they are the same measurement. A person's response is a position judgement. The macaque result is a location decoded from a neural population. A network's output is a model representation. Their matching or diverging directions are informative; the magnitudes are not a simple league table across species and methods.

What was varied and what was measured.

  1. Part
    Adaptor
    Input or measurement
    Moving grating to the left or right
    Why it matters
    Changes recent motion history
  2. Part
    Test
    Input or measurement
    Stationary object at the same pixel position
    Why it matters
    Controls physical target location
  3. Part
    Human result
    Input or measurement
    Reported object position
    Why it matters
    Tests perceptual consequence
  4. Part
    Macaque result
    Input or measurement
    Position decoded from IT neurons
    Why it matters
    Tests a biological object representation
  5. Part
    Model result
    Input or measurement
    Position encoded by tested artificial networks
    Why it matters
    Tests whether their representations change after adaptation

The baseline first established that location was measurable

The author preprint describes a baseline localisation task with 35 human participants. Objects from eight identities, including animals and everyday items, appeared against naturalistic backgrounds at varied positions. Participants saw a subset of 40 images for 100 milliseconds and clicked the perceived centre on a blank frame. The paper reports high repeatability of the position estimates, with correlations of 0.96 horizontally and 0.99 vertically across observers. That makes a later systematic shift easier to distinguish from noisy clicking.

For macaque IT, researchers recorded from implanted electrode arrays during image viewing. They trained cross-validated linear decoders on spike counts 70 to 170 milliseconds after image onset to predict horizontal and vertical position. With more than 150 units, reported correlations reached about 0.60 for x and 0.68 for y. Some image networks also carried decodable position information: the preprint reports horizontal correlations of 0.84 for VGG-16 and 0.81 for ResNet-18. Those baseline results establish that the test is not about an inability of artificial networks to encode location at all. The more discriminating test is whether their location code changes with motion history.

Baseline numbers in the March author preprint; these are different measurement tasks, not a cross-species accuracy league.

  1. Measurement
    Human localisation
    Reported setup
    35 participants; 40 images; 100 ms presentation
    What it establishes
    Stable reports of target position
  2. Measurement
    Macaque IT decoding
    Reported setup
    70–170 ms neural window; more than 150 units
    What it establishes
    Object position can be read from IT activity
  3. Measurement
    Model decoding
    Reported setup
    VGG-16 and ResNet-18 position information
    What it establishes
    A model can encode where without reproducing the aftereffect

The adaptation test changed only the preceding motion

In the human aftereffect task, 22 participants watched a drifting grating moving left or right. The preprint specifies 30 seconds of initial adaptation and three-second top-ups, followed by a stationary test image shown for 100 milliseconds. Participants then marked the apparent centre. In the macaque recording experiment, the motion adaptor lasted three seconds before the stationary test; the neural population was read after the target appeared. These are matched in logic, not identical in timing or response method.

The control that carries the argument is the unchanged test target. The same object at the same pixel position can follow different motion directions. The researchers compared position reports and pre-trained neural decoders across those histories. A decoder trained before adaptation was applied to responses after adaptation without retraining. That prevents a post-hoc fitting step from creating the observed shift by learning a new coordinate system for each condition.

The direction matters more than a leaderboard

In the preprint, rightward adaptation made the human target appear on average 0.20° leftward; leftward adaptation produced an average 0.13° rightward bias. Those are angular visual-field shifts, not screen pixels. The paper reports statistical tests for the human effects and a comparable direction-opponent shift in macaque IT population decoding. The IT leftward-adaptor condition was weaker and not significant in the reported preprint analysis, so the result should not be reduced to perfect symmetry.

The authors tested feedforward image networks and systems with recurrence or video processing. These could encode object position, yet their representations generally failed to acquire the adaptation-induced position shifts seen in the biological measurements. When the researchers imposed transformations based on IT data, models could show the bias. That is an intervention demonstrating one computational route, not evidence the unmodified networks already learned it.

Qualitative comparison. Human angles are from the author preprint; cross-system magnitudes are not directly comparable.

  1. System
    Human observers
    Stationary location available?
    Yes
    History-dependent shift?
    Yes; opposite to prior motion
  2. System
    Macaque IT population
    Stationary location available?
    Yes, by neural decoding
    History-dependent shift?
    Direction-opponent shift, with condition-specific strength
  3. System
    Tested artificial networks
    Stationary location available?
    Yes
    History-dependent shift?
    Generally absent without an imposed IT-like transformation

The measured shifts and their asymmetry

The author preprint reports a mean human horizontal shift of 0.20° left after rightward adaptation and 0.13° right after leftward adaptation. Both were statistically significant in its analysis. It reports no corresponding systematic vertical shift, consistent with a horizontally moving adaptor. The unit is a degree of visual angle, which depends on the viewing geometry; treating these values as screen pixels would be wrong.

The macaque decoder moved in the same directional pattern, but the conditions were not equally strong. After rightward adaptation, its mean x estimate shifted 0.71° left and was significant. After leftward adaptation, the mean was 0.14° right, but the paper reports p = 0.33, so that individual condition was not statistically established. The preprint also found changes in IT population geometry after adaptation, using centred kernel alignment, and retained substantial position-decoding reliability. This argues against a simple story in which the neural signal just became random or weak. It still does not prove IT alone caused the human perceptual effect: the human behaviour and macaque recordings came from different subjects and protocols.

Direction-specific preprint results. These are not directly comparable effect sizes across measurement systems.

  1. Measure
    Human reported x-position
    After rightward motion
    0.20° left; p < 0.001
    After leftward motion
    0.13° right; p = 0.0036
  2. Measure
    Macaque IT decoded x-position
    After rightward motion
    0.71° left; p < 0.001
    After leftward motion
    0.14° right; p = 0.33, not significant
  3. Measure
    Vertical position
    After rightward motion
    No systematic human or IT shift reported
    After leftward motion
    No systematic human or IT shift reported

A missing computation, not a universal verdict on AI

The study points to adaptation-driven reshaping of object representations as a candidate ingredient of biological vision. It does not prove that IT alone creates the illusion, and a correlated neural shift is not a complete causal account of perception. It also does not show that no future model could learn the effect. The comparison covers the architectures and training regimes the authors actually tested.

For model evaluation, the useful lesson is to test sequences as well as isolated images. An object-location benchmark that resets state for every frame can hide whether a system uses recent motion to update its representation. That may matter in video interpretation, robotics and human-like perceptual modelling, although this paper does not quantify a deployment failure in those systems. A fair follow-up would expose models to the same adaptor and target sequence, measure the direction and uncertainty of any shift, and compare that with human reports rather than simply asking whether the final image is classified correctly.

Which model tests failed, and what the successful intervention means

The researchers did more than pass one final image through one feedforward network. They examined standard object-recognition models, added unit-level suppression fitted to recorded IT adaptation dynamics, and tested systems with temporal processing. Suppression reduced activity and changed feature reliability, but did not create the direction-opponent position shift. In the preprint's model comparison, simulated-adaptation deltas for AlexNet, VGG-16, ResNet-18, ViT-L32 and SimCLR ResNet-50 stayed near zero, while the human and IT comparisons were much larger. The authors also tested the video architecture SlowFast and a recurrent network, ConvRNN; temporal input or feedback alone did not reproduce the biological pattern in those tests.

A separate intervention did create the effect: the authors estimated transformations from IT's own before-and-after adaptation activity and applied those transformations to model features. Position decoders on transformed VGG-16, ResNet-18 and SimCLR features then showed analogous directional biases. This is a useful sufficiency result: existing feature spaces can be made to express the shift when an empirically derived adaptation transform is imposed. It is not evidence that those networks learned such a transform on their own, nor a proposal that hallucinated object displacement is always desirable in a deployed vision system.

Three distinct model questions in the author preprint.

  1. Test
    Unmodified image networks
    Finding
    Location decodable; history-linked bias absent
    Interpretation
    Knowing pixel position is not the same as adapting perceived position
  2. Test
    Suppression, video and recurrence
    Finding
    No systematic direction-opponent shift in tested implementations
    Interpretation
    Generic temporal processing or lower activity was insufficient
  3. Test
    IT-derived transform applied
    Finding
    Shift appeared in transformed feature spaces
    Interpretation
    A biologically measured transform can induce the effect

A better benchmark would preserve the sequence

The overlooked design lesson is to preserve the adaptor, delay and target as a sequence when evaluating models. A benchmark that resets the model before every target can only test its response to the final pixels. A stronger follow-up would present the same leftward and rightward histories, use a frozen position decoder, measure uncertainty and vertical controls, and compare the direction of the change with the biological reports. It would also test whether a model's shift improves a task that matters outside this illusion. Until such work exists, the result is a precise gap in the tested systems, not a universal failure of artificial vision.

Sources

  1. PrimaryYork University release on motion adaptation and vision networksYork Universityaccessed 2026-10-02
  2. PrimaryResearch preprint and methodsYork University researchersaccessed 2026-10-02
  3. PrimaryPublished study DOICurrent Biologyaccessed 2026-10-02

Share this briefing

Know someone who owns this problem? Send it to them.

Related briefings

The briefing, in your inbox

Practitioner analysis of cyber and AI security news. No vendor noise.

How often

Every new briefing in one email, at 7am, or at 7am, 12:30pm and 6pm. Nothing is sent when nothing is new. Unsubscribe any time.