The same image, a different position: a vision test AI models missed
After motion adaptation, people and macaque inferior temporal cortex shifted their representation of a stationary object. The tested artificial vision networks located objects but did not reproduce the history-dependent bias.
By Parminder Kumar Sharma · · 10 min read
The test image did not move
A stationary object can appear displaced after a person has watched motion in one direction. In a study announced by York University and published in Current Biology, researchers used that motion aftereffect to ask a sharper question than whether humans or AI can locate an object: does recent visual history change the position code even when the test image's pixels are unchanged?
Human observers reported a shift opposite to the adapting motion. Neural population activity in macaque inferior temporal (IT) cortex also carried a shifted position signal. The artificial vision systems the team tested could represent object location but did not spontaneously show the corresponding adaptation-driven displacement. That is a difference in a specific computation, not a claim that all AI vision is incapable of seeing or that biological perception is always more accurate.
The hidden ambiguity in a normal location test
A vision system can appear to know where an object is because its internal features preserve enough of the image's pixel coordinates for a decoder to recover them. That alone says little about whether the system represents where an observer perceives the object. In an ordinary test, physical and perceived positions agree, so the two explanations make the same prediction. The researchers needed a case in which they diverged. Motion adaptation provides it: after watching movement, a stationary target can appear displaced although the target itself has not moved.
This is why the study does not simply compare humans with a model on a static image benchmark. It manipulates the history before the target. If the target image is identical after leftward and rightward adaptation, a position change cannot be attributed to a changed target pixel. The comparison asks whether the preceding motion altered a perceptual or neural position code. The biological and model measurements are still different kinds of evidence: a person's click is a report, IT activity is a neural signal from which researchers decode position, and a network feature is a mathematical representation. Agreement in direction is informative, but none is a direct readout of another system's subjective experience.
One stationary target, different histories
The researchers first established that human observers can report object location and that position is decodable from macaque IT activity. They then used drifting gratings as adaptors, followed by a briefly shown stationary object. Because the object's physical location stayed fixed, a shift in reported or decoded position reflects the preceding motion condition rather than a moved target. The research preprint describes an initial 30-second adaptation and three-second top-ups during the human task; its methods and figures should be read alongside the final journal version for exact replication.
This design connects behaviour and neural signals without pretending they are the same measurement. A person's response is a position judgement. The macaque result is a location decoded from a neural population. A network's output is a model representation. Their matching or diverging directions are informative; the magnitudes are not a simple league table across species and methods.
What was varied and what was measured.
- Part
- Adaptor
- Input or measurement
- Moving grating to the left or right
- Why it matters
- Changes recent motion history
- Part
- Test
- Input or measurement
- Stationary object at the same pixel position
- Why it matters
- Controls physical target location
- Part
- Human result
- Input or measurement
- Reported object position
- Why it matters
- Tests perceptual consequence
- Part
- Macaque result
- Input or measurement
- Position decoded from IT neurons
- Why it matters
- Tests a biological object representation
- Part
- Model result
- Input or measurement
- Position encoded by tested artificial networks
- Why it matters
- Tests whether their representations change after adaptation
| Part | Input or measurement | Why it matters |
|---|---|---|
| Adaptor | Moving grating to the left or right | Changes recent motion history |
| Test | Stationary object at the same pixel position | Controls physical target location |
| Human result | Reported object position | Tests perceptual consequence |
| Macaque result | Position decoded from IT neurons | Tests a biological object representation |
| Model result | Position encoded by tested artificial networks | Tests whether their representations change after adaptation |
The baseline first established that location was measurable
The author preprint describes a baseline localisation task with 35 human participants. Objects from eight identities, including animals and everyday items, appeared against naturalistic backgrounds at varied positions. Participants saw a subset of 40 images for 100 milliseconds and clicked the perceived centre on a blank frame. The paper reports high repeatability of the position estimates, with correlations of 0.96 horizontally and 0.99 vertically across observers. That makes a later systematic shift easier to distinguish from noisy clicking.
For macaque IT, researchers recorded from implanted electrode arrays during image viewing. They trained cross-validated linear decoders on spike counts 70 to 170 milliseconds after image onset to predict horizontal and vertical position. With more than 150 units, reported correlations reached about 0.60 for x and 0.68 for y. Some image networks also carried decodable position information: the preprint reports horizontal correlations of 0.84 for VGG-16 and 0.81 for ResNet-18. Those baseline results establish that the test is not about an inability of artificial networks to encode location at all. The more discriminating test is whether their location code changes with motion history.
Baseline numbers in the March author preprint; these are different measurement tasks, not a cross-species accuracy league.
- Measurement
- Human localisation
- Reported setup
- 35 participants; 40 images; 100 ms presentation
- What it establishes
- Stable reports of target position
- Measurement
- Macaque IT decoding
- Reported setup
- 70–170 ms neural window; more than 150 units
- What it establishes
- Object position can be read from IT activity
- Measurement
- Model decoding
- Reported setup
- VGG-16 and ResNet-18 position information
- What it establishes
- A model can encode where without reproducing the aftereffect
| Measurement | Reported setup | What it establishes |
|---|---|---|
| Human localisation | 35 participants; 40 images; 100 ms presentation | Stable reports of target position |
| Macaque IT decoding | 70–170 ms neural window; more than 150 units | Object position can be read from IT activity |
| Model decoding | VGG-16 and ResNet-18 position information | A model can encode where without reproducing the aftereffect |
The adaptation test changed only the preceding motion
In the human aftereffect task, 22 participants watched a drifting grating moving left or right. The preprint specifies 30 seconds of initial adaptation and three-second top-ups, followed by a stationary test image shown for 100 milliseconds. Participants then marked the apparent centre. In the macaque recording experiment, the motion adaptor lasted three seconds before the stationary test; the neural population was read after the target appeared. These are matched in logic, not identical in timing or response method.
The control that carries the argument is the unchanged test target. The same object at the same pixel position can follow different motion directions. The researchers compared position reports and pre-trained neural decoders across those histories. A decoder trained before adaptation was applied to responses after adaptation without retraining. That prevents a post-hoc fitting step from creating the observed shift by learning a new coordinate system for each condition.
The direction matters more than a leaderboard
In the preprint, rightward adaptation made the human target appear on average 0.20° leftward; leftward adaptation produced an average 0.13° rightward bias. Those are angular visual-field shifts, not screen pixels. The paper reports statistical tests for the human effects and a comparable direction-opponent shift in macaque IT population decoding. The IT leftward-adaptor condition was weaker and not significant in the reported preprint analysis, so the result should not be reduced to perfect symmetry.
The authors tested feedforward image networks and systems with recurrence or video processing. These could encode object position, yet their representations generally failed to acquire the adaptation-induced position shifts seen in the biological measurements. When the researchers imposed transformations based on IT data, models could show the bias. That is an intervention demonstrating one computational route, not evidence the unmodified networks already learned it.
Qualitative comparison. Human angles are from the author preprint; cross-system magnitudes are not directly comparable.
- System
- Human observers
- Stationary location available?
- Yes
- History-dependent shift?
- Yes; opposite to prior motion
- System
- Macaque IT population
- Stationary location available?
- Yes, by neural decoding
- History-dependent shift?
- Direction-opponent shift, with condition-specific strength
- System
- Tested artificial networks
- Stationary location available?
- Yes
- History-dependent shift?
- Generally absent without an imposed IT-like transformation
| System | Stationary location available? | History-dependent shift? |
|---|---|---|
| Human observers | Yes | Yes; opposite to prior motion |
| Macaque IT population | Yes, by neural decoding | Direction-opponent shift, with condition-specific strength |
| Tested artificial networks | Yes | Generally absent without an imposed IT-like transformation |
The measured shifts and their asymmetry
The author preprint reports a mean human horizontal shift of 0.20° left after rightward adaptation and 0.13° right after leftward adaptation. Both were statistically significant in its analysis. It reports no corresponding systematic vertical shift, consistent with a horizontally moving adaptor. The unit is a degree of visual angle, which depends on the viewing geometry; treating these values as screen pixels would be wrong.
The macaque decoder moved in the same directional pattern, but the conditions were not equally strong. After rightward adaptation, its mean x estimate shifted 0.71° left and was significant. After leftward adaptation, the mean was 0.14° right, but the paper reports p = 0.33, so that individual condition was not statistically established. The preprint also found changes in IT population geometry after adaptation, using centred kernel alignment, and retained substantial position-decoding reliability. This argues against a simple story in which the neural signal just became random or weak. It still does not prove IT alone caused the human perceptual effect: the human behaviour and macaque recordings came from different subjects and protocols.
Direction-specific preprint results. These are not directly comparable effect sizes across measurement systems.
- Measure
- Human reported x-position
- After rightward motion
- 0.20° left; p < 0.001
- After leftward motion
- 0.13° right; p = 0.0036
- Measure
- Macaque IT decoded x-position
- After rightward motion
- 0.71° left; p < 0.001
- After leftward motion
- 0.14° right; p = 0.33, not significant
- Measure
- Vertical position
- After rightward motion
- No systematic human or IT shift reported
- After leftward motion
- No systematic human or IT shift reported
| Measure | After rightward motion | After leftward motion |
|---|---|---|
| Human reported x-position | 0.20° left; p < 0.001 | 0.13° right; p = 0.0036 |
| Macaque IT decoded x-position | 0.71° left; p < 0.001 | 0.14° right; p = 0.33, not significant |
| Vertical position | No systematic human or IT shift reported | No systematic human or IT shift reported |
A missing computation, not a universal verdict on AI
The study points to adaptation-driven reshaping of object representations as a candidate ingredient of biological vision. It does not prove that IT alone creates the illusion, and a correlated neural shift is not a complete causal account of perception. It also does not show that no future model could learn the effect. The comparison covers the architectures and training regimes the authors actually tested.
For model evaluation, the useful lesson is to test sequences as well as isolated images. An object-location benchmark that resets state for every frame can hide whether a system uses recent motion to update its representation. That may matter in video interpretation, robotics and human-like perceptual modelling, although this paper does not quantify a deployment failure in those systems. A fair follow-up would expose models to the same adaptor and target sequence, measure the direction and uncertainty of any shift, and compare that with human reports rather than simply asking whether the final image is classified correctly.
Which model tests failed, and what the successful intervention means
The researchers did more than pass one final image through one feedforward network. They examined standard object-recognition models, added unit-level suppression fitted to recorded IT adaptation dynamics, and tested systems with temporal processing. Suppression reduced activity and changed feature reliability, but did not create the direction-opponent position shift. In the preprint's model comparison, simulated-adaptation deltas for AlexNet, VGG-16, ResNet-18, ViT-L32 and SimCLR ResNet-50 stayed near zero, while the human and IT comparisons were much larger. The authors also tested the video architecture SlowFast and a recurrent network, ConvRNN; temporal input or feedback alone did not reproduce the biological pattern in those tests.
A separate intervention did create the effect: the authors estimated transformations from IT's own before-and-after adaptation activity and applied those transformations to model features. Position decoders on transformed VGG-16, ResNet-18 and SimCLR features then showed analogous directional biases. This is a useful sufficiency result: existing feature spaces can be made to express the shift when an empirically derived adaptation transform is imposed. It is not evidence that those networks learned such a transform on their own, nor a proposal that hallucinated object displacement is always desirable in a deployed vision system.
Three distinct model questions in the author preprint.
- Test
- Unmodified image networks
- Finding
- Location decodable; history-linked bias absent
- Interpretation
- Knowing pixel position is not the same as adapting perceived position
- Test
- Suppression, video and recurrence
- Finding
- No systematic direction-opponent shift in tested implementations
- Interpretation
- Generic temporal processing or lower activity was insufficient
- Test
- IT-derived transform applied
- Finding
- Shift appeared in transformed feature spaces
- Interpretation
- A biologically measured transform can induce the effect
| Test | Finding | Interpretation |
|---|---|---|
| Unmodified image networks | Location decodable; history-linked bias absent | Knowing pixel position is not the same as adapting perceived position |
| Suppression, video and recurrence | No systematic direction-opponent shift in tested implementations | Generic temporal processing or lower activity was insufficient |
| IT-derived transform applied | Shift appeared in transformed feature spaces | A biologically measured transform can induce the effect |
A better benchmark would preserve the sequence
The overlooked design lesson is to preserve the adaptor, delay and target as a sequence when evaluating models. A benchmark that resets the model before every target can only test its response to the final pixels. A stronger follow-up would present the same leftward and rightward histories, use a frozen position decoder, measure uncertainty and vertical controls, and compare the direction of the change with the biological reports. It would also test whether a model's shift improves a task that matters outside this illusion. Until such work exists, the result is a precise gap in the tested systems, not a universal failure of artificial vision.
Sources
- PrimaryYork University release on motion adaptation and vision networksYork Universityaccessed 2026-10-02
- PrimaryResearch preprint and methodsYork University researchersaccessed 2026-10-02
- PrimaryPublished study DOICurrent Biologyaccessed 2026-10-02


