Humans have a distinct limit on the number of moving objects they can track simultaneously
the verdict
SUPPORTED
the evidence backs this
refutedsupported
the weight of evidence
8 sources for · 0 against
Peer-reviewed studies on multiple-object tracking consistently report that human observers have a distinct capacity limit, typically around three to four moving objects, although performance can vary depending on speed and task demands.
Attention allows us to monitor objects or regions of visual space and select information from them for report or storage. Classical theories of attention assumed a single focus of selection but many everyday activities, such as video games, navigating busy intersections, or watching over children at a swimming pool, require attention to multiple regions of interest. Laboratory tracking tasks have indeed demonstrated the ability to track four or more targets simultaneously. Although the mechanisms by which attention maintains contact with several targets are not yet established, recent studies have identified several characteristics of the tracking process, including properties defining a 'trackable' target, the maximum number of targets that can be tracked, and the hemifield independence of the tracking process. This research also has implications for computer vision, where there is a growing demand for multiple-object tracking.
The ability to divide attention enables people to keep track of up to four independently moving objects. We now show that this tracking capacity is independently constrained in the left and right visual fields as if separate tracking systems were engaged, one in each field. Specifically, twice as many targets can be successfully tracked when they are divided between the left and right hemifields as when they are all presented within the same hemifield. This finding places broad constraints on the anatomy and mechanisms of attentive tracking, ruling out a single attentional focus, even one that moves quickly from target to target.
Much of our interaction with the visual world requires us to isolate some currently important objects from other less important objects. This task becomes more difficult when objects move, or when our field of view moves relative to the world, requiring us to track these objects over space and time. Previous experiments have shown that observers can track a maximum of about 4 moving objects. A natural explanation for this capacity limit is that the visual system is architecturally limited to handling a fixed number of objects at once, a so-called magical number 4 on visual attention. In contrast to this view, Experiment 1 shows that tracking capacity is not fixed. At slow speeds it is possible to track up to 8 objects, and yet there are fast speeds at which only a single object can be tracked. Experiment 2 suggests that that the limit on tracking is related to the spatial resolution of attention. These findings suggest that the number of objects that can be tracked is primarily set by a flexibly allocated resource, which has important implications for the mechanisms of object tracking and for the relationship between object tracking and other cognitive processes.
Mounting evidence suggests that visual attention may be simultaneously deployed to multiple distinct object locations, but the constraints upon this multi-object attentional system are still debated. Results from multiple object tracking (MOT) experiments have been interpreted as revealing a fixed attentional capacity limit of 4 objects, while other evidence has suggested that attentional capacity may be more fluid. Here, we investigated the influence of target stimulus factors, such as speed and size, and of distractor filtering factors, such as number of distractors and screen density, on MOT performance. Each factor had significant effects on capacity, producing values that ranged from above 6 objects down to one object, depending on the task demands. Although our results support the view that crowding effects modulate the effective capacity of attention, we also find evidence that central processes related to distractor suppression and target enhancement modulate capacity.
Overall performance when tracking moving targets is known to be poorer for larger numbers of targets, but the specific effect on tracking's temporal resolution has never been investigated. We document a broad range of display parameters for which visual tracking is limited by temporal frequency (the interval between when a target is at each location and a distracter moves in and replaces it) rather than by object speed. We tested tracking of one, two, and three moving targets while the eyes remained fixed. Variation of the number of distracters and their speed revealed both speed limits and temporal frequency limits on tracking. The temporal frequency limit fell from 7 Hz with one target to 4 Hz with two targets and 2.6 Hz with three targets. The large size of this performance decrease implies that in the two-target condition participants would have done better by tracking only one of the two targets and ignoring the other. These effects are predicted by serial models involving a single tracking focus that must switch among the targets, sampling the position of only one target at a time. If parallel processing theories are to explain why dividing the tracking resource reduces temporal resolution so markedly, supplemental assumptions will be required.
When tracking multiple moving targets among visually similar distractors, human observers are capable of distributing attention over several spatial locations. It is unclear, however, whether capacity limitations or perceptual–cognitive abilities are responsible for the development of expertise in multiple object tracking. Across two experiments, we examined the role of working memory and visual attention in tracking expertise. In Experiment 1, individuals who regularly engaged in object tracking sports (soccer and rugby) displayed improved tracking performance, relative to non-tracking sports (swimming, rowing, running) (p = 0.02, ηp2 = 0.163), but no differences in gaze strategy (ps > 0.31). In Experiment 2, participants trained on an adaptive object tracking task showed improved tracking performance (p = 0.005, d = 0.817), but no changes in gaze strategy (ps > 0.07). They did, however, show significant improvement in a working memory transfer task (p < 0.001, d = 0.970). These findings indicate that the development of tracking expertise is more closely linked to processing capacity limits than perceptual–cognitive strategies.
Driving on a busy road, eluding a group of predators, or playing a team sport involves keeping track of multiple moving objects. In typical laboratory tasks, the number of visual targets that humans can track is about four. Three types of theories have been advanced to explain this limit. The fixed-limit theory posits a set number of attentional pointers available to follow objects. Spatial interference theory proposes that when targets are near each other, their attentional spotlights mutually interfere. Resource theory asserts that a limited resource is divided among targets, and performance reflects the amount available per target. Utilising widely separated objects to avoid spatial interference, the present experiments validated the predictions of resource theory. The fastest target speed at which two targets could be tracked was much slower than the fastest speed at which one target could be tracked. This speed limit for tracking two targets was approximately that predicted if at high speeds, only a single target could be tracked. This result cannot be accommodated by the fixed-limit or interference theories. Evidently a fast target, if it moves fast enough, can exhaust attentional resources.
Human visual processing is limited-we can only track a few moving objects at a time and store a few items in visual working memory (WM). A shared mechanism that may underlie these performance limits is how the visual system parses a scene into representational units. In the present study, we explored whether multiple-object tracking (MOT) and WM rely on a common item-based indexing mechanism. We measured the contralateral delay activity (CDA), an event-related slow wave that tracks load in an item-based manner, as participants completed a combined WM and MOT task, concurrently tracking items and remembering visual information. In Experiment 1, participants tracked one or two moving discs without needing to remember the discs' colors (track and ignore condition) or while also remembering the discs' colors (two or four colors in total; track and remember condition). In Experiment 2, participants attended either two static discs or two moving discs, while remembering the discs' colors (two or four colors). In both experiments, the CDA was largely determined by the tracking task-CDA amplitudes reflected the number of tracked discs and not the number of to-be-remembered colors. However, when the discs were static, the CDA amplitudes did reflect color load. We discuss this set of findings in relation to longstanding theories of visual cognition (fingers of instantiation and object files) and the implications for cognitive models of representation of visual information-that how a scene is parsed into item-based representations is a key mechanism in the operation of WM.
Everything we examined (8)
This check searched the claim as stated. It did not run a separate search for evidence against it.