12Classically, visual processing is described as a cascade of local feedforward computations. Feedforward 13Convolutional Neural Networks (ffCNNs) have shown how powerful such models can be. Previously, 14 using visual crowding as a well-controlled challenge, we showed that no classic model of vision, 15including ffCNNs, can explain human global shape processing (1). Here, we show that Capsule Neural 16 Networks (CapsNets; 2), combining ffCNNs with a grouping and segmentation mechanism, solve this 17 challenge. We also show that ffCNNs and standard recurrent networks do not, suggesting that the 18 grouping and segmentation capabilities of CapsNets are crucial. Furthermore, we provide 19 psychophysical evidence that grouping and segmentation is implemented recurrently in humans, and 20show that CapsNets reproduce these results well. We discuss why recurrence seems needed to 21 implement grouping and segmentation efficiently. Together, we provide mutually reinforcing 22 psychophysical and computational evidence that a recurrent grouping and segmentation process is 23 essential to understand the visual system and create better models that harness global shape 24 computations. 25
26Author Summary 27 Feedforward Convolutional Neural Networks (ffCNNs) have revolutionized computer vision and are 28 deeply transforming neuroscience. However, ffCNNs only roughly mimic human vision. There is a 29 rapidly expanding literature investigating differences between humans and ffCNNs. Several findings 30 suggest that, unlike humans, ffCNNs rely mostly on local visual features. Furthermore, ffCNNs lack 31 recurrent connections, which abound in the brain. Here, we use visual crowding, a well-known 32 psychophysical phenomenon, to investigate recurrent computations in global shape processing. 33Previously, we showed that no model based on the classic feedforward framework of vision, including 34 ffCNNs, can explain global effects in crowding. Here, we show that Capsule Networks (CapsNets), 35combining ffCNNs with recurrent grouping and segmentation, solve this challenge. Lateral and top-36 down recurrent connections do not, suggesting that grouping and segmentation are crucial for 37 human-like global computations. Based on these results, we hypothesize that one computational 38 function of recurrence is to efficiently implement grouping and segmentation. We provide 39 psychophysical evidence that, indeed, recurrent processes implement grouping and segmentation in 40 humans. CapsNets reproduce these results too. Together, we provide mutually reinforcing 41 computational and psychophysical evidence that a recurrent grouping and segmentation process is 42 essential to understand the visual system and create better models that harness global shape 43 computations. 44