Really cool work — learning over sequential experiences that contain the embodied cue of viewpoint as well as visual inputs, can give rise to human-like 3D shape perception!
tyler bonnenexcited to share some recent work! neural networks trained on multi-view sensory data are the first to match human-level 3D shape perception we predict human accuracy, error patterns, and reaction time—all zero-shot, no training on experimental data arxiv.org/abs/2602.17650 1/🧠