Shared Representations of Stability in Humans, Supervised & Unsupervised Deep Neural Networks

preprint OA: closed
View at publisher

Abstract

At a glance, the human visual system transforms complex retinal images into generic feature representations useful for guiding a wide range of behaviors. Here, we provide evidence that the feature representations embedded in purely feed- forward neural networks are sufficient to explain seemingly high-level human judgments: in this case, the stability of randomly arranged block towers. To do this, we first show that we can linearly decode stability from the features of two deep neural networks – a supervised network trained on ImageNet, and a varia- tional autoencoder trained only to reconstruct images of block towers from various perspectives – neither of which were ever taught stability per se. We then derive a set of image-computable features to use as predictors of performance, finding that the stability judgments of both human subjects and the neural net decoders are best predicted by the same feature. Our findings suggest overall that at least some aspects of seemingly higher-level reasoning in a now paradigmatic intuitive physics task may be grounded in direct readouts of purely perceptual features.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00