Systematic comparison of color representations between humans and deep neural networks: towards predicting human color perception in a vast color space

preprint OA: closed CC-BY-4.0
🔓 Open OA copy View at publisher
AI-generated summary by claude@2026-07, 2026-07-15

Contrastive language-image pre-training (CLIP) models align best with human color perception, enabling predictions for the large-scale structure of color representation across thousands of colors.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

The representational structure of large-scale human color perception remains incompletely understood. While classical studies measured numerous color pairs, these measurements compared only similar colors, and exploring the global relationships among thousands of colors has been infeasible due to the time costs of psychophysical experiments. Given these constraints, deep neural networks (DNNs) have attracted attention as a promising tool for providing proxies or predictions of human perception beyond the scope of psychophysical experiments. However, it remains unclear which DNNs possess embeddings that geometrically align with human color perception. Furthermore, it is unclear which learning paradigm enables DNNs to acquire a color representation that aligns with that of humans. Here, we systematically investigate which learning paradigm enables DNNs to produce a color representation that is structurally congruent with that of humans, with a focus on three types: self-supervised learning (SSL) that trains on images alone, supervised learning (SL) that trains on images with category labels, and contrastive language-image pre-training (CLIP) that trains on image-text pairs. We compared the embeddings of DNNs with the human similarity judgments of 93 colors using a rigorous unsupervised method termed Gromov-Wasserstein Optimal Transport (GWOT). Our results show that, while each learning paradigm acquires color representations that strongly align with human data at the fine-item level in early layers, only CLIP sustains such a representation at the output. Furthermore, when we leveraged a key advantage of DNNs and investigated the representational structure of 4096 colors, the early layers of each learning paradigm and the output of CLIP consistently converged on their own characteristic structures. These structures present plausible predictions for the large-scale human color representation. Our work demonstrates an approach for exploring unknown territories of human perception through the use of computational models validated in a limited empirical space, and provides predictions for future large-scale psychophysical experiments. Author summary How do we perceive the vast world of color? Despite extensive research into human color perception, studies evaluating many colors have mostly captured differences between similar colors, while those mapping global relationships are restricted to a few dozen. Consequently, we still do not know the global structure of the massive “color map” that might underlie our perception of thousands of colors, as testing this directly is practically impossible. To explore this space, we turned to deep neural networks, a form of AI. Our first step was to identify models that “see” color in a way that matches humans. We compared models against human data capturing the global relationships among all possible pairs of 93 colors. Using a powerful geometric comparison method, we found the models that matched the human color map. This allowed us to use these models as reliable computational proxies. We then used them to do what human experiments currently cannot: chart a vast global map of 4,096 colors. The human-aligned models consistently converged on two distinct structures. Our work provides the first plausible, testable predictions for the large-scale structure of human color perception and demonstrates a new way to explore otherwise unreachable territories of our perceptual world.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-29T02:00:03.542394+00:00
License: CC-BY-4.0