What can be learned about color from language?

preprint OA: closed
View at publisher

Abstract

Certain colors are strongly associated with certain adjectives (e.g. red is hot, blue is cold). Some of these associations are grounded in visual experiences such as seeing glowing red embers. Surprisingly, despite having no visual experience, many congenitally blind people show very similar color associations which are likely learned through language. We show that these associations are indeed embedded in the statistical structure of language. We apply a projection method to word embeddings trained on corpora of spoken and written language to identify color-adjective associations as they are represented in English. These projections were predictive of color-adjective associations reported by blind and sighted English speakers. The most predictive projections were generated by embeddings derived from a corpus of fiction, which outperformed even the state-of-the-art large language model, GPT-4. By augmenting the training corpora in various ways we discover the types of sentences most responsible for conveying the color-adjective associations to the models. We find that word embedding models learn these associations from indirect (second-order) co-occurrences, and that when prompted, people are able to identify some of the words that are most informative for associating colors with specific adjectives. Learning through linguistic co-occurrences is one way word meanings can be continually aligned across language users despite large variations in perceptual experience.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00