Semantic Information Orthogonal to Visual Features Peaks in Lateral Occipitotemporal Cortex
preprint
OA: closed
CC-BY-4.0
Abstract
Language model embeddings of scene descriptions predict responses in the human higher visual cortex. However, a fundamental question remains: does this alignment reflect truly visually-independent semantic content, or does it occur because language models better mimic the complex visual features that drive these areas? We used 7T fMRI data from the Natural Scenes Dataset to directly address this by removing the influence of visual feature embeddings from language model embeddings, isolating semantic content that is separate from the visual signal. We then used these visually-independent embeddings to predict brain responses in individual voxels through cross-validated ridge regression. After adjusting for visual signals, we found a clear difference in brain regions: the lateral occipitotemporal cortex, especially in areas selective for body perception, showed significantly more visually-independent semantic variance compared to ventral stream regions. In contrast, the early visual cortex displayed notably negative predictions after adjustment, confirming that our method effectively removed visually-driven signals. This pattern was consistent across all eight subjects, both hemispheres, and six combinations of language models and visual feature architectures. These findings suggest that the lateral stream retains substantially more variance from language models unrelated to various visual feature models than the ventral stream does. This suggests that visually independent semantic coding is organized heterogenously along the occipital cortex. Highlights Body-selective lateral occipitotemporal cortex (EBA) contains the strongest visually-independent semantic representations in human visual cortex. After removing visual feature variance, semantic encoding is significantly greater in lateral stream regions than in canonical ventral stream areas (FFA, PPA, RSC).The lateral-over-ventral dissociation is architecture-invariant, replicating across six combinations of language models (BERT, GPT-2, CLIP-text) and visual feature sets, with GPT-2 > BERT > CLIP-text ordering validating the pipeline.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00
- unpaywall
- last seen: 2026-05-28T02:00:01.590549+00:00
License: CC-BY-4.0