Sparse mixed codes on shared manifolds for human-like spatial attention in artificial neural networks

preprint OA: closed CC-BY-NC-ND-4.0

Abstract

Spatial attention is often partitioned into endogenous, exogenous, and social forms, yet it remains unclear whether a single neural circuit can support all three and how their population codes are organized. Here we trained recurrent artificial neural networks (ANNs) with convolutional sensory front-ends on three classic cueing paradigms (central, peripheral, and gaze cues) to reproduce human reaction time (RT) profiles across cue–target onset asynchronies. Despite differences in sensory architecture and visual experience, all ANNs captured the full facilitation–inhibition time course for all three attention types. Model-based targeted dimensionality reduction (mTDR) revealed that cue- and choice-related activity in the advanced cognitive module evolved as rotations within a shared low-dimensional manifold, with angular deflections that mirrored the distinct temporal dynamics of endogenous, exogenous, and social attention. Attentional signals were encoded by highly sparse, distributed population activity: a small subset of recurrent units explained most task-related variance, was sufficient to recover human-like RT patterns after virtual lesioning, and became progressively sparser as training improved performance. At the same time, single unit responses displayed pervasive mixed selectivity, dominated by nonlinear conjunctions of cue type, cue direction, and validity, whose strength and heterogeneity robustly predicted model performance. Together, these results identified low-dimensional geometric rotations, sparse coding, and nonlinear mixed selectivity as core computational principles through which a single recurrent circuit could generate human-like temporal dynamics across endogenous, exogenous, and social orienting, and provided testable predictions for population-level mechanisms of spatial attention in the brain.
Full text 1,938 characters · extracted from oa-html · click to expand
Abstract Spatial attention is often partitioned into endogenous, exogenous, and social forms, yet it remains unclear whether a single neural circuit can support all three and how their population codes are organized. Here we trained recurrent artificial neural networks (ANNs) with convolutional sensory front-ends on three classic cueing paradigms (central, peripheral, and gaze cues) to reproduce human reaction time (RT) profiles across cue–target onset asynchronies. Despite differences in sensory architecture and visual experience, all ANNs captured the full facilitation–inhibition time course for all three attention types. Model-based targeted dimensionality reduction (mTDR) revealed that cue- and choice-related activity in the advanced cognitive module evolved as rotations within a shared low-dimensional manifold, with angular deflections that mirrored the distinct temporal dynamics of endogenous, exogenous, and social attention. Attentional signals were encoded by highly sparse, distributed population activity: a small subset of recurrent units explained most task-related variance, was sufficient to recover human-like RT patterns after virtual lesioning, and became progressively sparser as training improved performance. At the same time, single unit responses displayed pervasive mixed selectivity, dominated by nonlinear conjunctions of cue type, cue direction, and validity, whose strength and heterogeneity robustly predicted model performance. Together, these results identified low-dimensional geometric rotations, sparse coding, and nonlinear mixed selectivity as core computational principles through which a single recurrent circuit could generate human-like temporal dynamics across endogenous, exogenous, and social orienting, and provided testable predictions for population-level mechanisms of spatial attention in the brain. Competing Interest Statement The authors have declared no competing interest.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-NC-ND-4.0