Can Centaur Truly Simulate Human Cognition? The Fundamental Limitation of Instruction Understanding
preprint
OA: closed
AI-generated summary
Centaur's high performance across cognitive tasks is maintained even when instructions are removed or misleading, suggesting reliance on statistical cues rather than true comprehension.
One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works
Abstract
Recent advances in cognitive modeling have demonstrated the potential of large language models (LLMs) to unify diverse aspects of human cognition. The Centaur model, an LLM fine-tuned on cognitive tasks, achieves high performance across 160 psychological experiments, suggesting that a single model may capture multiple cognitive processes. However, whether this success stems from genuine task understanding or exploitation of superficial statistical cues remains unclear. To test this, we systematically manipulated Centaur’s input by (1) removing task instructions, (2) removing all contextual information, and (3) providing misleading instructions. All three manipulations remove information necessary for humans to perform the tasks. Results show that Centaur often maintains high performance under these manipulations, outperforming both baseline cognitive models and the unfine-tuned LLM (Llama) that receives correct instructions. These findings indicate that Centaur’s success likely relies on superficial statistical cues rather than true instruction comprehension. Our study highlights the need for more diverse out-of-distribution tests for LLM-based cognitive models.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00