Improving crowdsourcing for AI through cognitive-inspired data engineering
preprint
OA: closed
CC-BY-4.0
Abstract
Crowdsourcing offers a fast and cost-efficient approach to obtaining human labeled datasets. However, crowdsourced datasets and the models trained on them can inherit the cognitive constraints and biases of their annotators. In a process we refer to as cognitive-inspired data engineering, we investigate whether ideas from cognitive science can be applied to mitigate the presence of cognitive constraints and cognitive biases in crowdsourced datasets and, as a result, improve the performance of models trained on these datasets. We evaluate our approach by crowdsourcing labels for medical image diagnostic tasks using two different crowdsourcing platforms across two experiments. In Experiment 1, we collect subjective probability judgments from novice annotators through Amazon Mechanical Turk and, in Experiment 2, we collect subjective probability judgments and binary classifications from skilled annotators though DiagnosUs, a crowdsourcing platform specializing in medical and scientific data annotation. In both experiments, we find that de-biasing subjective probability judgments via recalibration leads to more accurate crowdsourced datasets and more accurate models trained on these datasets. Our results suggest that cognitive-inspired data engineering offers a promising avenue to improve the quality of crowdsourced datasets\textcolor{black}{, with consistent downstream benefits for machine learning models
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00
- unpaywall
- last seen: 2026-06-02T02:00:03.124865+00:00
License: CC-BY-4.0