Automatic Scoring of Creative Problem-Solving with Large Language Models: A Comparison of Originality and Quality Ratings
preprint
OA: closed
CC-BY-4.0
AI-generated summary
Large language models RoBERTa and GPT-2 effectively scored both originality and quality of creative problem-solving task responses, outperforming other methods and generalizing to new items.
One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works
Abstract
Creative problem-solving is a naturalistic form of creative thinking involving the generation of solutions that are not only original but also of high quality (i.e., plausible and effective). Naturalistic tasks that evaluate both originality and quality are vital for the promotion of creativity in real-world settings—yet scoring such tasks remains challenging, due to costly human labor required to manually rate task responses. Past work has shown that large language models (LLMs) can be trained to predict human originality ratings of responses to tests of divergent thinking. In the present research, we extend this work to creative problem-solving, examining whether both originality and quality can be automatically scored for a naturalistic creativity task. We gathered data from 10 studies, amounting to 3,235 participants who completed a creative problem-solving task (CPST). We then fine-tuned two open-source LLMs, RoBERTa and GPT-2, to predict human ratings of originality and quality on the CPST, and compared their performance to two other scoring methods: elaboration (i.e., word count) and semantic distance. We found that RoBERTa and GPT-2 models predict solution quality (RoBERTa, r = .83; GPT-2, r = .83) better than solution originality (RoBERTa, r = .79; GPT-2, r = .80). Moreover, we found that both models outperformed elaboration and semantic distance methods and generalized to new CPST items not present in their training set. We therefore show for the first time that naturalistic creativity tasks can be automatically scored for both originality and quality. Open access is provided to the models and training data.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00
- unpaywall
- last seen: 2026-05-28T02:00:01.590549+00:00
License: CC-BY-4.0