Enhancing Scale Development: Pseudo Factor Analysis of Language Embedding Similarity Matrices

preprint OA: closed
View at publisher

Abstract

This article builds and extends on recent work using Large Language Models (LLMs) in psychometrics to generate pseudo-discrimination parameters. While earlier work looked at pseudo-discrimination on an item-by-construct basis, we introduce pseudo-factor analysis to enhance scale design. Pseudo-factor analysis is a data-less, model-based approach to evaluating aspects of a latent construct’s measurement model, such as dimensionality and the relations between factors and their indicators. Across two studies using Five and Six-factor personality frameworks, a variety of sentence transformer models, and three encoding approaches (i.e., atomic, atomic reversed, and macro), pseudo-factor analyses recovered theoretically expected structures. These pseudo-factor structures were strongly related to their established empirical factor structures based on factor analyses of human ratings in prior published research. We suggest that Pseudo-Factor Analysis is a viable method for checking and potentially modifying scale items after item generation and before item trialing. We provide a Shiny application for calculating pseudo-factor analysis parameters and related psychometric estimates.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00