Test linking by pairwise comparisons of item difficulty

preprint OA: closed
View at publisher

Abstract

To obtain comparability between scores on different forms of tests and exams, one can use test equating or linking methods. Most equating and linking methods require either common (anchor) items or equivalent groups of test takers to link the tests. However, in some circumstances neither of these two situations may occur. In these situations, expert judgment is required to assess the difficulty of the items of the two tests. One method to use expert judgment for linking tests is pairwise linking. This is done by asking experts to compare one item from Test 1 to one item from Test 2, do this for several different item pairs, and rate which of the items of the pair they think is most difficult on the attribute that the tests measure. In this chapter, we introduce an adjusted Bradley-Terry-Luce (BTL) model for pairwise linking that uses both information from the expert judgments and from the items of the separate tests. We tested the performance of the model in terms of bias and precision of the shift parameters that are used to link the tests. The adjusted BTL model for pairwise linking can be used to link tests without common items or equivalent groups. The model can be used for linking tests with dichotomous items and provides a starting point for model adaptations allowing for more general situations like test with polytomous scored items. We advise to use at least 800 pairwise item comparisons from experts about the relative difficulty of the item for a student with average ability as input to obtain a sufficiently precise link between the two tests.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00