Human annotation is a key step in data-driven modeling, yet traditional approaches seek consensus among raters, treating disagreement as error and failing to capture the complexity of human interpretation. This has given rise to the perspectivist approach, which explicitly models annotator variability and embraces multiple viewpoints. In this study, we apply a Bayesian Nonparametric Multidimensional Item Response Theory model to multi-rater annotation, adopting a formulation where annotated texts are treated as persons carrying latent traits and annotators function as items. This allows us to automatically identify groups of annotators and assign to each text a set of scores over latent dimensions whose number is inferred directly from the data. We demonstrate the approach through a case study involving social media comments on immigration, annotated independently by multiple raters for the presence of racist content. The model uncovers the structure of annotator heterogeneity, offering a model-based alternative to consensus-based labeling. We identified two distinct annotator clusters with systematically different perspectives, yielding group-specific severity scores. Disagreement was found to concentrate around politically charged language, where the boundary between opinion and hateful rhetoric emerged as contested.
Modeling Multi-rater Behavior with Bayesian Nonparametric MIRT: Inferring Latent Traits and Group Structure
Alex Cucco
;Lara Fontanella;Pasquale Valentini;
2026-01-01
Abstract
Human annotation is a key step in data-driven modeling, yet traditional approaches seek consensus among raters, treating disagreement as error and failing to capture the complexity of human interpretation. This has given rise to the perspectivist approach, which explicitly models annotator variability and embraces multiple viewpoints. In this study, we apply a Bayesian Nonparametric Multidimensional Item Response Theory model to multi-rater annotation, adopting a formulation where annotated texts are treated as persons carrying latent traits and annotators function as items. This allows us to automatically identify groups of annotators and assign to each text a set of scores over latent dimensions whose number is inferred directly from the data. We demonstrate the approach through a case study involving social media comments on immigration, annotated independently by multiple raters for the presence of racist content. The model uncovers the structure of annotator heterogeneity, offering a model-based alternative to consensus-based labeling. We identified two distinct annotator clusters with systematically different perspectives, yielding group-specific severity scores. Disagreement was found to concentrate around politically charged language, where the boundary between opinion and hateful rhetoric emerged as contested.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


