- Main
Parameter-Driven Consensus-based Filtering Improves Collective Judgment Reliability in Crowdsourced Annotation
Abstract
Crowdsourced annotation can be viewed as a form of distributed human judgment in which individual reliability and task difficulty jointly shape collective decisions. Probabilistic aggregation models estimate latent annotator reliability and task difficulty, but are typically used only to weight judgments or infer labels, not to guide selective dataset refinement. We propose a replicated-dataset consensus filtering framework that improves collective judgment reliability by removing unreliable annotators and difficult tasks based on latent reliability-difficulty parameters. Instead of relying on a single dataset estimate, the method constructs multiple replicated datasets by small random label removal, re-estimates latent parameters on each replicated dataset, and removes components that are consistently selected across replicated datasets. Filtering is applied iteratively with entropy-based stopping conditions. The framework is model-agnostic and can be combined with any probabilistic aggregation model that estimates reliability and difficulty parameters. Using a standard reliability-difficulty model instantiation, experiments on multiple crowdsourced datasets show that parameter-driven consensus filtering improves aggregation accuracy and removes incorrect tasks more selectively than random removal. These results support a cognitively grounded, parameter-based approach to improving collective judgment reliability in distributed annotation settings.