Logo image
On the Convergent Validity of Offline Evaluation Designs for Recommender Systems
Preprint   Open access

On the Convergent Validity of Offline Evaluation Designs for Recommender Systems

Sushobhan Parajuli, Samira Vaez Barenji and Michael D Ekstrand
arXiv.org
27 Jul 2026
url
https://doi.org/10.48550/arXiv.2607.25097View
Preprint (Author's original) Open arXiv.org - Non-exclusive license to distribute

Abstract

Computer Science - Information Retrieval
Offline evaluation on historical interaction logs is the most common evaluation methodology for recommender systems. However, such evaluations depend on sparse, incomplete, or biased data, which raises concerns about whether commonly used evaluation setups reliably reflect true user preferences. In this work, we study how offline evaluation design choices affect the validity of recommender system comparisons. We evaluate a set of recommendation models across several evaluation setups that vary key factors such as data filtering thresholds and candidate set construction. To assess the validity of these configurations, we measure the correlation between model rankings obtained from conventional train-test splits on sparse interaction data and rankings from evaluations based on dense ground-truth user feedback. We use this agreement as an indication of their validity with respect to true user preferences. Our results show that the validity of sparse evaluation depends on the dataset and the specific dense evaluation targets, and that there is no uniformly best offline evaluation design.

Metrics

1 Record Views

Details

Logo image