Article detail · 2006
Estimating average precision with incomplete and imperfect judgments
- Year
- 2006
- Type
- conference-abstract
Abstract
OpenAlex · English
We consider the problem of evaluating retrieval systems using incomplete judgment information. Buckley and Voorhees recently demonstrated that retrieval systems can be efficiently and effectively evaluated using incomplete judgments via the bpref measure [6]. When relevance judgments are complete, the value of bpref is an approximation to the value of average precision using complete judgments. However, when relevance judgments are incomplete, the value of bpref deviates from this value, though it continues to rank systems in a manner similar to average precision evaluated with a complete judgment set. In this work, we propose three evaluation measures that (1) are approximations to average precision even when the relevance judgments are incomplete and (2) are more robust to incomplete or imperfect relevance judgments than bpref. The proposed estimates of average precision are simple and accurate, and we demonstrate the utility of these estimates using TREC data.
Topics
Citations
OpenAlex cited_by_count. Not a WoS or Scopus citation count; those sources have no separate column here.
330 citations
OpenAlex cited_by_count (cache / database)
24 publications in the local catalog that cite this work (OpenAlex reference match; not the full global list).
- A new rank correlation coefficient for information retrieval 2008
- A simple and efficient sampling method for estimating AP and NDCG 2008
- Relevance assessment 2008
- Extending average precision to graded relevance judgments 2010
- Document selection methodologies for efficient and effective learning-to-rank 2009
- Estimating average precision when judgments are incomplete 2007
- Estimation of Fair Ranking Metrics with Incomplete Judgments 2021
- Inferring document relevance from incomplete information 2007
- An uncertainty-aware query selection model for evaluation of IR systems 2012
- Intelligent Topic Selection for Low-Cost Information Retrieval Evaluation: A New Perspective on Deep vs. Shallow Judging 2018