Towards optimal cardinality estimation of unions and intersections with sketches.

Description: 

Estimating the cardinality of unions and intersections of sets is a problem of interest in OLAP. Large data applications often require the use of approximate methods based on small sketches of the data. We give new estimators for the cardinality of unions and intersection and show they approximate an optimal estimation procedure. These estimators enable the improved accuracy of the streaming MinCount sketch to be exploited in distributed settings. Both theoretical and empirical results demonstrate substantial improvements over existing methods.

Authors: 
Daniel Ting
Publication Date: 
Tuesday, August 2, 2016
Publication Information: 
ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. (KDD), 2016.