Fast Estimation of the Pattern Frequency Spectrum

van Leeuwen, Matthijs; Ukkonen, Antti

doi:10.1007/978-3-662-44851-9_8

Matthijs van Leeuwen²³ &
Antti Ukkonen²⁴

Part of the book series: Lecture Notes in Computer Science ((LNAI,volume 8725))

Included in the following conference series:

Joint European Conference on Machine Learning and Knowledge Discovery in Databases

4171 Accesses
5 Citations

Abstract

Both exact and approximate counting of the number of frequent patterns for a given frequency threshold are hard problems. Still, having even coarse prior estimates of the number of patterns is useful, as these can be used to appropriately set the threshold and avoid waiting endlessly for an unmanageable number of patterns. Moreover, we argue that the number of patterns for different thresholds is an interesting summary statistic of the data: the pattern frequency spectrum.

To enable fast estimation of the number of frequent patterns, we adapt the classical algorithm by Knuth for estimating the size of a search tree. Although the method is known to be theoretically suboptimal, we demonstrate that in practice it not only produces very accurate estimates, but is also very efficient. Moreover, we introduce a small variation that can be used to estimate the number of patterns under constraints for which the Apriori property does not hold. The empirical evaluation shows that this approach obtains good estimates for closed itemsets.

Finally, we show how the method, together with isotonic regression, can be used to quickly and accurately estimate the frequency pattern spectrum: the curve that shows the number of patterns for every possible value of the frequency threshold. Comparing such a spectrum to one that was constructed using a random data model immediately reveals whether the dataset contains any structure of interest.

Download to read the full chapter text

Chapter PDF

Introduction to Pattern Mining

The pattern frequency distribution theory: a mathematic establishment toward rational and reliable pattern mining

Article 20 August 2022

Computing Theoretically-Sound Upper Bounds to Expected Support for Frequent Pattern Mining Problems over Uncertain Big Data

Keywords

These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.

References

Agrawal, R., Srikant, R.: Fast algorithms for mining association rules in large databases. In: Proceedings of VLDB 1994, pp. 487–499 (1994)
Google Scholar
Agrawal, R., Srikant, R.: Mining sequential patterns. In: Proceedings of ICDE 1995, pp. 3–14 (1995)
Google Scholar
Barlow, R.E., Brunk, H.D.: The isotonic regression problem and its dual. Journal of the American Statistical Association 67(337), 140–147
Google Scholar
De Bie, T.: Maximum entropy models and subjective interestingness: An application to tiles in binary databases. Data Min. Knowl. Discov. 23(3), 407–446 (2011)
Article MATH MathSciNet Google Scholar
Boley, M., Gärtner, T., Grosskreutz, H.: Formal concept sampling for counting and threshold-free local pattern mining. In: Proc. of SDM 2010, pp. 177–188 (2010)
Google Scholar
Boley, M., Grosskreutz, H.: A randomized approach for approximating the number of frequent sets. In: Proceedings of ICDM 2008, pp. 43–52 (2008)
Google Scholar
Bringmann, B., Nijssen, S., Tatti, N., Vreeken, J., Zimmermann, A.: Mining sets of patterns: Next generation pattern mining. In: Tutorial at ICDM 2011 (2011)
Google Scholar
Chen, P.C.: Heuristic sampling: A method for predicting the performance of tree searching programs. SIAM Journal on Computing 21(2), 295–315 (1992)
Article MATH Google Scholar
Cheng, H., Yan, X., Han, J., Hsu, C.-W.: Discriminative frequent pattern analysis for effective classification. In: Proceedings of the ICDE, pp. 716–725 (2007)
Google Scholar
Goethals, B., Moens, S., Vreeken, J.: MIME: A framework for interactive visual pattern mining. In: Proceedings of KDD 2011, pp. 757–760 (2011)
Google Scholar
Gunopulos, D., Khardon, R., Mannila, H., Saluja, S., Toivonen, H., Sharm, R.S.: Discovering all most specific sentences. ACM Trans. Database Syst. 28(2), 140–174 (2003)
Article Google Scholar
Jerrum, M., Sinclair, A.: The markov chain monte carlo method: An approach to approximate counting and integration. Approximation Algorithms for NP-hard Problems, 482–520 (1996)
Google Scholar
Kilby, P., Slaney, J.K., Thiébaux, S., Walsh, T.: Estimating search tree size. In: Proceedings of AAAI 2006, pp. 1014–1019 (2006)
Google Scholar
Knuth, D.E.: Estimating the efficiency of backtrack programs. Mathematics of Computation 29(129), 122–136 (1975)
Article MathSciNet Google Scholar
Liu, G., Lu, H., Yu, J.X., Wang, W., Xiao, X.: AFOPT: An efficient implementation of pattern growth approach. In: Proc. of FIMI at ICDM 2003 (2003)
Google Scholar
Purdom, P.W.: Tree size by partial backtracking. SIAM Journal on Computing 7(4), 481–491 (1978)
Article MATH MathSciNet Google Scholar
Vreeken, J., van Leeuwen, M., Siebes, A.: Krimp: mining itemsets that compress. Data Min. Knowl. Discov. 23(1), 169–214 (2011)
Article MATH MathSciNet Google Scholar
Yan, X., Han, J.: gSpan: Graph-based substructure pattern mining. In: Proceedings of ICDM 2002, pp. 721–724 (2002)
Google Scholar
Yang, G.: The complexity of mining maximal frequent itemsets and maximal frequent patterns. In: Proceedings of KDD 2004, pp. 344–353 (2004)
Google Scholar

Download references

Author information

Authors and Affiliations

Department of Computer Science, KU Leuven, Belgium
Matthijs van Leeuwen
Helsinki Institute for Information Technology HIIT, Aalto University, Finland
Antti Ukkonen

Authors

Matthijs van Leeuwen
View author publications
You can also search for this author in PubMed Google Scholar
Antti Ukkonen
View author publications
You can also search for this author in PubMed Google Scholar

Editor information

Editors and Affiliations

Faculty of Applied Sciences,Department of Computer and Decision Engineering, Université Libre de Bruxelles, Av. F. Roosevelt, CP 165/15, 1050, Brussels, Belgium
Toon Calders
Dipartimento di Informatica, Università degli Studi “Aldo Moro”, via Orabona 4, 70125, Bari, Italy
Floriana Esposito
Department of Computer Science, Universität Paderborn, Warburger Str. 100, 33098, Paderborn, Germany
Eyke Hüllermeier
Dipartimento di Informatica, Università degli Studi di Torino, Corso Svizzera 185, 10149, Torino, Italy
Rosa Meo

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

van Leeuwen, M., Ukkonen, A. (2014). Fast Estimation of the Pattern Frequency Spectrum. In: Calders, T., Esposito, F., Hüllermeier, E., Meo, R. (eds) Machine Learning and Knowledge Discovery in Databases. ECML PKDD 2014. Lecture Notes in Computer Science(), vol 8725. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-662-44851-9_8

Download citation

DOI: https://doi.org/10.1007/978-3-662-44851-9_8
Publisher Name: Springer, Berlin, Heidelberg
Print ISBN: 978-3-662-44850-2
Online ISBN: 978-3-662-44851-9
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics

Fast Estimation of the Pattern Frequency Spectrum

Abstract

Chapter PDF

Similar content being viewed by others

Introduction to Pattern Mining

The pattern frequency distribution theory: a mathematic establishment toward rational and reliable pattern mining

Computing Theoretically-Sound Upper Bounds to Expected Support for Frequent Pattern Mining Problems over Uncertain Big Data

Keywords

References

Author information

Authors and Affiliations

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Publish with us

Navigation

Fast Estimation of the Pattern Frequency Spectrum

Abstract

Chapter PDF

Similar content being viewed by others

Introduction to Pattern Mining

The pattern frequency distribution theory: a mathematic establishment toward rational and reliable pattern mining

Computing Theoretically-Sound Upper Bounds to Expected Support for Frequent Pattern Mining Problems over Uncertain Big Data

Keywords

References

Author information

Authors and Affiliations

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Share this paper

Publish with us

Search

Navigation