Introducing Skew into the TPC-H Benchmark
While uniform data distributions were a design choice for the TPC-D benchmark and its successor TPC-H, it has been universally recognized that data skew is prevalent in data warehousing. A modern benchmark should therefore provide a test bed to evaluate the ability of database engines to handle skew. This paper introduces a concrete and practical way to introduce skew in the TPC-H data model by modifying the customer and supplier tables to reflect non-uniform customer and supplier populations. The first proposal consists in defining customer and supplier populations by nation that are roughly proportional to the actual nation populations. In a second proposal, nations are divided into two groups, one with large and equal populations and the other with equal and small populations. We then experiment with the proposed skew models to show how the optimizer of a parallel system can recognize skew and potentially produce different plans depending on the presence of skew. A comparison is made between query performance with the proposed method vs. the original uniform TPC-H distributions. Finally, an approach is presented to introduce skew into TPC-H with the current query set that is compatible with the current benchmark specification rules and could be implemented today.
Unable to display preview. Download preview PDF.
- 1.Lakshmi, S.M., Yu, P.S.: Effect of Skew on Join Performance in Parallel Architectures. In: International Symposium on Databases in Parallel and Distributed Systems (1988)Google Scholar
- 2.Walton, C.B., Dale, A.G., Jenevein, R.M.: A Taxonomy and Performance Model of Data Skew Effects in Parallel Joins. In: Proceedings of VLDB, pp. 537–548 (1991)Google Scholar
- 3.Wolf, J.L., Dias, D.M., Yu, P.S., Turek, J.: An Effective Algorithm for Parallelizing Hash Joins in the Presence of Data Skew. In: Proceedings of ICDE 1991 (1991)Google Scholar
- 4.DeWitt, D.J., Naughton, J.F., Schneider, D.A., Seshadri, S.: Practical Skew Handling in Parallel Joins. In: Proceedings of VLDB 1992, pp. 27–40 (1992)Google Scholar
- 5.Xu, Y., Kostamaa, P.: Efficient Outer Join Data Skew Handling in Parallel DBMS. In: Proceedings of VLDB, pp. 1390–1396 (2009)Google Scholar
- 6.TPC Benchmark H (Decision Support) Standard Specification Revision 2.14.0, www.tpc.org