Scheduling Strategies and Their Evaluation in a Data Stream Management System
MavStream, a Data Stream Management System (DSMS), has been developed for processing stream data from applications such as network monitoring, sensor monitoring and traffic management systems that require near-real time results and have to process unbounded streams of data. In order to be useful, a result produced by MavStream has to meet certain Quality of Service (QoS) requirements on tuple latency, memory usage, and throughput. Strategies used for scheduling the operators of continuous query (CQ) significantly affect the QoS metrics and hence are of interest. This paper discusses scheduling strategies used in MavStream, their design, implementation, and evaluation. Scheduling is done in MavStream at the operator level. The scheduler maintains a ready queue of operators and decides on the operators to be scheduled based on the scheduling strategy. We first introduce the path capacity scheduling strategy with the goal of minimizing tuple latency by scheduling operator paths with maximum processing capacity. Later we discuss segment-scheduling strategy that aims at minimization of total memory requirement by scheduling operator segments with maximum memory release capacity. We then discuss simplified segment strategy, which splits operator path into just two segments providing better tuple latency performance than segment scheduling strategy and lower memory utilization than path capacity scheduling strategy. Extensive set of experiments have been designed and performed to evaluate the proposed scheduling strategies by simulating real time streams. The performance metrics of average tuple latency, memory utilization and throughput are compared with each other for different strategies and with round robin strategy to validate the analytical conclusions.
KeywordsSchedule Strategy Query Execution Operator Path Query Plan Continuous Query
Unable to display preview. Download preview PDF.
- 1.Gilani, A.: Design and Implementation of Stream Operators, Query Instantiator and Stream Buffer Manager. MS Thesis, CSE Dept. The University of Texas at Arlington (2003) [online], http://www.cse.uta.edu/research/publications/Downloads/CSE-2003-37.pdf
- 2.Motwani, R., Widom, J., Arasu, A., Babcock, B., Babu, S., Datar, M., Manku, G.S., Olston, C., Rosenstein, J., Varma, R.: Query processing, approximation, and resource management in a data stream management system. In: Proc. of CIDR 2003, pp. 245–256 (January 2003)Google Scholar
- 3.Babcock, B., Babu, S., Datar, M., Motwani, R., Widom, J.: Models and issues in data stream systems. In: Proc. of ACM PODS, pp. 1–16 (June 2002)Google Scholar
- 4.Sonune, S.: Design and Implementation of Windowed Operators and Scheduler for Stream Data. MS Thesis CSE Department. The University of Texas at Arlington (2003) [online], http://www.cse.uta.edu/research/publications/Downloads/CSE-2003-38.pdf
- 5.Carney, D., Cetintemel, U., Cherniack, M., Convey, C., Lee, S., Seidman, G., Stonebraker, M., Tatbul, N., Zdonik, S.: Monitoring streams - a new class of data management applications. In: Proc. of the VLDB (2002)Google Scholar
- 6.Cook, D., et al.: MavHome: An Agent-Based Smart Home. In: Proc. of the Conference on Pervasive Computing (2003), http://mavhome.uta.edu
- 8.Carney, D., Cetintemel, U., Rasin, A., Zdonik, S., Cherniack, M., Stonebraker, M.: Operator scheduling in a data stream manager. In: Proc. of the VLDB (2003)Google Scholar
- 9.Babcock, B., Babu, S., Datar, M., Motwani, R.: Chain: Operators Scheduling for Memory Minimization in Stream Systems. In: Proc. of the ACM SIGMOD (2003)Google Scholar
- 10.Viglas, S., Naughton, J.: Rate-based Query Optimization for Streaming Information Sources. In: Proc. of the ACM SIGMOD (2002)Google Scholar
- 11.Avnur, R., Hellerstein, J.: Eddies: Continuously adaptive query processing. In: Proc. of the ACM SIGMOD, pp. 261–272 (200)Google Scholar
- 12.Tatbul, N., Cetintemel, U., Zdonik, S., Cherniack, M., Stonebraker, M.: Load Shedding in a Data Stream Manager. In: Proc. of the VLDB (2003)Google Scholar
- 13.Pajjuri, V.: Design and implementation of scheduling strategies and their evaluation in MavStream, MS Thesis CSE Department. The University of Texas at Arlington (2004) [online], http://itlab.uta.edu/ITLABWEB/Students/sharma/theses/Vamshi.pdf