Scheduling of Intermittent Query Processing

التفاصيل البيبلوغرافية
العنوان: Scheduling of Intermittent Query Processing
المؤلفون: Chandrasekaran, Saranya, Sudarshan, S.
سنة النشر: 2023
المجموعة: Computer Science
مصطلحات موضوعية: Computer Science - Databases
الوصف: Stream processing is usually done either on a tuple-by-tuple basis or in micro-batches. There are many applications where tuples over a predefined duration/window must be processed within certain deadlines. Processing such queries using stream processing engines can be very inefficient since there is often a significant overhead per tuple or micro-batch. The cost of computation can be significantly reduced by using the wider window available for computation. In this work, we present scheduling schemes where the overhead cost is minimized while meeting the query deadline constraints. For such queries, since the result is needed only at the deadline, tuples can be processed in larger batches, instead of using micro-batches. We present scheduling schemes for single and multi query scenarios. The proposed scheduling algorithms have been implemented as a Custom Query Scheduler, on top of Apache Spark. Our performance study with TPC-H data, under single and multi query modes, shows orders of magnitude improvement as compared to naively using Spark streaming.
نوع الوثيقة: Working Paper
URL الوصول: http://arxiv.org/abs/2306.06678
رقم الأكسشن: edsarx.2306.06678
قاعدة البيانات: arXiv