Team betrachtet Datenauswertungen auf großen Bildschirmen

Projekt

Optimization Techniques for Query Processing in Distributed Big Data Environments

The fast-growing demand for data-intensive applications has led to the wide adoption of distributed big data platforms like Hadoop, Spark, and cloud-native data processing systems. While much has been written on query optimization, including survey literature, many of them lack a coherent analysis, concentrating inste…

The fast-growing demand for data-intensive applications has led to the wide adoption of distributed big data platforms like Hadoop, Spark, and cloud-native data processing systems. While much has been written on query optimization, including survey literature, many of them lack a coherent analysis, concentrating instead on descriptive characteristics and individual techniques. In this paper, we provide a comprehensive review of existing query optimization techniques in distributed big data settings. In this paper, we outline a transparent taxonomy that categorizes optimization strategies into various layers query-level, data-level, execution-level and system level while discussing their interplay and trade-offs. This review goes beyond summarizing well-established approaches such as cost-based optimization and parallel execution, integrating recent trends that include adaptive query optimization, workload-aware partitioning, and resource-aware scheduling in heterogeneous (both cloud-based and not cloud-based) environments. Our focus is on the open research challenges in the areas of dynamic data distribution, edge cloud synergy, and privacy-based query optimization. This paper is the first to review QP metrics, compare it with the various methods enabling efficient application designs with extreme scalability to place the performance measures on equal footing and future research directions in an attempt to devise distributed query processing systems.