With modern enterprises two to five times larger than they were 20 or even ten
years ago, producing enormous volumes of structured, semi-structured and unstructured data that needs more sophisticated processing and analysis so as to exploit this data in a way strategic decision making can be made, the demand for more capable storage and analytics also is growing rapidly. While traditional data processing platforms have limitations in their ability to scale, resource management, workload isolation and operational flexibility when it comes to large-scale analytics workloads. This paper describes the design, implementation and execution on one such a scalable analytics framework based on Apache Spark and Kubernetes to process enterprise data. The pipeline built in this architecture integrates the distributed processing capability of Apache Spark along with the orchestration and resource management capabilities of
Kubernetes to achieve a cloud-native analytics environment that supports variable workloads. It supports elastic resource allocation, automatic deployment and fault tolerance, along with rapid execution of complex analytical queries. We perform a thorough empirical analysis over wide-ranging enterprise datasets of 10+ million records to measure execution time, throughput, resource utilization and scalability. Results show
that the integrated Spark-Kubernetes architecture achieves better processing efficiency, workload manageability, and fault resilience in contrast to conventional deployment methods. The results confirm that cloud-native data analytics systems are successful at resolving modern enterprise computing requirements and provide rich insights to organizations in need of flexible and economical big data per processing platforms
Keywords : Apache Spark, Kubernetes, Enterprise Data Processing, Cloud-Native Analytics, Big Data Analytics, Distributed Data Processing, Resource Management, Scalability, Performance Evaluation, Container Orchestration, Data Engineering, RealTime Analytics, Cloud Computing, Enterprise Intelligence, Elastic Computing.
Author : Processing Naga Hemanth Badabagni Staff Data Engineer | AI Data Platform Architect | Enterprise Data Engineering Leader
Title : Design and Performance Evaluation of Scalable Analytics Systems Using Apache Spark and Kubernetes for Enterprise Data
Volume/Issue : 2023;05(02)
Page No : 31-56