This intensive three-day course is dedicated to constructing and fine-tuning high-efficiency data processing workloads within Kubernetes-based environments, leveraging the combined power of PySpark, Pandas, and Polars.
Learners will cultivate a hands-on grasp of how Spark applications function on Kubernetes, exploring how specific configuration choices directly impact performance, scalability, resource utilization, and operational costs. The curriculum delves into critical optimization domains such as executor sizing, memory allocation, dynamic resource allocation, partitioning strategies, shuffle mechanics, the small-file challenge, and the efficient handling of Parquet data.
Additionally, the course tackles typical hurdles encountered when using Pandas, such as memory constraints and out-of-memory errors, while introducing Polars as a robust, high-performance alternative for specific processing tasks. Through practical, hands-on exercises, participants will learn to diagnose performance and memory bottlenecks, evaluate various configuration approaches, and implement optimization techniques in realistic ETL and machine learning contexts.
The core focus of the course remains on practical decision-making: mastering the ability to pinpoint bottlenecks, choose the right tools, configure Spark effectively, and strike a balance between performance and infrastructure resource usage and cost.
Read more...