Курс от CourseraLearn how to optimize distributed data processing and streaming workflows through this advanced course in the Data Engineering Skill Path. You will build essential competencies including tuning Spark jobs, applying distributed computing frameworks, optimizing data lake storage for query performance, performing stateful stream processing with windowing and watermarks, troubleshooting pipeline failures, and reconciling datasets across multiple systems. Through hands-on practice with Spark, Python, SQL, spreadsheets, and generative AI tools, you will learn to design and refine high-performance data workflows for large-scale analytics. This course brings together expertise from IBM and Google, offering multiple perspectives across distributed processing, warehouse modeling, stream optimization, and advanced feature engineering. You will progress from foundational Spark SQL operations to modeling optimizations, stream processing techniques, performance monitoring, text processing, reconciliation, and AI-assisted performance improvement. The curriculum blends conceptual depth with practical exercises, allowing you to build confidence in diagnosing and tuning distributed workloads. Perfect for emerging data engineers and practitioners who want to strengthen their skills in distributed optimization, streaming systems, and large-scale data performance tuning across modern cloud and open-source environments.
10 модулей · 153 учебных материалов

Преподаватель курса