К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Apache Spark: Design & Execute ETL Pipelines Hands-On · LearnSpace
Назад в каталог
courseraАнализ данных

Apache Spark: Design & Execute ETL Pipelines Hands-On

Курс от EDUCBA
Средний≈ 4.7 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

Build practical data engineering skills by learning how to design, develop, and execute end-to-end ETL (Extract, Transform, Load) pipelines using Apache Spark. In this hands-on course, you will begin by setting up a Spark development environment, installing and configuring PySpark, Hadoop, and MySQL, organizing ETL project structures, and exploring real-world datasets. As you progress, you will implement complete and incremental ETL workflows using Apache Spark. You'll integrate Spark with MySQL through JDBC, apply data transformation logic with Spark SQL, perform business-rule filtering, and address common issues such as data type compatibility and project structure challenges. Through guided, practical exercises, you'll gain experience building scalable ETL workflows in a PySpark environment. This course is designed for aspiring data engineers, big data practitioners, and learners who want practical experience with Apache Spark-based ETL development. By the end of the course, you will be able to construct, execute, and optimize Spark ETL pipelines, implement full and incremental data loading strategies, and integrate Spark applications with relational databases using JDBC for real-world data engineering workflows.

Навыки, которые вы освоите

Apache SparkPySparkData TransformationExtract, Transform, LoadData AnalysisData EngineeringExploratory Data AnalysisSoftware InstallationData PipelinesData ProcessingMySQLMySQL WorkbenchData StoreDevelopment EnvironmentSQLApache HadoopData Import/Export

Программа курса

2 модулей · 20 учебных материалов

01Setting Up the Foundation10 материалов

Getting Started with the ETL Project

Introduction to ProjectВидеоInstallation of PackagesВидеоInstallation of Packages ContinueВидеоGetting Started with the ETL ProjectЗадание

Building the Project Structure and Understanding Data

Учитесь у экспертов

EDUCBA

Преподаватель курса

Apache Spark: Design & Execute ETL Pipelines Hands-On
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 4.7 ч

2 модулей

Язык: Английский

Субтитры: Арабский, Французский, Итальянский, Бразильский португальский, Корейский, Немецкий, Испанский, Японский, Казахский, Венгерский

Часть программы вашего университета
Setting up Project StructureВидео
Exploring DatasetВидео
Building the Project Structure and Understanding DataЗадание
Building Your First Spark ETL Foundation: Environment, Structure & Data UnderstandingDIALOGUE
Graded Quiz - Setting Up the FoundationЗадание
Setting Up a Spark ETL Project Environment and Preparing Data for ProcessingDIALOGUE
02Building ETL Workflows in Apache Spark10 материалов

Complete Load and Transformations

Entire Load and Transformations Part 1ВидеоEntire Load and Transformations Part 2ВидеоEntire Load and Transformations Part 3ВидеоEntire Load and Transformations Part 4ВидеоComplete Load and TransformationsЗадание

Handling Incremental Loads

Incremental LoadВидеоIncremental Load ContinueВидеоHandling Incremental LoadsЗаданиеGraded Quiz – Building ETL Workflows in Apache SparkЗаданиеImplementing an End-to-End Spark ETL Framework from Setup to Incremental LoadsDIALOGUE