К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Data Engineering and Spark Foundations for AI and ML · LearnSpace
Назад в каталог
courseraАнализ данных

Data Engineering and Spark Foundations for AI and ML

Курс от Edureka
Начальный≈ 6 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

This course builds a strong foundation in modern data engineering using Databricks, Apache Spark, and Lakehouse architecture. You will learn how data engineers design, process, organise, and ingest data to support analytics, AI, and machine learning workflows. You will begin by exploring the role of data engineering in modern organisations, including how it differs from data science and why reliable data pipelines are essential for AI/ML success. You will then get introduced to Databricks, its workspace environment, notebooks, and core platform capabilities. Next, you will work with Apache Spark and PySpark to process data at scale. You will explore DataFrames, transformations, schemas, supported file formats, and practical workflows for reading, transforming, and writing data in formats such as CSV, JSON, and Parquet. You will also examine modern storage concepts, including data lakes, data warehouses, and Lakehouse architecture. Through Databricks, you will learn how to organise data using catalogs, schemas, and volumes, and understand how structured data management supports scalable engineering workflows. By the end of this course, you will be able to: - Explain core data engineering concepts and the role of Databricks. - Describe Lakehouse architecture and modern data storage patterns. - Process and transform data using Apache Spark and PySpark. - Organise project data using catalogs, schemas, and volumes. - Ingest data through batch, API-based, and real-time workflows. Designed for aspiring data engineers, data analysts, software developers, and students entering the data field, this course prepares you to build a strong technical foundation for advanced Databricks, Delta Lake, and AI/ML pipeline development.

Навыки, которые вы освоите

Apache SparkDatabricksData LakesData PipelinesPySparkDistributed ComputingData WarehousingData CollectionApplied Machine LearningJSONMachine LearningPython ProgrammingDatabase Architecture and AdministrationData ManagementData EntryData ArchitectureData Import/ExportArtificial IntelligenceSQLData Processing

Программа курса

3 модулей · 45 учебных материалов

01Introduction to Data Engineering and Databricks19 материалов
Specialization IntroductionВидеоCourse Introduction ВидеоCourse Syllabus: Data Engineering and Spark Foundations for AI and MLЧтениеIntroduction to Data Engineering Видео

Учитесь у экспертов

Edureka

Преподаватель курса

Data Engineering and Spark Foundations for AI and ML
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 6 ч

3 модулей

Язык: Английский

Часть программы вашего университета
Role of Data Engineering in AI/ML PipelinesВидео
Data Engineering vs. Data Science: Understanding the DistinctionsЧтение
Exploring Data Engineering LifecycleВидео
Undercurrents of the Data Engineering LifecycleВидео
Introduction to Databricks Видео
Exploring the Databricks PlatformВидео
Building First Notebook in DatabricksВидео
Industry Use Cases of Databricks for AI/MLЧтение
Revisiting the Fundamentals of Data Engineering and DatabricksЗадание
Data Lake vs Data Warehouse: Foundational ConceptsВидео
Lakehouse Architecture in Modern Data SystemsВидео
Exploring the Lakehouse Structure in DatabricksВидео
Databricks Free Edition GuideЧтение
Introduction to Data Engineering and DatabricksЗадание
Exploring the Foundations of Data Engineering and DatabricksDIALOGUE
02Data Processing with Apache Spark and PySpark10 материалов
Apache Spark: Engine Behind Big Data ProcessingВидеоSpark Core APIs and Execution Concepts ВидеоPySpark DataFrames and Core OperationsВидеоSupported Data Formats in SparkВидеоCore Concepts Check: Apache Spark and PySparkЗаданиеUsing Built-in Datasets in DatabricksВидеоReading CSV and JSON in SparkВидеоWriting Transformed Data in Parquet FormatВидеоData Processing with Apache Spark and PySparkЗаданиеUnderstanding Big Data Processing with Apache Spark and PySparkDIALOGUE
03Organizing and Ingesting Data in the Databricks16 материалов
Evolution of Storage in DatabricksВидеоWorking with Catalogs, Schemas, and VolumesВидеоOrganizing Project Files in Unity Catalog VolumesВидеоIntroduction to Data Ingestion ВидеоData Source Systems in Modern PipelinesВидеоData Ingestion Patterns in Data EngineeringЧтениеData Organization and Ingestion FundamentalsЗаданиеExploring Data Ingestion Options in DatabricksВидеоIngesting Data from a Public REST APIВидеоSimulating Batch Ingestion in DatabricksВидеоReal-Time Ingestion in DatabricksВидеоEnterprise Data Ingestion vs. Databricks Free EditionЧтениеData Engineering Simulation: Designing Modern Data PlatformsDIALOGUEPractice project: Retail Order Pipeline Analytics for UrbanCart RetailЧтениеEnd Course Knowledge Check: Data Engineering and Spark Foundations for AI and MLЗаданиеCourse summaryВидео