К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Delta Lake and Medallion Architecture for AI and ML · LearnSpace
Назад в каталог
courseraАнализ данных

Delta Lake and Medallion Architecture for AI and ML

Курс от Edureka
Средний≈ 6.6 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

This course expands your data engineering skills by focusing on Delta Lake and Medallion Architecture for building dependable, scalable, and analytics-ready data platforms. You will learn how modern pipelines maintain data quality, support version control, and prepare trusted datasets for AI/ML workloads. You will start with Delta Lake fundamentals, including ACID transactions, schema enforcement, schema evolution, transaction logs, and Time Travel. Through practical demonstrations, you will create Delta tables, manage updates and deletes, apply MERGE operations, and restore earlier versions of data when needed. You will then apply PySpark transformation techniques to clean, reshape, join, and validate datasets. You will also explore performance-focused practices such as OPTIMIZE, VACUUM, and Z-Ordering to make Delta tables more efficient for large-scale processing. Next, you will design Medallion Architecture pipelines using Bronze, Silver, and Gold layers. This helps convert raw data into clean, validated, and business-ready datasets for reporting, analytics, and machine learning. The course also introduces structured streaming, Change Data Feed, and Delta constraints for improving pipeline reliability. By the end of this course, you will be able to: - Use Delta Lake for reliable and versioned data management. - Transform and validate datasets using PySpark. - Optimise Delta tables for performance and storage efficiency. - Build Bronze, Silver, and Gold data layers. - Develop dependable batch and streaming pipelines for AI/ML use cases. Designed for data engineers, analytics engineers, software developers, and AI/ML professionals, this course prepares you to build modern data pipelines that are reliable, scalable, and ready for production workloads.

Навыки, которые вы освоите

Data ValidationData PipelinesData TransformationData CleansingData ArchitectureData ManipulationPySparkPython ProgrammingData EntryApache SparkExtract, Transform, LoadData CollectionSQLMLOps (Machine Learning Operations)Data LakesData InfrastructureData ProcessingData MaintenanceDatabricksData Management

Программа курса

4 модулей · 49 учебных материалов

01Delta Lake Fundamentals 17 материалов
Specialization IntroductionВидеоCourse IntroductionВидеоCourse syllabus: Delta Lake and Medallion Architecture for Modern Data EngineeringЧтениеIntroduction to Delta Lake on DatabricksВидео

Учитесь у экспертов

Edureka

Преподаватель курса

Delta Lake and Medallion Architecture for AI and ML
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 6.6 ч

4 модулей

Язык: Английский

Часть программы вашего университета
ACID Transactions in Delta Lake Видео
Creating and Exploring Delta TablesВидео
Tracking Changes with the Delta Transaction LogВидео
Understanding the Delta Transaction Log Location in DatabricksЧтение
Schema Enforcement and Schema Evolution in Delta LakeВидео
Delta Lake FundamentalsЗадание
Performing CRUD Operations on Delta TablesВидео
Managing Inserts and Updates with MERGE INTOВидео
Time Travel and Versioned Data AccessВидео
Using Delta Lake Time Travel for Data RecoveryВидео
Delta Lake Table Properties and Configuration OptionsЧтение
Delta Lake Fundamentals Задание
Understanding Reliable Data Management with Delta LakeDIALOGUE
02Data Transformation with Pyspark 12 материалов
PySpark Transformations for Data EngineeringВидеоFiltering, Selecting, and Renaming ColumnsВидеоAggregations and GroupBy OperationsВидеоJoining Datasets and Handling DuplicatesВидеоPySpark Data TransformationsЗаданиеHandling Nulls and Applying Data Quality ChecksВидеоTransforming Text Data for AI ApplicationsВидеоWorking with Complex Types: Arrays, Structs, and MapsЧтениеDelta Table Optimization: OPTIMIZE, VACUUM, and Z-OrderВидеоOptimizing Delta Tables for Query PerformanceВидеоData Transformation with Pyspark ЗаданиеMastering Data Transformation and Optimization with PySparkDIALOGUE
03Medallion Architecture for AI/ML 12 материалов
Medallion Architecture: The Foundation of Modern Data EngineeringВидеоThe Three Layers of the Medallion ArchitectureВидеоDesigning a Medallion Pipeline for an AI/ML Use CaseЧтениеSetting Up the Medallion WorkspaceВидеоMedallion Layers and Data RefinementЗаданиеIngesting Raw Data into the Bronze LayerВидеоBuilding the Silver Layer: Cleaning and Validating DataВидеоBuilding the Gold Layer: Aggregated and Business-Ready DataВидеоMedallion Architecture as a Best Practice for ML Data Pipelines ЧтениеWhy ML Data Pipelines Extend Naturally to AI WorkloadsЧтениеMedallion Architecture for AI/ML ЗаданиеDesigning Scalable AI Pipelines with Medallion ArchitectureDIALOGUE
04Streaming and Pipeline reliability8 материалов
Introduction to Structured Streaming with Delta LakeВидеоDelta Change Data Feed for Incremental ProcessingВидеоStreaming Constraints in Databricks Free EditionЧтениеEnforcing Data Quality with Delta ConstraintsВидеоEnterprise Data Engineering Simulation: Reliable Lakehouse PipelinesDIALOGUEPractice project: Building a Reliable Delta Lake Pipeline for MediPlus ClinicsЧтениеEnd Course Knowledge Check: Delta Lake & Medallion ArchitectureЗаданиеCourse SummaryВидео