К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Data Engineering with Databricks Cookbook · LearnSpace
Назад в каталог
courseraАнализ данных

Data Engineering with Databricks Cookbook

Курс от Packt
Средний≈ 9.8 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

This course offers a hands-on approach to mastering data engineering using Apache Spark, Delta Lake, and Databricks. By combining these technologies, you will learn how to build robust, scalable data pipelines and implement effective data management strategies in real-world applications. With a focus on performance optimization, data orchestration, and modern data engineering practices, this course provides essential skills for professionals working in the data engineering space. You’ll start by exploring data ingestion techniques using Apache Spark, followed by methods for transforming and managing data within a data lakehouse. Each section builds on the last, providing learners with actionable insights that can be directly applied to their workflows. The course also covers DataOps and DevOps practices to help you streamline and automate your data processes. What sets this course apart is its emphasis on practical, real-world applications. You’ll work through concrete examples and recipes for managing data, from ingestion to transformation, ensuring that you can tackle data engineering challenges with confidence. Ideal for data engineers, data scientists, and IT professionals with a background in SQL and Python, this course will help you enhance your skills in data pipeline orchestration and optimization.

Навыки, которые вы освоите

DevOpsDevops ToolsData ManipulationData TransformationData LakesData GovernanceData AccessData ProcessingApachePerformance TuningPySparkData CaptureApache SparkData EngineeringDatabricksReal Time DataData Pipelines

Программа курса

11 модулей · 92 учебных материалов

01Data Ingestion and Data Extraction with Apache Spark10 материалов

Mastering Data Loading and Transformation in Apache Spark

OverviewВидеоIntroductionЧтениеCommon Issues Faced While Working with CSV DataЧтениеReading JSON Data with Apache SparkЧтение

Учитесь у экспертов

Packt - Course Instructors

Преподаватель курса

Data Engineering with Databricks Cookbook
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 9.8 ч

11 модулей

Язык: Английский

Часть программы вашего университета
The Flatten() and Collect_list() FunctionsЧтение
Parsing XML Data with Apache SparkЧтение
Working with Nested Data Structures in Apache SparkЧтение
The Map Keys and Map Values FunctionsЧтение
Using the regexp_extract() FunctionЧтение
Data Ingestion and Extraction with Apache SparkЗадание
02Data Transformation and Data Manipulation with Apache Spark9 материалов

Mastering Data Operations and Advanced Analytics with PySpark

OverviewВидеоIntroductionЧтениеFiltering Data with Apache SparkЧтениеPerforming Joins with Apache SparkЧтениеPerforming Aggregations with Apache SparkЧтениеApproximate AggregationsЧтениеNested Window FunctionsЧтениеHandling Null Values with Apache SparkЧтениеMastering Data Processing in Apache SparkЗадание
03Data Management with Delta Lake8 материалов

Mastering Data Integrity and Performance with Delta Lake

OverviewВидеоIntroductionЧтениеReading a Delta Lake TableЧтениеMerging Data into Delta TablesЧтениеChange Data Capture in Delta LakeЧтениеOptimizing Delta Lake TablesЧтениеVersioning and Time Travel for Delta Lake TablesЧтениеMastering Delta Lake Data ManagementЗадание
04Ingesting Streaming Data8 материалов

Mastering Real-Time Data Pipelines with Spark Structured Streaming

OverviewВидеоIntroductionЧтениеReading Data from Real-Time Sources, Such as Apache Kafka, with Apache Spark Structured StreamingЧтениеDefining Transformations and Filters on a Streaming DataFrameЧтениеConfiguring Checkpoints for Structured Streaming in Apache SparkЧтениеConfiguring Triggers for Structured Streaming in Apache SparkЧтениеApplying Window Aggregations to Streaming Data with Apache Spark Structured StreamingЧтениеExploring Streaming Data Processing with Apache SparkЗадание
05Processing Streaming Data8 материалов

Mastering Real-Time Data Integration and Monitoring with Spark Streaming

OverviewВидеоIntroductionЧтениеIdempotent Stream Writing with Delta Lake and Apache Spark Structured StreamingЧтениеMerging or Applying Change Data Capture on Apache Spark Structured Streaming and Delta LakeЧтениеJoining Streaming Data with Static Data in Apache Spark Structured Streaming and Delta LakeЧтениеJoining Streaming Data with Streaming Data in Apache Spark Structured Streaming and Delta LakeЧтениеMonitoring Real-Time Data Processing with Apache Spark Structured StreamingЧтениеStreaming Data Processing FundamentalsЗадание
06Performance Tuning with Apache Spark9 материалов

Mastering Spark Performance: Strategies for Efficient Data Processing

OverviewВидеоIntroductionЧтениеUsing Broadcast VariablesЧтениеOptimizing Spark Jobs by Minimizing Data ShufflingЧтениеAvoiding Data SkewЧтениеCaching and PersistenceЧтениеPartitioning and RepartitioningЧтениеOptimizing Join StrategiesЧтениеMastering Spark Performance TuningЗадание
07Performance Tuning in Delta Lake6 материалов

Mastering Delta Lake Optimization: Partitioning, Z-ordering, and Compression

OverviewВидеоIntroductionЧтениеOrganizing Data with Z-ordering for Efficient Query ExecutionЧтениеSkipping Data for Faster Query ExecutionЧтениеReducing Delta Lake Table Size and I/O Cost with CompressionЧтениеPerformance Tuning in Delta LakeЗадание
08Orchestration and Scheduling Data Pipeline with Databricks Workflows7 материалов

Mastering Automated Data Pipelines with Databricks Workflows

OverviewВидеоIntroductionЧтениеRunning and Managing Databricks WorkflowsЧтениеPassing Task and Job Parameters Within a Databricks WorkflowЧтениеConditional Branching in Databricks WorkflowsЧтениеTriggering Jobs Based on File ArrivalЧтениеMastering Databricks Workflow OrchestrationЗадание
09Building Data Pipelines with Delta Live Tables9 материалов

Mastering Real-Time Data Pipelines with Delta Live Tables

OverviewВидеоIntroductionЧтениеBuilding a Data Pipeline with Delta Live Tables on DatabricksЧтениеImplementing Data Quality and Validation Rules with Delta Live Tables in DatabricksЧтениеQuarantining Bad Data with Delta Live Tables in DatabricksЧтениеMonitoring Delta Live Tables PipelinesЧтениеDeploying Delta Live Tables Pipelines with Databricks Asset BundlesЧтениеApplying Changes (CDC) to Delta Tables with Delta Live TablesЧтениеData Pipeline Fundamentals with Delta Live TablesЗадание
10Data Governance with Unity Catalog11 материалов

Mastering Secure Data Management with Unity Catalog

OverviewВидеоIntroductionЧтениеCreating a CatalogЧтениеDefining and Applying Fine-Grained Access Control Policies Using Unity CatalogЧтениеTagging, Commenting, and Capturing Metadata About Data and AI Assets Using Databricks Unity CatalogЧтениеUsing the Unity Catalog UIЧтениеApply Row FiltersЧтениеApply Column MasksЧтениеUsing Unity Catalogs Lineage Data for Debugging, Root Cause Analysis, and Impact AssessmentЧтениеAccessing and Querying System Tables Using Unity CatalogЧтениеData Governance with Unity CatalogЗадание
11Implementing DataOps and DevOps on Databricks7 материалов

Automating Data Workflows and CI/CD Pipelines on Databricks

OverviewВидеоIntroductionЧтениеAutomating Tasks by Using the Databricks CLIЧтениеUsing the Databricks VSCode Extension for Local Development and TestingЧтениеUsing Databricks Asset Bundles (DABs)ЧтениеLeveraging GitHub Actions with Databricks Asset Bundles (DABs)ЧтениеDataOps and DevOps Implementation on DatabricksЗадание