Курс от CourseraLearn how to evaluate, validate, and monitor data quality across multiple systems in this comprehensive course within the Data Engineering Skill Path. You will develop essential competencies, including executing predefined quality checks, profiling datasets using descriptive statistics, identifying common data quality issues, reconciling data across systems, and validating warehouse data for completeness and accuracy. Through hands-on practice using Google ETL frameworks, Meta’s statistical methods, and IBM’s Python and warehouse tools, you will gain the skills needed to ensure trustworthy, reliable datasets for downstream analytics. This course combines expertise from Google, Meta, and IBM, offering multiple perspectives on data quality management across pipelines, statistical workflows, and enterprise data warehouses. You will progress from detecting issues in ETL processes, to profiling and summarizing datasets, to programmatically auditing data with Python, and finally to applying structural and business-rule validations in a warehouse environment. The curriculum balances conceptual understanding with practical exercises, preparing you to confidently assess and improve data quality throughout the data lifecycle. Perfect for aspiring data engineers, data analysts, and learners seeking strong foundational skills in data auditing, profiling, and validation across modern data environments.
6 модулей · 85 учебных материалов

Преподаватель курса