К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Building Resilient Systems · LearnSpace
Назад в каталог
courseraIT и технологии

Building Resilient Systems

Курс от STARWEAVER
Средний≈ 11.1 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

Building resilient systems requires more than knowing individual tools. It demands the ability to design architectures that anticipate failure and recover effectively. In this intermediate course, you will learn how to apply resilience engineering principles to modern distributed systems, focusing on high availability architecture, fault tolerance, and disaster recovery planning. You will analyze how and why systems fail, identify hidden risks in resilient system design, and build strategies that improve uptime and system reliability. The course connects key concepts such as load balancing, redundancy, observability, and incident response into a cohesive system resilience strategy aligned with business goals like RTO and RPO. Designed for IT professionals, DevOps engineers, and system architects, this course emphasizes practical decision-making, trade-offs, and operational readiness. By the end, you will be able to design resilient systems architectures, strengthen system reliability, and lead effective incident management and continuous improvement practices.

Навыки, которые вы освоите

Incident ManagementContinuous MonitoringIncident ResponseDisaster RecoverySystem MonitoringDistributed ComputingService LevelLoad BalancingSystem Design and ImplementationSite Reliability EngineeringBusiness ContinuitySystems DesignSystem ImplementationSystems ArchitectureNetwork Monitoring

Программа курса

4 модулей · 68 учебных материалов

01Foundations of Resilient Systems18 материалов

Lesson 1: Understanding System Failures

Foundations of Resilient Systems DIALOGUEWelcome to Building Resilient SystemsВидеоWelcome to the Course: Course OverviewЧтениеModule Introduction Видео

Учитесь у экспертов

Ahmed Elhenedy

Network instructor

Starweaver

Global Leaders in Professional & Technology Education

Building Resilient Systems
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 11.1 ч

4 модулей

Язык: Английский

Часть программы вашего университета
Why Systems Fail Видео
Failure Types and Their Impact Видео
Learning from Real-World Outages Видео

Lesson 2: What Makes a System Resilient

Defining Resilience in Modern Systems ВидеоKey Characteristics of Resilient Architectures ВидеоResilience v/s. Traditional Design Approaches Видео

Lesson 3: Core Principles of Resilience Engineering

Core Principles of Resilience EngineeringВидеоRedundancy, Diversity, and Isolation ВидеоTrade-offs in Resilient Design ВидеоDesigning Resilient Systems ЧтениеDesigning for Failure Before It HappensОбсуждениеHands-On-Learning: Identifying Failure Risks in a System Design Взаимная проверкаPost-Outage Architecture Review Meeting DIALOGUEFoundations of Resilient SystemsЗадание
02High Availability and Fault Tolerance Design16 материалов

Lesson 1: High Availability Fundamentals

High Availability and Fault Tolerance Design DIALOGUEModule Introduction ВидеоWhat High Availability Really MeansВидеоAvailability Metrics and SLAs ВидеоEliminating Single Points of Failure Видео

Lesson 2: Redundancy and Failover Strategies

Active-Active v/s. Active-Passive Designs ВидеоLoad Balancing and Traffic Distribution ВидеоFailover Mechanisms and Health Checks Видео

Lesson 3: Fault Tolerance Patterns

Designing for Partial Failures ВидеоGraceful Degradation and Backpressure ВидеоContaining Failures with Isolation ВидеоHigh Availability and Fault-Tolerant Architecture ЧтениеBalancing Availability, Cost, and ComplexityОбсуждениеHands-On-Learning: Designing a High Availability Architecture Взаимная проверка
03Disaster Recovery Planning and Operational Readiness16 материалов

Lesson 1: Backup and Recovery Planning

Disaster Recovery Planning and Operational Readiness DIALOGUEModule Introduction ВидеоBackup Strategies and Recovery Models ВидеоUnderstanding RTO and RPO ВидеоDesigning Backup and Recovery Solutions Видео

Lesson 2: Disaster Recovery Testing

Why Disaster Recovery Testing Matters ВидеоTypes of Disaster Recovery Tests ВидеоDeveloping Disaster Recovery Testing Procedures Видео

Lesson 3: Runbooks and Operational Readiness

What is an Operational Runbook ВидеоRunbook Structure and Best Practices ВидеоCreating Effective Recovery Runbooks ВидеоDisaster Recovery Planning and RTO or RPO Concepts ЧтениеEvaluating Recovery Readiness in Real-World EnvironmentsОбсуждениеHands-On-Learning: Creating a Disaster Recovery PlanВзаимная проверка
04Monitoring, Observability, and Incident Management18 материалов

Lesson 1: Observability Foundations

Monitoring, Observability, and Incident Management DIALOGUEModule Introduction ВидеоMonitoring v/s. Observability ВидеоObservability Pillars: Logs, Metrics, and Traces ВидеоImplementing Comprehensive Observability Видео

Lesson 2: Alerting and Escalation Strategies

Principles of Effective Alerting ВидеоAlert Thresholds and Escalation Paths ВидеоDesigning Effective Alerting Strategies Видео

Lesson 3: Incident Review and Continuous Improvement

Incident Lifecycle and Response Review ВидеоConducting Productive Post-Incident Reviews ВидеоDriving Continuous Improvement from Incidents ВидеоObservability and Incident Management Fundamentals ЧтениеDesigning Observability and Alerting for Real ImpactОбсуждениеHands-On-Learning: Incident Analysis and Post-Incident Review Взаимная проверка
Designing a High Availability Strategy Under SLA Pressure DIALOGUE
High Availability and Fault Tolerance DesignЗадание
Presenting a Comprehensive Disaster Recovery Strategy to LeadershipDIALOGUE
Disaster Recovery Planning and Operational ReadinessЗадание
Leading a Blameless Post-Incident Review After a Major Outage DIALOGUE
Monitoring, Observability, and Incident ManagementЗадание
Project: Designing and Defending a Resilient System ArchitectureВзаимная проверка
Course Wrap-UpВидео