К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Foundations of Site Reliability Engineering Training · LearnSpace
Назад в каталог
courseraIT и технологии

Foundations of Site Reliability Engineering Training

Курс от Simplilearn
Начальный≈ 24.1 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

This Advanced Site Reliability Engineering Training builds strong expertise in designing, operating, and scaling highly reliable cloud systems using modern SRE and DevOps practices. You learn SLIs, SLOs, SLAs, error budgets, observability, incident management, alerting, RCA, CI CD, chaos engineering, Infrastructure as Code, and performance testing through hands on labs and real world demos using Prometheus, Grafana, Jenkins, Docker, Kubernetes, and Ansible. The course shows how to reduce toil, automate operations, improve resilience, and maintain production ready systems at scale. By the end of this course, you will be able to: - Implement Reliability Metrics: Define SLIs, SLOs, SLAs, and manage error budgets - Build Observability Systems: Configure Prometheus, Grafana, and advanced alerting - Automate Incident Response: Apply RCA, blameless postmortems, and toil reduction - Design Resilient Deployments: Use blue green, canary, and CI CD pipelines - Apply Chaos Engineering: Test system resilience in Kubernetes environments - Optimize Performance at Scale: Conduct load testing and improve reliability Ideal for DevOps engineers, cloud professionals, SRE aspirants, system administrators, and IT practitioners.

Навыки, которые вы освоите

Site Reliability EngineeringPrometheus (Software)Incident ResponseCI/CDInfrastructure as Code (IaC)Performance TestingDocker (Software)KubernetesJenkinsGrafanaAnsibleCloud ComputingIncident ManagementConfiguration ManagementAmazon Elastic Compute CloudProblem ManagementContinuous DeploymentMachine LearningService LevelArtificial Intelligence

Программа курса

7 модулей · 107 учебных материалов

01SRE Foundations13 материалов

What is SRE?

Course SyllabusЧтениеCourse Introduction: Site Reliability Engineering (SRE)ВидеоLearning ObjectivesВидеоIntroduction to Site Reliability Engineering (SRE)Видео

Учитесь у экспертов

Priyanka Mehta

Преподаватель курса

Foundations of Site Reliability Engineering Training
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 24.1 ч

7 модулей

Язык: Английский

Часть программы вашего университета
Core Concepts in SREВидео
Quiz on What is SRE?Задание

Reliability Metrics

Demo: Creating an EC2 InstanceВидеоDemo: Creating SLIs, SLOs, and SLAs for a Sample ServiceВидеоUnderstanding Error Budgets: Concepts and BenefitsВидеоApplying Error Budgets: Examples and Advanced PracticesВидеоShip Fast or Stay Stable? The SRE Reliability DilemmaDIALOGUEQuiz on Reliability MetricsЗаданиеAssessment for SRE FoundationsЗадание
02Error Budgets & Observability11 материалов

Error Budgets in Practice

Demo: Calculating and Simulating Error BudgetВидеоMonitoring and Observability​ВидеоOverview of Alert FatigueВидеоCorrelating Observability DataВидеоQuiz on Error Budgets in PracticeЗадание

Modern Observability

AI/ML in ObservabilityВидеоDemo: Setting up Prometheus and Grafana for Monitoring - Part 1ВидеоDemo: Setting up Prometheus and Grafana for Monitoring - Part 2ВидеоRunning Out of Error Budget: What Should You Fix First?DIALOGUEQuiz on Modern ObservabilityЗаданиеAssessment for Error Budgets & ObservabilityЗадание
03Incident Management & Toil Reduction15 материалов

Incident Response Fundamentals

Incident ManagementВидеоBlameless PostmortemВидеоOverview and Types of Incident CommunicationВидеоMetrics and Automation in Incident ResponseВидеоQuiz on Incident Response FundamentalsЗадание

Incident Automation & Toil

Demo: Implementing Incident Management with Prometheus - Part 1ВидеоDemo: Implementing Incident Management with Prometheus - Part 2ВидеоToil ReductionВидеоDemo: Implementing Toil Reduction with Automated Service Recovery Using Shell Script - Part 1ВидеоDemo: Implementing Toil Reduction with Automated Service Recovery Using Shell Script - Part 2ВидеоSRE CultureВидеоKey TakeawaysВидеоThe 2 AM Incident: Fix Fast or Fix ForeverDIALOGUEQuiz on Incident Automation & ToilЗаданиеAssessment for Incident Management & Toil ReductionЗадание
04Reliability Engineering & Deployments14 материалов

Reliability Engineering Basics

Learning ObjectivesВидеоIntroduction to Reliability EngineeringВидеоDeployment Strategies in Reliability EngineeringВидеоDemo: Implementing Site Reliability Engineering (SRE) with Blue-Green and Canary DeploymentВидеоQuiz on Reliability Engineering BasicsЗадание

SRE Automation Foundations

Introduction to SRE AutomationВидеоInfrastructure as Code (IaC): Concepts, Benefits, Tools, and Best PracticesВидеоConfiguration Management in SRE: Concepts, Practices, and BenefitsВидеоSRE Automation: Key Areas and TypesВидеоSRE Automation: Pipelines, Monitoring, Scaling, and Incident ResponseВидеоDemo: Automating SRE with Ansible and HTTPS NginxВидеоDeploy Without Downtime: Choosing the Safest Release StrategyDIALOGUEQuiz on SRE Automation FoundationsЗаданиеAssessment for Reliability Engineering & DeploymentsЗадание
05Alerting, Automation & RCA21 материалов

Alert Design and Implementation

Principles of Good AlertingВидеоManaging Alert Fatigue: Actionable Alerts and Prioritization FrameworkВидеоCommon Alerting ToolsВидеоDesigning Effective Alerts: Multi-Level and SLO-Based AlertingВидеоDemo: Monitoring EC2 Instance and Alerting Strategy with Prometheus, Node Exporter, and Alertmanager - Part 1ВидеоDemo: Monitoring EC2 Instance and Alerting Strategy with Prometheus, Node Exporter, and Alertmanager - Part 2ВидеоQuiz on Alert Design and ImplementationЗадание

RCA & Postmortems

Incident Response: Process, Escalation Paths, and the Incident Commander RoleВидеоRoot Cause Analysis (RCA) and Its Importance in SREВидеоRoot Cause Analysis in SRE: Techniques and ImplementationВидеоEffective Postmortems: Blameless Practices and Continuous ImprovementВидеоDemo: Setting Up System Monitoring, Incident Alerts, and Response with Prometheus and Alertmanager - Part 1ВидеоDemo: Setting Up System Monitoring, Incident Alerts, and Response with Prometheus and Alertmanager - Part 2Видео
06CI/CD & Chaos Engineering16 материалов

CI/CD for SRE

Learning ObjectivesВидеоCI/CD Fundamentals for SREВидеоOperationalizing CI/CD for SRE TeamsВидеоCI/CD Tooling and Automation for SRE TeamsВидеоDemo: Setting up CI/CD Pipeline with Jenkins and Docker - Part 1ВидеоDemo: Setting up CI/CD Pipeline with Jenkins and Docker - Part 2ВидеоDemo: Setting up CI/CD Pipeline with Jenkins and Docker - Part 3ВидеоQuiz on CI/CD for SREЗадание

Chaos Engineering

Choas Engineering FundamentalsВидеоChaos Engineering PracticesВидеоChaos Engineering in Kubernetes and Use CasesВидеоDemo: Implementing Chaos Engineering with Pumba​ - Part 1ВидеоDemo: Implementing Chaos Engineering with Pumba​ - Part 2ВидеоDeploy, Disrupt, Repeat: Engineering Resilience with CI/CD and ChaosDIALOGUE
07Performance Testing & Advanced SRE17 материалов

Performance Engineering

Introduction to Performance TestingВидеоRealistic Load ProfilesВидеоPerformance Testing in CI/CDВидеоDemo: Multi-User Load Testing with Chaos - Part 1ВидеоDemo: Multi-User Load Testing with Chaos - Part 2ВидеоQuiz on Performance EngineeringЗадание

SRE at scale

SRE Fundamentals: Core Principles and Supporting PracticesВидеоImplementing SRE: Workflow, Team Structure, Tools, and MetricsВидеоImplementing Error Budgets and Building a Learning CultureВидеоUse Case: Integrated SRE approachВидеоSRE Implementation: Challenges, Strategies, and Future TrendsВидеоDemo: Implementing Container Restart Detection and Alerting with Docker - Part 1Видео
Demo: Setting Up System Monitoring Incident Alerts and Response with Prometheus and Alertmanager - Part 3Видео
SRE ReliabilityВидео
Managing Reliability with Error BudgetsВидео
Measuring and Improving ReliabilityВидео
Key TakeawaysВидео
Too Many Alerts, One Real Problem: Can You Find the Root Cause?DIALOGUE
Quiz on RCA & PostmortemsЗадание
Assessment for Alerting, Automation & RCAЗадание
Quiz on Chaos EngineeringЗадание
Assessment for CI/CD & Chaos EngineeringЗадание
Demo: Implementing Container Restart Detection and Alerting with Docker - Part 2Видео
Key TakeawaysВидео
Can Your System Handle the Surge? Scaling Reliability Under PressureDIALOGUE
Quiz on SRE at scaleЗадание
Assessment for Performance Testing & Advanced SREЗадание