К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Site Reliability Engineering (SRE) Principles · LearnSpace
Назад в каталог
courseraПрограммирование

Site Reliability Engineering (SRE) Principles

Курс от Edureka
Начальный≈ 9.7 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

This course equips you with practical Site Reliability Engineering (SRE) skills for modern cloud-native and DevOps environments. You will begin with SRE fundamentals, including reliability principles, the relationship between SRE and DevOps, and key reliability metrics such as SLIs, SLOs, and error budgets. You will then explore observability and operations using Prometheus, Grafana, and Argo CD for monitoring, alerting, dashboards, GitOps deployments, incident management, on-call practices, and blameless postmortems. The course concludes with SRE automation and recovery, covering runbooks, Ansible playbooks, Pyrra, burn-rate alerts, GitOps-based rollbacks, and anomaly detection. By the end of the course, you will be able to define and implement reliability objectives, build monitoring and SLO dashboards, configure effective alerts, manage incidents and postmortems, automate operational tasks, track error budgets, and apply recovery strategies using GitOps workflows. Designed for DevOps engineers, SREs, platform engineers, cloud engineers, Kubernetes administrators, and operations teams, this course requires a basic understanding of Linux, Git, YAML, and Kubernetes fundamentals. Enroll today and take the next step toward becoming a skilled Site Reliability Engineer capable of building resilient, observable, and highly automated cloud-native systems that scale with confidence.

Навыки, которые вы освоите

AutomationGit (Version Control System)System MonitoringDashboard CreationIT AutomationPrometheus (Software)Service LevelContinuous MonitoringAnsibleDevOpsRelease ManagementCloud-Native ComputingInfrastructure as Code (IaC)Incident ResponseAI literacyAnomaly DetectionKubernetes

Программа курса

4 модулей · 63 учебных материалов

01Foundations of Site Reliability Engineering18 материалов

Introduction to SRE and Reliability Thinking

Course IntroductionВидеоCourse SyllabusЧтениеWhat is Site Reliability Engineering?ВидеоSRE vs DevOps: Operational Alignment for Reliable DeliveryВидео

Учитесь у экспертов

Edureka

Преподаватель курса

Site Reliability Engineering (SRE) Principles
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 9.7 ч

4 модулей

Язык: Английский

Часть программы вашего университета
Reliability, Availability, and Resilience in Modern SystemsВидео
Developing a Reliability Mindset in SREЧтение
Knowledge Check: Introduction to SRE and Reliability ThinkingЗадание

Service Health Metrics and Reliability Targets

Service Level Indicators and Service Level ObjectivesВидеоError Budgets and Budget PoliciesВидеоHands-On: Defining SLIs and SLOs for a Sample Application - Sample Application Setup and Initial SLI MeasurementВидеоHands-On: Defining SLIs and SLOs for a Sample Application - Defining SLOs and Improving Reliability PerformanceВидеоHands-On: Calculating an Error Budget from an SLO - Error Budget CalculationВидеоHands-On: Calculating an Error Budget from an SLO - Budget Usage and Release DecisionВидеоDesigning Service Health Metrics and Reliability TargetsЧтениеKnowledge Check: Service Health Metrics and Reliability Targets Задание

Module Wrap-Up and Assessment

Foundation of Site Reliability EngineeringDIALOGUEKnowledge Check: Foundations of Site Reliability EngineeringЗаданиеModule Summary: Foundations of Site Reliability Engineering Чтение
02Monitoring, Alerting, and Incident Operations21 материалов

Observability and GitOps Based Application Deployment

Observability Fundamentals for Reliable SystemsВидеоMetrics, Logs, and Traces in SREВидеоHands-On: Setting Up the SRE Lab EnvironmentВидеоConnecting Observability with GitOps-Based DeploymentЧтениеKnowledge Check: Observability and GitOps Based Application DeploymentЗадание

Dashboards, Alerts, and SLO Monitoring

Building Effective SRE DashboardsВидеоAlerting Principles and Reducing Alert FatigueВидеоHands-On: Installing Prometheus and Grafana - Monitoring Stack InstallationВидеоHands-On: Installing Prometheus and Grafana - Prometheus and Grafana VerificationВидеоHands-On: Building an SLO Dashboard in GrafanaВидеоTurning Monitoring Data into ReliabilityЧтениеKnowledge Check: Dashboards, Alerts, and SLO MonitoringЗадание

Incident Management and On Call Practices

Managing Incidents for Reliable ServicesВидеоOn-Call Best Practices and Escalations Workflows ВидеоHands-On: Managing On - Call Escalations with GoAlertВидеоHands-On: Writing a Blameless Incident PostmortemВидеоBuilding Incident Readiness and Response CultureЧтениеKnowledge Check: Incident Management and On Call Practices Задание

Module Wrap-Up and Assessment

Monitoring, Alerting, and Incident OperationsDIALOGUEKnowledge Check: Monitoring, Alerting, and Incident Operations ЗаданиеModule Summary: Monitoring, Alerting, and Incident OperationsЧтение
03Automation, SLO Tracking, GitOps Recovery, and AI for SRE19 материалов

Toil Reduction and SRE Automation

Toil Reduction Strategies and Runbook StandardizationВидеоHands-On: Creating a Basic SRE RunbookВидеоHands-On: Automating SRE Tasks with Ansible PlaybooksВидеоReducing Operational Toil Through Safe AutomationЧтениеKnowledge Check: Toil Reduction and SRE AutomationЗадание

SLO Tracking and Reliability Decisions

Reliability Reviews and Release DecisionsВидеоHands-On: Tracking SLOs and Error Budgets with PyrraВидеоHands-On: Configuring Burn Rate AlertsВидеоUsing SLO Data to Guide Reliability DecisionsЧтениеKnowledge Check: SLO Tracking and Reliability DecisionsЗадание

GitOps Recovery and AI-Enhanced SRE

GitOps for Reliable OperationsВидеоAI for SRE: Anomaly Detection and Intelligent ResponseВидеоHands-On: Performing GitOps-Based Rollback with Argo CDВидеоHands-On: Detecting Metric Anomalies with Prometheus and GrafanaВидеоStrengthening Recovery with GitOps and AI-Assisted ReliabilityЧтениеKnowledge Check: GitOps Recovery and AI-Enhanced SREЗадание

Module Wrap-Up and Assessment

Automation, SLO Tracking, GitOps Recovery, and AI for SRE DIALOGUEKnowledge Check: Automation, SLO Tracking, GitOps Recovery, and AI for SRE ЗаданиеModule Summary: Automation, SLO Tracking, GitOps Recovery, and AI for SREЧтение
04Course Wrap-Up and Assessments5 материалов
Practice Project: Building an SRE Operating Model for Reliable Cloud Native ServicesЧтениеApplying Site Reliability Engineering Practices in a Production IncidentDIALOGUEEnd Course Knowledge Check: Site Reliability Engineering (SRE) PrinciplesЗаданиеDesigning Automated SRE Core Principles for Reliable OperationsЗаданиеCourse SummaryВидео