К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Fundamentals of Reinforcement Learning · LearnSpace
Назад в каталог
courseraАнализ данных

Fundamentals of Reinforcement Learning

Курс от University of Alberta, Alberta Machine Intelligence Institute
Средний≈ 15.2 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

Reinforcement Learning is a subfield of Machine Learning, but is also a general purpose formalism for automated decision-making and AI. This course introduces you to statistical learning techniques where an agent explicitly takes actions and interacts with the world. Understanding the importance and challenges of learning agents that make decisions is of vital importance today, with more and more companies interested in interactive agents and intelligent decision-making. This course introduces you to the fundamentals of Reinforcement Learning. When you finish this course, you will: - Formalize problems as Markov Decision Processes - Understand basic exploration methods and the exploration/exploitation tradeoff - Understand value functions, as a general-purpose tool for optimal decision-making - Know how to implement dynamic programming as an efficient solution approach to an industrial control problem This course teaches you the key concepts of Reinforcement Learning, underlying classic and modern algorithms in RL. After completing this course, you will be able to start using RL for real problems, where you have or can specify the MDP. This is the first course of the Reinforcement Learning Specialization.

Навыки, которые вы освоите

Reinforcement LearningMachine Learning AlgorithmsMarkov ModelAgentic systemsAlgorithmsMachine LearningArtificial IntelligenceDecision Intelligence

Программа курса

5 модулей · 66 учебных материалов

01Welcome to the Course! 7 материалов

Course Introduction

Specialization IntroductionВидеоCourse IntroductionВидеоMeet your instructors!ВидеоYour Specialization RoadmapВидео

Учитесь у экспертов

Martha White

Assistant Professor

Adam White

Assistant Professor

Fundamentals of Reinforcement Learning
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 15.2 ч

5 модулей

Язык: Английский

Субтитры: Арабский, Французский, Бенгальский, Узбекский, Украинский, Китайский (Китай), Греческий, Итальянский, Бразильский португальский, Вьетнамский, Нидерландский, Корейский, Немецкий, Пушту, Урду, Русский, Тайский, Индонезийский, Шведский, Турецкий, Азербайджанский, Испанский, Дари, Хинди, Японский, Казахский, Венгерский, Польский

Часть программы вашего университета
Reinforcement Learning TextbookЧтение
Read Me: Pre-requisites and Learning ObjectivesЧтение
Meet and Greet!Обсуждение
02An Introduction to Sequential Decision-Making16 материалов

The K-Armed Bandit Problem

Module 1 Learning ObjectivesЧтениеWeekly ReadingЧтениеLet's play a game!PLUGINSequential Decision Making with Evaluative FeedbackВидеоCompare bandits to supervised learningОбсуждение

What to Learn? Estimating Action Values

Learning Action ValuesВидеоWhat's underneath?PLUGINEstimating Action Values IncrementallyВидео

Exploration vs. Exploitation Tradeoff

What is the trade-off?ВидеоOptimistic Initial ValuesВидеоUpper-Confidence Bound (UCB) Action SelectionВидеоJonathan Langford: Contextual Bandits for Real World Reinforcement LearningВидеоWeek 1 SummaryВидеоChapter SummaryЧтение

Weekly Assessment

Sequential Decision-MakingЗаданиеBandits and Exploration/ExploitationПрограммирование
03Markov Decision Processes12 материалов

Introduction to Markov Decision Processes

Module 2 Learning ObjectivesЧтениеWeekly ReadingЧтениеMarkov Decision ProcessesВидеоExamples of MDPsВидео

Goal of Reinforcement Learning

The Goal of Reinforcement LearningВидеоIs the reward hypothesis sufficient?ОбсуждениеMichael Littman: The Reward HypothesisВидео

Continuing Tasks

Continuing TasksВидеоExamples of Episodic and Continuing TasksВидеоWeek 2 SummaryВидео

Weekly Assesment

MDPsЗаданиеGraded Assignment: Describe Three MDPsВзаимная проверка
04Value Functions & Bellman Equations 15 материалов

Policies and Value Functions

Module 3 Learning ObjectivesЧтениеWeekly ReadingЧтениеSpecifying PoliciesВидеоValue FunctionsВидеоRich Sutton and Andy Barto: A brief History of RLВидео

Bellman Equations

Bellman Equation DerivationВидеоWhy Bellman Equations?Видео

Optimality (Optimal Policies & Value Functions)

Optimal PoliciesВидеоOptimal Value FunctionsВидеоUsing Optimal Value Functions to Get Optimal PoliciesВидеоCheck-inОбсуждениеWeek 3 SummaryВидеоChapter SummaryЧтение

Weekly Assessment

[Practice] Value Functions and Bellman EquationsЗадание[Graded] Value Functions and Bellman EquationsЗадание
05Dynamic Programming16 материалов

Policy Evaluation (Prediction)

Module 4 Learning ObjectivesЧтениеWeekly ReadingЧтениеPolicy Evaluation vs. ControlВидеоIterative Policy EvaluationВидео

Policy Iteration (Control)

Policy ImprovementВидеоPolicy IterationВидео

Generalized Policy Iteration

Flexibility of the Policy Iteration FrameworkВидеоEfficiency of Dynamic ProgrammingВидеоWarren Powell: Approximate Dynamic Programming for Fleet Management (Short)ВидеоWarren Powell: Approximate Dynamic Programming for Fleet Management (Long)ВидеоWeek 4 SummaryВидеоChapter SummaryЧтение

Weekly Assessment

Dynamic ProgrammingЗаданиеOptimal Policies with Dynamic ProgrammingПрограммированиеWhere can you use dynamic programming?Обсуждение

Course Wrap-up

Congratulations!Видео