К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Generative AI Advanced Fine-Tuning for LLMs · LearnSpace
Назад в каталог
courseraАнализ данных

Generative AI Advanced Fine-Tuning for LLMs

Курс от IBM
Средний≈ 9 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

"Fine-tuning large language models (LLMs) is essential for aligning them with specific business needs, improving accuracy, and optimizing performance. In today’s AI-driven world, organizations rely on fine-tuned models to generate precise, actionable insights that drive innovation and efficiency. This course equips aspiring generative AI engineers with the in-demand skills employers are actively seeking. You’ll explore advanced fine-tuning techniques for causal LLMs, including instruction tuning, reward modeling, and direct preference optimization. Learn how LLMs act as probabilistic policies for generating responses and how to align them with human preferences using tools such as Hugging Face. You’ll dive into reward calculation, reinforcement learning from human feedback (RLHF), proximal policy optimization (PPO), the PPO trainer, and optimal strategies for direct preference optimization (DPO). The hands-on labs in the course will provide real-world experience with instruction tuning, reward modeling, PPO, and DPO, giving you the tools to confidently fine-tune LLMs for high-impact applications. Build job-ready generative AI skills in just two weeks! Enroll today and advance your career in AI!"

Навыки, которые вы освоите

Fine-tuningReinforcement LearningGenerative AILarge Language ModelingModel OptimizationLLM ApplicationModel TrainingMachine Learning MethodsModel Evaluation

Программа курса

2 модулей · 41 учебных материалов

01Different Approaches to Fine-Tuning 17 материалов

Welcome to the Course

Course IntroductionВидеоCourse OverviewЧтениеSpecialization OverviewЧтениеHelpful tips for Course CompletionPLUGIN

Instruction-Tuning and Reward Modeling

Учитесь у экспертов

Joseph Santarcangelo

Ph.D., Data Scientist at IBM

Ashutosh Sagar

Преподаватель курса

Wojciech 'Victor' Fulmyk

Преподаватель курса

Fateme Akbari

Преподаватель курса

Generative AI Advanced Fine-Tuning for LLMs
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 9 ч

2 модулей

Язык: Английский

Субтитры: Арабский, Французский, Узбекский, Украинский, Китайский (Китай), Греческий, Итальянский, Бразильский португальский, Вьетнамский, Нидерландский, Корейский, Немецкий, Пушту, Русский, Тайский, Индонезийский, Шведский, Турецкий, Азербайджанский, Испанский, Хинди, Японский, Казахский, Венгерский, Польский

Часть программы вашего университета
Basics of Instruction-TuningВидео
Instruction-Tuning with Hugging FaceВидео
Instruction TuningPLUGIN
Instruction Fine-Tuning LLMsВнешний инструмент
Best Practices for Instruction-Tuning Large Language Models Чтение
Reward Modeling: Response EvaluationВидео
Reward Model Training Видео
Reward Modeling with Hugging FaceВидео
Reward Modeling & Response EvaluationPLUGIN
Lab: Reward ModelingВнешний инструмент
Practice Quiz: Instruction-Tuning and Reward Modeling  Задание
Summary and Highlights Чтение
Different Approaches to Instruction-TuningЗадание
02Fine-Tuning Causal LLMs with Human Feedback and Direct Preference24 материалов

PPO

Large Language Models (LLMs) as DistributionsВидеоFrom Distributions to PoliciesВидеоReinforcement Learning from Human Feedback (RLHF)ВидеоProximal Policy Optimization (PPO)ВидеоPPO with Hugging FaceВидеоPPO TrainerВидеоLog-derivative TrickPLUGINLab: Reinforcement Learning from Human Feedback using PPOВнешний инструментSummary and Highlights ЧтениеPractice Quiz: Proximal Policy Optimization (PPO)Задание

DPO

DPO: Partition FunctionВидеоDPO: Optimal SolutionВидеоFrom Optimal Policy to DPOВидеоDPO with Hugging FaceВидеоLab: Direct Preference Optimization (DPO) using Hugging FaceВнешний инструментFine-tune LLMs Locally with InstructLabPLUGIN

Fine-Tuning Causal LLMs with Human Feedback and Direct Preference Assessment

Fine-Tuning Causal LLMs with Human Feedback and Direct PreferenceЗадание

Course Cheatsheet and Glossary

Cheat Sheet: Generative AI Advanced Fine-Tuning for LLMsPLUGIN Glossary: Generative AI Advance Fine-Tuning for LLMsPLUGIN

Course Wrap-Up

Course ConclusionЧтениеCongratulations and Next StepsЧтениеThanks from the course teamЧтение
Summary and HighlightsЧтение
Practice Quiz: Direct Preference Optimization (DPO)Задание