К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Programming Generative AI: Multimodal AI and Model Customization · LearnSpace
Назад в каталог
courseraПрограммирование

Programming Generative AI: Multimodal AI and Model Customization

Курс от Pearson
Средний≈ 8.3 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

Unlock the full potential of generative AI with our advanced course module focused on state-of-the-art multimodal models. This course is designed for learners eager to bridge the gap between images and text, and to master the latest techniques in AI-driven content generation. You’ll begin by exploring the foundational concepts behind multimodal models, learning how contrastive language-image pre-training enables seamless integration of visual and textual data. Discover how these models power innovative applications like semantic image search, allowing you to query image content without manual labeling. Dive deeper into the mechanics of latent diffusion models and unravel the inner workings of stable diffusion, gaining the skills to transform text prompts into entirely new, never-before-seen images. The course also covers essential strategies for evaluating generative models and introduces efficient methods for fine-tuning and adapting pre-trained models to new styles and subjects. By the end, you’ll be equipped to build, adapt, and optimize cutting-edge text-to-image systems—ready to innovate in creative, research, or commercial settings.

Навыки, которые вы освоите

Model OptimizationFine-tuningModel EvaluationGenerative AIGenerative Model ArchitecturesImage AnalysisAutoencodersImage QualityComputer VisionMultimodal PromptsModel TrainingEmbeddingsPrompt Engineering

Программа курса

1 модулей · 47 учебных материалов

01Programming Generative AI: Unit 347 материалов

Connecting Text and Images

TopicsВидеоComponents of a Multimodal ModelВидеоVision-Language UnderstandingВидеоContrastive Language-Image PretrainingВидеоEmbedding Text and Images with CLIPВидео

Учитесь у экспертов

Pearson

Преподаватель курса

Jonathan Dinu

Freelance ML Engineer

Programming Generative AI: Multimodal AI and Model Customization
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 8.3 ч

1 модулей

Язык: Английский

Субтитры: Американский английский

Часть программы вашего университета
Zero-Shot Image Classification with CLIPВидео
Semantic Image Search with CLIPВидео
Conditional Generative ModelsВидео
Introduction to Latent Diffusion ModelsВидео
The Latent Diffusion Model ArchitectureВидео
Failure Modes and Additional ToolsВидео
Stable Diffusion DeconstructedВидео
Writing Our Own Stable Diffusion PipelineВидео
Decoding Images from the Stable Diffusion Latent SpaceВидео
Improving Generation with GuidanceВидео
Playing with PromptsВидео
Connecting Text and Images QuizЗадание

Post-Training Procedures for Diffusion Models

TopicsВидеоMethods and Metrics for Evaluating Generative AIВидеоManual Evaluation of Stable Diffusion with DrawBenchВидеоQuantitative Evaluation of Diffusion Models with Human Preference PredictorsВидеоOverview of Methods for Fine-Tuning Diffusion ModelsВидеоSourcing and Preparing Image Datasets for Fine-TuningВидеоGenerating Automatic Captions with BLIP-2ВидеоParameter Efficient Fine-Tuning with LoRAВидеоInspecting the Results of Fine-TuningВидеоInference with LoRAs for Style-Specific GenerationВидеоConceptual Overview of Textual InversionВидеоSubject-Specific Personalization with DreamboothВидеоDreambooth versus LoRA Fine-TuningВидеоDreambooth Fine-Tuning with Hugging FaceВидеоInference with Dreambooth to Create Personalized AI AvatarsВидеоAdding Conditional Control to Text-to-Image Diffusion ModelsВидеоCreating Edge and Depth Maps for ConditioningВидеоDepth and Edge-Guided Stable Diffusion with ControlNetВидеоUnderstanding and Experimenting with ControlNet ParametersВидеоGenerative Text Effects with Font Depth MapsВидеоFew Step Generation with Adversarial Diffusion Distillation (ADD)ВидеоReasons to DistillВидеоComparing SDXL and SDXL TurboВидеоText-Guided Image-to-Image TranslationВидеоVideo-Driven Frame-by-Frame Generation with SDXL TurboВидеоNear Real-Time Inference with PyTorch Performance OptimizationsВидеоProgramming Generative AI: SummaryВидеоCourse SummaryВидеоPost-Training Procedures for Diffusion Models QuizЗаданиеEnd of Assessment QuizЗадание