К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Build Multimodal Generative AI Applications · LearnSpace
Назад в каталог
courseraПрограммирование

Build Multimodal Generative AI Applications

Курс от IBM
Средний≈ 7.7 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

Ready to level up your GenAI skills? Step into the exciting world of multimodal AI, where language, images, and speech come together to build smarter, more interactive applications. In this hands-on course, you’ll learn how to build systems that work across multiple modalities, from creating AI-powered storytellers and meeting assistants to developing image captioning tools and video generation apps. You’ll gain experience with real-world tools like IBM’s Granite, OpenAI’s Whisper, Sora and DALL·E, Meta’s Llama, Mistral’s Mixtral, and Gradio. Plus, you'll explore multimodal search, question answering, and retrieval systems that combine text, speech, and visual data. By the end of the course, you’ll be able to design and build full-stack multimodal AI solutions using Python and frameworks like Flask and Gradio. If you’re looking to gain in-demand skills for building the next generation of AI applications, enroll today and power up your AI career!

Навыки, которые вы освоите

Prompt EngineeringMultimodal PromptsAI PersonalizationWeb ApplicationsAI IntegrationsLLM ApplicationFlask (Web Framework)Retrieval-Augmented GenerationLarge Language ModelingWeb DevelopmentAI powered creativityEmbeddingsOpenAI APISoftware Development

Программа курса

3 модулей · 38 учебных материалов

01Foundations of Multimodal AI17 материалов

Welcome to the Course

Video: Course IntroductionВидеоReading: Course OverviewЧтениеRAG and Agentic AI Professional Certificate OverviewВидеоHelpful Tips for Course CompletionPLUGIN

Introduction to Multimodal AI: Text and Speech Processing

Учитесь у экспертов

Hailey Quach

Преподаватель курса

Ricky Shi

Data Scientist

Build Multimodal Generative AI Applications
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 7.7 ч

3 модулей

Язык: Английский

Субтитры: Арабский, Французский, Узбекский, Украинский, Китайский (Китай), Греческий, Итальянский, Бразильский португальский, Вьетнамский, Нидерландский, Корейский, Немецкий, Пушту, Русский, Тайский, Индонезийский, Шведский, Турецкий, Азербайджанский, Испанский, Хинди, Японский, Казахский, Венгерский, Польский

Часть программы вашего университета
Introduction to Multimodal AI Видео
Reading: What is Multimodal Generative AI and Why Does It Matter? PLUGIN
Reading: What is Computer Vision? PLUGIN
Text-to-Speech Technologies Видео
Speech-to-Text Technologies Видео
Reading: Text Processing, Speech Processing, and Text-to-Speech PLUGIN
Lab: Use Mistral and gTTS to Create Your Personal StorytellerВнешний инструмент
Reading: Challenges in Multimodal AI Integration PLUGIN
Lab: Build a Meeting Assistant with Whisper, LangChain, & GradioВнешний инструмент
Practice Quiz: Introduction to Multimodal AI: Text and Speech ProcessingЗадание

Module Summary and Evaluation

Reading: Summary and Highlights ЧтениеCheat Sheet: Foundations of Multimodal AI PLUGINGraded Quiz: Foundations of Multimodal AIЗадание
02Integrating Visual and Video Modalities10 материалов

Image Generation and Captioning

Understanding Image Captioning with Meta's LlamaВидеоReading: Introduction to Text-to-Video and Image-to-Video TechnologiesPLUGINDemo: Text-to-Video Generation with OpenAI's SoraВидеоLab: DALL·E Image Generation Guide for BeginnersВнешний инструментReading: Strengths, Limitations, and Practical Applications of Multimodal Vision Models in Real World ScenariosPLUGINLab: Build an Image Captioning System with watsonx and IBM's GraniteВнешний инструментImage Generation and Captioning Задание

Module Summary and Evaluation

Reading: Summary and Highlights ЧтениеCheat Sheet: Integrating Visual and Video Modalities PLUGINGraded Quiz: Integrating Visual and Video Modalities Задание
03Advanced Multimodal Applications11 материалов

Build Advanced Multimodal Applications

Introduction to Multimodal Retrieval-Augmented Generation (MM-RAG)ВидеоLab: Build a Style Finder Using Multimodal Retrieval and SearchВнешний инструментMultimodal Chatbots and QA Systems ВидеоLab: Building Your First GenAI-Powered Image-Based Web Application: AI Nutrition CoachВнешний инструментBuild Advanced Multimodal ApplicationsЗадание

Module Summary and Evaluation

Summary and Highlights ЧтениеCheat Sheet: Advanced Multimodal Applications PLUGINGraded Quiz: Advanced Multimodal Applications Задание

Course Wrap-Up

Course Wrap-upВидеоReading: Congratulations and Next Steps ЧтениеThanks from the Course Team Чтение