К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Transformers for Vision AI, Multimodal & Generative AI · LearnSpace
Назад в каталог
courseraПрограммирование

Transformers for Vision AI, Multimodal & Generative AI

Курс от Packt
Средний≈ 6 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

Discover how transformer architectures are revolutionizing computer vision and multimodal AI, and explore the latest advancements toward general artificial intelligence. Learn to apply transformers to image, video, and cross-modal tasks. This course explores the expanding frontier of transformer models beyond natural language, focusing on their applications in computer vision, multimodal AI, and generative ideation. Learners will investigate vision transformers, text-to-image and text-to-video generation, and the integration of multiple AI models for advanced tasks. The course also addresses risk mitigation in large models and looks ahead to the future of AI with functional AGI and creative generative systems. By the end, you will understand how to harness transformers for cutting-edge vision and multimodal applications. With a focus on emerging trends and practical implementations, this course guides learners through the latest research and real-world use cases in vision and multimodal AI. Concepts are introduced progressively, enabling learners to build expertise in applying transformers across diverse domains. This course is part three of a three-course Specialization designed to build a complete and cohesive understanding of the subject. While it offers valuable skills on its own, you'll gain the most benefit by progressing through all three courses as a structured learning journey. This course is based on Transformers for Natural Language Processing and Computer Vision, by Denis Rothman. Packt is one of the world's most prolific publishers of cutting-edge technical content. For over two decades we've made it our mission to curate and publish the knowledge of only the very best technical experts. We focus on real-world courses that help our customers get the job done, with coverage that extends across a wide range of established and cutting-edge technical topics. If you're an individual or an organisation that embraces learning by doing, Packt is the perfect fit for you.

Навыки, которые вы освоите

AI OrchestrationVision Transformer (ViT)Generative Model ArchitecturesPrompt EngineeringHugging FaceChatGPTOpenAIComputer VisionLarge Language ModelingAI WorkflowsResponsible AIModel TrainingMultimodal PromptsImage AnalysisArtificial Intelligence and Machine Learning (AI/ML)Model EvaluationAI powered creativityGenerative AI AgentsGenerative AIEmbeddings

Программа курса

5 модулей · 43 учебных материалов

01Beyond Text: Vision Transformers in the Dawn of Revolutionary AI10 материалов

Unveiling Vision Transformers: From Image Patches to Multimodal AI

OverviewВидеоIntroductionЧтениеA Feature Extractor SimulatorЧтениеConfiguration and ShapesЧтение

Учитесь у экспертов

Packt - Course Instructors

Преподаватель курса

Transformers for Vision AI, Multimodal & Generative AI
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 6 ч

5 модулей

Язык: Английский

Часть программы вашего университета
Explaining Vision Transformers to a Product StakeholderDIALOGUE
CLIPЧтение
DALL-E 2 and DALL-E 3Чтение
From Research to Mainstream AI with DALL-EЧтение
Example 1: A standard image and textЧтение
Exploring Multimodal AI Models and Their CapabilitiesЗадание
02Transcending the Image-Text Boundary with Stable Diffusion6 материалов

Unleashing Creativity: Hands-On with Stable Diffusion for Images and Video

OverviewВидеоIntroductionЧтениеExplaining Stable Diffusion's Text-to-Image GenerationDIALOGUERunning the Keras Stable Diffusion ImplementationЧтениеText-to-Video with a Variation of OpenAI CLIPЧтениеExploring Diffusion Models and Image-Video SynthesisЗадание
03Hugging Face AutoTrain: Training Vision Models without Coding8 материалов

No-Code Vision Model Training with AutoTrain: From Data to Deployment

OverviewВидеоIntroductionЧтениеTraining Models with AutoTrainЧтениеPresenting a No-Code Vision Model Solution to a Non-Technical ManagerDIALOGUEViT For Image ClassificationЧтениеBeit For Image ClassificationЧтениеConvNext for Image ClassificationЧтениеExploring Vision Model Training with Hugging Face AutoTrainЗадание
04On the Road to Functional AGI with HuggingGPT and its Peers9 материалов

Building Intelligent Pipelines: Integrating AI Models for Advanced Task Automation

OverviewВидеоIntroductionЧтениеDefining F-AGIЧтениеInstalling and ImportingЧтениеExplaining F-AGI, Model Chaining, and AI PipelinesDIALOGUELevel 2: DifficultЧтениеCustomGPTЧтениеModel Chaining: Chaining Google Cloud Vision to ChatGPTЧтениеExploring the Path to Advanced AI SystemsЗадание
05Beyond Human-Designed Prompts with Generative Ideation10 материалов

Automating Creativity: Building AI-Driven Content and Image Pipelines

OverviewВидеоIntroductionЧтениеChatGPT with GPT-4 Provides a Graph to Illustrate the PresentationЧтениеImplementing Llama 2 with Hugging FaceЧтениеAdvising on Automated Generative Ideation ImplementationDIALOGUEMidjourneyЧтениеMicrosoft DesignerЧтениеThe Future Is Yours!ЧтениеExploring Generative Ideation and AI IntegrationЗаданиеTransformers Beyond Text: Vision, Multimodal AI, and the Future of Generative Models Final AssessmentЗадание