К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Modern AI Models for Vision and Multimodal Understanding · LearnSpace
Назад в каталог
courseraПрограммирование

Modern AI Models for Vision and Multimodal Understanding

Курс от University of Colorado Boulder
Продвинутый≈ 12.7 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

Step into the frontier of artificial intelligence with this advanced course designed to explore the latest models powering visual and multimodal intelligence. From foundational mathematical tools to state-of-the-art architectures, you'll gain the skills to understand and build systems that interpret images, text, and more—just like today’s leading AI models. You'll begin by discovering how Nonlinear Support Vector Machines (NSVMs) and Fourier transforms lay the groundwork for signal processing and pattern recognition in visual data. You'll then build a strong foundation in probabilistic reasoning and temporal modeling with RNNs, enabling AI systems to understand sequences and context. After, you'll learn how transformer architectures revolutionize both language and vision tasks. Finally, you'll dive into multimodal learning with CLIP, which connects images and text, and explore diffusion models that generate high-fidelity images through iterative refinement. This course is ideal for learners who want to go beyond traditional deep learning and explore the models shaping the future of AI. With a blend of theory, code, and real-world applications, you'll be equipped to tackle cutting-edge challenges in computer vision and multimodal AI. This course can be taken for academic credit as part of CU Boulder’s MS in Data Science or MS in Computer Science degrees offered on the Coursera platform. These fully accredited graduate degrees offer targeted courses, short 8-week sessions, and pay-as-you-go tuition. Admission is based on performance in three preliminary courses, not academic history. CU degrees on Coursera are ideal for recent graduates or working professionals. Learn more: MS in Data Science: https://www.coursera.org/degrees/master-of-science-data-science-boulder MS in Computer Science: https://coursera.org/degrees/ms-computer-science-boulder

Навыки, которые вы освоите

Classification AlgorithmsEmbeddingsRecurrent Neural Networks (RNNs)Digital Signal ProcessingGenerative Model ArchitecturesVision Transformer (ViT)Supervised LearningTransfer LearningMachine Learning Methods

Программа курса

4 модулей · 88 учебных материалов

01SMV and Fourier27 материалов

Course Introduction

Course Updates and Accessibility SupportЧтениеEarn Academic Credit for your Work!ЧтениеCourse SupportЧтениеMeet Your Instructor Видео

Учитесь у экспертов

Tom Yeh

Associate Professor

Modern AI Models for Vision and Multimodal Understanding
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 12.7 ч

4 модулей

Язык: Английский

Субтитры: Казахский, Венгерский, Узбекский

Часть программы вашего университета
Inside the CourseЧтение
Assessment ExpectationsЧтение
AI Citation and AcknowledgementЧтение

SVM

Get the Workbook: SVMЧтениеLinear SVMВидеоVisualize LinearВидеоRadial Basis Function (RBF)ВидеоRBF KernelВидеоVisualize a RBF SVMВидеоSupport Vector Machine (SVM)Задание

Fourier 1D

Get the Workbook: Fourier 1D & 2DЧтение1D DFTВидео1D Inverse DFT Видео1D Basic FunctionsВидеоFrequency and TimeВидеоFourier 1DЗадание

Fourier 2D

2D DFTВидео2D Inverse DFTВидео2D Basic FunctionsВидеоFrequency and Spatial ВидеоFourier 2DЗадание

End-of-Module Knowledge Assessment

AI Policy QuizЗаданиеSMV and FourierЗадание
02Probability and RNN22 материалов

Probability Part One

Get the Workbook: ProbabilityЧтениеProbability in Language Models ВидеоConditional Probabilities ВидеоThe Chain Rule of ProbabilitiesВидеоCalculating Joint Probabilities ВидеоProbability Part OneЗадание

Probability Part Two

Pixel-Base Image ModelsВидеоAutoregressive Image ModelВидеоAttention Mechanisms in Transformer ModelsВидеоProbability Part TwoЗадание

RNN Part One

Get the Workbook: RNNЧтениеBatch vs RecurrentВидеоMLP vs RNNВидеоMany to OneВидеоOne to ManyВидеоRNN Part OneЗадание

RNN Part Two

One to OneВидеоSequence to SequenceВидеоDeep RNNВидеоAutoregressive RNNВидеоRNN Part TwoЗадание

End-of-Module Knowledge Assessment

Probability and RNNЗадание
03Transformer and ViT22 материалов

Transformer Part One

Get the Workbook: TransformerЧтениеBatch vs Recurrent vs AttentionВидеоAttention + MLPВидеоDot-Product Self-AttentionВидеоQKV Self-AttentionВидеоTransformer Part OneЗадание

Transformer Part Two

Transformer EncoderВидеоSelf vs Cross AttentionВидеоEncoder and Decoder for TransformerВидеоDecoder Output LayerВидеоTransformer Part TwoЗадание

ViT Part One

Get the Workbook: ViTЧтениеImage to TokensВидеоNormalization for ViTВидеоSelf-Attention for ViTВидеоViT Part OneЗадание

ViT Part Two

Multi-Head AttentionВидеоMLP Forward FeedВидеоViT Output LayerВидеоLoss Gradient for ViTВидеоViT Part TwoЗадание

End-of-Module Knowledge Assessment

Transformer and ViTЗадание
04CLIP and Diffusion17 материалов

CLIP Part One

Get the Workbook: CLIPЧтениеBatch of PairsВидеоImage Encoder (Batch)ВидеоText Encoder (Batch)ВидеоCLIP Part OneЗадание

CLIP Part Two

Joint EmbeddingВидеоContrastive Pre-TrainingВидеоZero-Shot Image ClassifierВидеоZero-Shot Image PredictionВидеоCLIP Part TwoЗадание

Diffusion

Get the Workbook: DiffusionЧтениеDiffusion IntroductionВидеоNoise PredictionВидеоTime Conditioning and Parallel TrainingВидеоReverse DiffusionВидеоDiffusionЗадание

End-of-Module Knowledge Assessment

CLIP and DiffusionЗадание