К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Transformer Architectures and Multimodal Models · LearnSpace
Назад в каталог
courseraАнализ данных

Transformer Architectures and Multimodal Models

Курс от Edureka
Средний≈ 11.2 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

This course explores the foundations and evolution of modern transformer architectures, taking you from early sequence models to advanced multimodal systems that power today’s AI breakthroughs. Combining strong conceptual depth with practical demonstrations, this course provides a structured journey through attention mechanisms, transformer design, efficiency innovations, and large-scale training strategies. You will begin by understanding Recurrent Neural Networks (RNNs), LSTMs, and GRUs—examining their strengths and limitations in modeling sequential data. From there, you’ll transition into attention mechanisms and multi-head attention, uncovering how transformers overcame long-standing challenges like vanishing gradients and long-term dependency modeling. As the course progresses, you’ll build a deep understanding of encoder-decoder architectures, positional encoding techniques such as sinusoidal embeddings and RoPE, and efficiency innovations like Flash Attention, GQA, and Mixture of Experts (MoE). The course then expands into multimodal learning and similarity-based systems. You’ll explore Vision Transformers (ViTs), embedding alignment techniques, contrastive learning, and large-scale distributed training strategies. Through demonstrations and analysis, you’ll see how modern transformer systems scale to massive datasets while maintaining performance and memory efficiency. By the end of this course, you will be able to: • Explain the limitations of traditional RNN-based sequence models and how attention mechanisms address them. • Implement and analyze multi-head attention and transformer encoder-decoder architectures. • Compare positional encoding strategies and understand their impact on model generalization. • Evaluate efficiency techniques such as Flash Attention, GQA, and MoE for scaling transformers. • Understand Vision Transformers and multimodal representation learning. • Apply similarity learning concepts using embeddings and distance metrics. • Design scalable transformer training systems using distributed and memory-optimized strategies. • Architect transformer-based systems for real-world NLP and multimodal applications. This course is ideal for AI engineers, machine learning practitioners, researchers, and advanced students who want a rigorous understanding of transformer systems beyond surface-level usage. A foundational understanding of Python and basic neural networks will be helpful. Join us to master transformer architectures, explore multimodal intelligence, and build the technical depth required to understand and scale the models shaping modern AI.

Навыки, которые вы освоите

Recurrent Neural Networks (RNNs)EmbeddingsDistributed ComputingVision Transformer (ViT)Software ArchitectureArtificial Intelligence and Machine Learning (AI/ML)Deep LearningNatural Language ProcessingGenerative Model ArchitecturesArtificial IntelligenceModel OptimizationComputer VisionScalabilityUnsupervised LearningModel TrainingMemory ManagementLarge Language ModelingArtificial Neural Networks

Программа курса

4 модулей · 68 учебных материалов

01Sequence Models and Attention Foundations20 материалов

Recurrent Neural Networks (RNN) Foundations

Specialization IntroductionВидеоWelcome to Transformer Architectures and Multimodal ModelsЧтениеCourse IntroductionВидеоRecurrent Neural Networks and BackpropagationВидео

Учитесь у экспертов

Edureka

Преподаватель курса

Transformer Architectures and Multimodal Models
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 11.2 ч

4 модулей

Язык: Английский

Субтитры: Арабский, Французский, Итальянский, Бразильский португальский, Корейский, Немецкий, Испанский, Японский

Часть программы вашего университета
Demonstration: Forward Pass in RNNsВидео
Demonstration: Vanishing Gradient Illustration in RNNВидео
Understanding RNNs: Sequence Modeling and Gradient ChallengesЧтение
Practice Knowledge Check: Recurrent Neural Networks (RNN) FoundationsЗадание

Long Short-Term Memory (LSTM) and Gated Recurrent Units (GRU)

LSTM and GRU: Gated ArchitecturesВидеоDemonstration: LSTM Networks for Sequence ModelingВидеоDemonstration: GRU Based Sequence ModelingВидеоGated Recurrent Networks: Solving Long-Term Dependency ProblemsЧтениеPractice Knowledge Check: Long Short-Term Memory (LSTM) and Gated Recurrent Units (GRU)Задание

Attention and Multi-Head Attention Mechanisms

Self-Attention and Multi-Head Attention ExplainedВидеоDemonstration: Multi-Head Attention in TransformerВидеоDemonstration : Head Contribution AnalysisВидеоAttention Mechanisms: From Context Weighting to Multi-Head RepresentationsЧтениеPractice Knowledge Check: Attention and Multi-Head Attention MechanismsЗадание

Module Wrap-Up and Assessment

Module Summary: Sequence Models and Attention FoundationsЧтениеKnowledge Check: Sequence Models and Attention FoundationsЗадание
02Complete Transformer Architectures22 материалов

Transformer Encoder and Decoder Models

Encoder and Decoder ArchitectureВидеоDemonstration: Encoder Forward Pass in Transformer Encoders: Attention FoundationsВидеоDemonstration: Encoder Forward Pass in Transformer Encoders: Encoder StackВидеоDemonstration: Autoregressive Decoding in Transformer Decoders: Core ComponentsВидеоDemonstration: Autoregressive Decoding in Transformer Decoders: Autoregressive GenerationВидеоTransformer Encoder Decoder ModelsЧтениеPractice Knowledge Check: Transformer BlocksЗадание

Positional Encoding Techniques

Sinusoidal and RoPE EncodingsВидеоDemonstration: RoPE ImplementationВидеоDemonstration: Encoding Comparison: Positional Encoding MechanismВидеоDemonstration: Encoding Comparison: Encoding Impact AnalysisВидеоPositional Encoding MethodsЧтениеPractice Knowledge Check: Positional Encoding TechniquesЗадание

Efficient Transformer Components

Flash Attention GQA and MoEВидеоDemonstration: Memory Efficient Attention: Standard Attention BaselineВидеоDemonstration: Memory Efficient Attention: Optimized AttentionВидеоDemonstration: Expert Routing Visualization: Token to Expert RoutingВидеоDemonstration: Expert Routing Visualization: Capacity and Load Balancing ВидеоEfficient Transformer DesignЧтение

Module Wrap-Up and Assessment

Module Summary: Complete Transformer ArchitecturesЧтениеKnowledge Check: Complete Transformer ArchitecturesЗадание
03Multimodal and Similarity-Based Models23 материалов

Vision Transformers and Multimodal Models

Vision Transformers and Multimodal LearningВидеоDemonstration: Image and Text Embedding Alignment: Similarity ComputationВидеоDemonstration: Image and Text Embedding Alignment: Retrieval VisualizationВидеоDemonstration: Multimodal Representation Analysis: Similarity EvaluationВидеоDemonstration: Multimodal Representation Analysis: Representation GeometryВидеоMultimodal Deep LearningЧтениеPractice Knowledge Check: Multimodal ModelsЗадание

Similarity Learning for Language Models

Text Embeddings and Similarity LearningВидеоDemonstration: Semantic Text Similarity: Computation and Heatmap Analysis ВидеоDemonstration: Semantic Text Similarity: Embedding Space GeometryВидеоDemonstration: Embedding Distance Metrics: Similarity FoundationsВидеоDemonstration: Embedding Distance Metrics: Visualizing and Ranking AnalysisВидеоSimilarity Learning for TextЧтение

Scaling Transformer Training Systems

Distributed Transformer TrainingВидеоDemonstration: Large Model Training Setup: Architecture SetupВидеоDemonstration: Large Model Training Setup: Training and OptimisationВидеоDemonstration: Memory Usage Optimization: Model Setup ВидеоDemonstration: Memory Usage Optimization: Benchmark and ComparisonВидеоScaling Transformer SystemsЧтение

Module Wrap-Up and Assessment

Module Summary: Multimodal and Similarity-Based ModelsЧтениеKnowledge Check: Multimodal and Similarity-Based ModelsЗадание
04Course Wrap-Up3 материалов

Course Wrap-up and Assessments

Practice Project: Building a Multimodal Transformer-Based Knowledge and Similarity EngineЧтениеEnd Knowledge Check: Transformer Architecture and Multimodal ModelsЗаданиеCourse SummaryВидео
Practice Knowledge Check: Efficient Transformer ComponentsЗадание
Practice Knowledge Check: Similarity ModelsЗадание
Practice Knowledge Check: Scaling StrategiesЗадание