К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Preparing Multimodal Data: Vision, Audio, and NLP Pipelines · LearnSpace
Назад в каталог
courseraПрограммирование

Preparing Multimodal Data: Vision, Audio, and NLP Pipelines

Курс от Coursera
Средний≈ 13.7 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

Raw images, audio clips, and text are only valuable when transformed into formats that AI models can actually use. This intermediate course equips you with the hands-on skills to build multimodal data processing pipelines across three core data types — visual, audio, and language — and to evaluate the AI models trained on them. You will preprocess and enhance image data using normalization, color-space conversion, and quality correction techniques. You will extract motion features from video using optical flow and frame differencing. On the audio side, you will apply spectral and cepstral feature extraction and build augmentation pipelines that improve model robustness. For language, you will fine-tune transformer models on domain-specific datasets and construct end-to-end text preprocessing pipelines using industry-standard tools. Grounded in real-world job tasks from machine learning and AI roles, this course prepares you to take raw, unstructured data and shape it into training-ready inputs — a skill in high demand across AI, computer vision, speech, and NLP teams.

Навыки, которые вы освоите

Data PreprocessingData TransformationComputer VisionNatural Language ProcessingModel EvaluationImage QualityModel TrainingData PipelinesFeature EngineeringData ProcessingMachine Learning MethodsHugging FaceArtificial Intelligence and Machine Learning (AI/ML)Artificial Neural NetworksLarge Language ModelingMachine Learning SoftwareFine-tuningData ArchitectureMachine Learning AlgorithmsImage Analysis

Программа курса

13 модулей · 85 учебных материалов

01Image Preprocessing and Normalization6 материалов
Why Image Preprocessing Determines Model SuccessDIALOGUENormalization Techniques and Color-Space FundamentalsВидеоImplementation Patterns for Image Preprocessing PipelinesЧтениеHow to Implement Image Normalization with NumPy and OpenCVЧтение

Учитесь у экспертов

Professionals from the Industry

Преподаватель курса

Preparing Multimodal Data: Vision, Audio, and NLP Pipelines
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 13.7 ч

13 модулей

Язык: Английский

Субтитры: Арабский, Французский, Итальянский, Бразильский португальский, Корейский, Немецкий, Пушту, Испанский, Дари, Японский

Часть программы вашего университета
Build Production Image Preprocessing PipelineЗадание
Image Preprocessing Knowledge CheckЗадание
02Motion Detection and Optical Flow7 материалов
Why Motion Analysis Is Critical for AI System SuccessDIALOGUEOptical Flow Algorithms and Frame Differencing MathematicsВидеоMotion Vector Analysis and Performance OptimizationЧтениеHow to Implement Optical Flow with OpenCV and NumPyЧтениеImplement Motion-Based Object Tracking SystemЛабораторнаяMotion Detection and Optical Flow Fundamentals Knowledge CheckЗаданиеComprehensive Motion Analysis AssessmentЗадание
03Image Quality Analysis Fundamentals6 материалов
Why Image Quality Analysis Matters in Production SystemsВидеоFundamentals of Image Quality AssessmentВидеоDiagnosing Image Quality Issues in Computer Vision DatasetsЧтениеDiagnostic Assessment Workflow DesignDIALOGUEComputer Vision Quality Diagnostic ReportЗаданиеImage Quality Diagnostic AssessmentЗадание
04Apply Targeted Mitigation Techniques7 материалов
Why Algorithmic Enhancement Saves Production DeploymentsВидеоAlgorithmic Enhancement Techniques OverviewВидеоImplementing Unsharp Masking for Blur CorrectionЧтениеOptimizing Enhancement Parameters for Production SystemsDIALOGUEAlgorithmic Image Enhancement: Deblurring, Denoising, and Histogram CorrectionЛабораторнаяApply Targeted Mitigation TechniquesЗаданиеImage Quality Enhancement Mastery AssessmentЗадание
05Spectral and Cepstral Feature Extraction for Audio Analysis7 материалов
Why Audio Feature Extraction Matters in Production ML SystemsВидеоSpectral Analysis Fundamentals: STFT and Mel-Scale FeaturesВидеоCepstral Analysis and MFCC Feature ExtractionЧтение Exploring MFCC Parameter Selection for Production SystemsDIALOGUEComputing MFCCs with Librosa: Step-by-Step ImplementationВидеоOptimizing MFCC Features for Environmental Sound RecognitionЗаданиеSpectral and Cepstral Feature Extraction Knowledge CheckЗадание
06Audio Augmentation Techniques for Real-World Model Generalization7 материалов
Audio Augmentation for Production ML SystemsDIALOGUEAudio Augmentation Techniques: Noise, Temporal, and Spectral TransformationsВидеоDesigning Robust Augmentation Pipelines for Production SystemsЧтениеBuilding Audio Augmentation Pipelines with Python and LibrosaВидеоBuild Production-Ready Audio Augmentation PipelinesЛабораторнаяAudio Augmentation Pipeline Design and ImplementationЗаданиеAudio Feature Extraction and Augmentation for Production ML SystemsЗадание
07Audio Model Performance Metrics & Analysis7 материалов
Why Audio Model Performance Monitoring Matters in ProductionВидеоEssential Audio Model Performance Metrics and Calculation MethodsВидеоPerformance Metrics in Production Audio Systems: Industry Applications and Best PracticesЧтениеCalculating Performance Metrics with Python for Audio Model Evaluation Видео Evaluating Performance Degradation Patterns Across User CohortsDIALOGUEAudio Model Performance Dashboard: Calculating WER and F1-Scores for User Cohort AnalysisЛабораторнаяPerformance Metrics Evaluation AssessmentЗадание
08Enhancing Audio Model Robustness through Augmentation Pipelines7 материалов
Why Systematic Root Cause Analysis Matters for Audio Model Reliability DIALOGUESystematic Root Cause Analysis Framework for Audio Model DebuggingЧтениеAudio Sample Error Analysis Using Spectrograms and Signal Processing ToolsВидеоImplementing Root Cause Investigation Workflow for Production Audio ModelsВидеоComplete Audio Model Debugging Investigation and Remediation Plan ЗаданиеRoot Cause Analysis and Systematic Debugging Assessment ЗаданиеComprehensive Audio Model Debugging and Root Cause Analysis EvaluationЗадание
09Fine-Tuning Transformer Language Models6 материалов
Why Domain-Specific Language Models Transform Business IntelligenceВидео Understanding Transformer Fine-Tuning Architecture and ProcessВидеоHugging Face Transformers Framework and Fine-Tuning ComponentsЧтение Implementing BERT Fine-Tuning with Hugging Face TrainerВидеоAnalyzing Fine-Tuning Decisions for Domain-Specific NLP Applications DIALOGUE Fine-Tuning Transformer Models Knowledge CheckЗадание
10Text Preprocessing Pipeline Development7 материалов
Why Text Preprocessing Pipelines Are Critical for NLP SuccessDIALOGUE spaCy Framework and Text Processing ComponentsЧтение Building Text Preprocessing Pipelines with spaCy ComponentsВидео Creating Automated Text Preprocessing Pipelines with spaCyВидеоBuild Production-Ready Text Preprocessing Pipelines with spaCyЛабораторная Text Preprocessing Pipeline Knowledge CheckЗадание Comprehensive NLP Fine-Tuning and Text Preprocessing AssessmentЗадание
11Introduction to Dual Evaluation Methodology6 материалов
Why Dual Evaluation Matters in Production AI SystemsВидеоAutomated Metrics Fundamentals for Language Model AssessmentВидеоHuman-in-the-Loop Evaluation Framework DesignЧтениеLanguage Model Evaluation: Automatic and Human-in-the-Loop MetricsВидеоAnalyzing Evaluation Trade-offs in Model Selection ScenariosDIALOGUEAutomated Metrics and Human Evaluation Concepts Knowledge CheckЗадание
12Implementing Comprehensive Model Assessment7 материалов
When Automated Metrics Miss Critical Quality IssuesВидеоIntegration Strategies for Automated and Human Evaluation MethodsВидеоDesigning Production-Ready Evaluation WorkflowsDIALOGUEComputing Automated Metrics with Python Evaluation LibrariesВидеоImplementing Comprehensive Language Model AssessmentЛабораторнаяIntegrated Evaluation Strategy AssessmentЗаданиеComprehensive Language Model Evaluation AssessmentЗадание
13Project: Preparing Multimodal Data: Vision, Audio, and NLP Pipelines5 материалов
Why This Project MattersЧтениеProject RequirementsЧтениеAssignment: Multimodal Data Processing PipelineЧтениеGraded Quiz: Multimodal Data Processing PipelineЗаданиеSolution KeyЧтение