К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Introduction to Multimodal AI with Hugging Face · LearnSpace
Назад в каталог
courseraБизнес

Introduction to Multimodal AI with Hugging Face

Курс от Hugging Face
Средний≈ 6.7 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

By the end of this course, you will be able to: • Explain how CLIP aligns image and text in a shared embedding space, use VLMs to perform visual question answering, image captioning, and document understanding, and navigate the Hub for multimodal models. • Build a pipeline that transcribes audio with Whisper and generates images with Diffusers, and describe how LoRA fine-tuning and multimodal RAG extend VLM capabilities. • Build an agentic workflow using smolagents with VLM support and MCP tool integration to automate multi-step tasks requiring vision and reasoning. • Apply ShieldGemma 2 to filter inputs and outputs of a VLM pipeline, test against adversarial inputs, and document failure modes for responsible deployment. AI that can only read text is already behind. This intermediate course assumes you're comfortable with the HF Transformers library and basic Gradio development. It opens with a practical challenge: 2,000 products with photos but no descriptions, and a stack of invoice PDFs that need structured data extraction. You’ll learn how CLIP aligned images and text in a shared space, then use modern vision-language models to caption products, answer questions about charts, and pull fields from invoices. Go wider: transcribe customer calls with Whisper, generate images from text briefs with Diffusers, and learn when to fine-tune a model versus when to give it better context through retrieval. Build agent workflows that can see screenshots, reason about what’s on screen, and connect to external tools through the Model Context Protocol (MCP) to act on what they find. The course closes with a deployment readiness review: your CTO wants to launch the AI pipeline next week, and you need to decide whether it’s safe to ship — with safety filtering, adversarial testing, and documented failure modes backing your recommendation.

Навыки, которые вы освоите

Retrieval-Augmented GenerationAgentic WorkflowsModel DeploymentComputer VisionResponsible AIMultimodal PromptsLLM ApplicationLarge Language ModelingFine-tuningPrompt EngineeringAI WorkflowsAgentic systemsGenerative AI AgentsImage AnalysisHugging FaceGenerative AIVision Transformer (ViT)Model Context ProtocolAI Security

Программа курса

4 модулей · 32 учебных материалов

01Multimodal Foundations and VLMs8 материалов
Welcome: From Text-Only to Multimodal AIВидео2,000 Products, No Descriptions — Can AI See What’s in the Photos?DIALOGUEHow CLIP Aligns Images and Text in a Shared SpaceВидеоVisual Question Answering and Image Captioning with VLMsВидео

Учитесь у экспертов

Hugging Face

Преподаватель курса

Introduction to Multimodal AI with Hugging Face
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 6.7 ч

4 модулей

Язык: Английский

Часть программы вашего университета
Multimodal Models and Hub Navigation ReferenceЧтение
Document AI — OCR, Layout Parsing, and Structured ExtractionВидео
Caption Products and Extract Invoice Data for BrightCartЛабораторная
Practice Assignment: Multimodal Foundations and VLMsЗадание
02Audio, Generation, and Adaptation Strategies7 материалов
The Captions Are Wrong for 30% of Products — Fine-Tune or Retrieve?DIALOGUETranscribing Audio with WhisperВидеоGenerating Images from Text with DiffusersВидеоAudio, Diffusers, and Adaptation Strategies ReferenceЧтениеWhen to Fine-Tune vs. When to Retrieve — LoRA and Multimodal RAGВидеоTranscribe Calls and Generate Visual Summaries for BrightCartЛабораторнаяPractice Assignment: Audio, Generation, and Adaptation StrategiesЗадание
03Agents, MCP, and Tool Use7 материалов
The Catalog Update Is Manual — Can an Agent Do It?DIALOGUEBuilding Your First Agent with smolagentsВидеоConnecting Agents to External Tools via MCPВидеоSmolagents, MCP, and Agent Design Patterns ReferenceЧтениеVision-Powered Agents — Screenshot, Reason, Act, IterateВидеоBuild an Agent That Automates BrightCart’s Catalog WorkflowЛабораторнаяPractice Assignment: Agents, MCP, and Tool UseЗадание
04Responsible Deployment10 материалов
Multimodal Safety Risks — What Can Go Wrong and WhyВидеоFiltering with ShieldGemma 2 — Input and Output SafetyВидеоMultimodal Safety and Responsible Deployment ReferenceЧтениеTesting Against Adversarial Inputs and Documenting Failure ModesВидеоWrap BrightCart’s VLM Pipeline with Safety FilteringЛабораторнаяIs BrightCart’s AI Pipeline Safe to Ship?DIALOGUEPractice Assignment: Responsible DeploymentЗаданиеFinal Assessment: Introduction to Multimodal AI with HF ЗаданиеWhat You Can See, Build, and Ship SafelyВидеоApplying Your Multimodal AI and Agent SkillsЧтение