К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Multimodal Agents with Vision Language Models · LearnSpace
Назад в каталог
courseraАнализ данных

Multimodal Agents with Vision Language Models

Курс от Edureka
Средний≈ 5.9 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

This course covers vision-language models and multi-agent coordination: how agents interpret images alongside text and divide work between them. Together they take an agent beyond single inputs. You explore how vision-language models such as CLIP, BLIP, and LLaVA align images with text, and use those embeddings to build a multimodal search system. You then construct a multimodal RAG pipeline that answers questions about charts and tables inside PDF documents, where text-only retrieval fails. The course closes with orchestration: planner, executor, and critic patterns that break complex tasks into steps, reliable tool schemas with error handling, and collaborative multi-agent systems where specialized agents pass context between each other. By the end of this course, you will be able to: - Explain how vision-language models align image and text representations. - Build a multimodal search system using shared embedding spaces. - Construct a multimodal RAG pipeline over documents containing charts. - Apply planner, executor, and critic patterns to decompose complex tasks. - Implement tool calling with reliable schemas and error handling. - Coordinate multiple agents through defined roles, handoffs, and shared context. Intended for learners who have completed Multimodal AI and Agent Fundamentals. Enroll now to give your agents sight, retrieval, and the ability to work as a team.

Навыки, которые вы освоите

AI OrchestrationPython ProgrammingGenerative AI AgentsAgentic systemsPrompt EngineeringNatural Language ProcessingComputer VisionImage AnalysisContext EngineeringLLM ApplicationDocument ManagementLarge Language ModelingAPI Design

Программа курса

3 модулей · 40 учебных материалов

01Vision-Language Foundations and Multimodal RAG 16 материалов
Specialization OverviewВидеоCourse IntroductionВидеоCourse SyllabusЧтениеHow Vision-Language Models "See"Видео

Учитесь у экспертов

Edureka

Преподаватель курса

Multimodal Agents with Vision Language Models
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 5.9 ч

3 модулей

Язык: Английский

Часть программы вашего университета
Vision-Language Model Comparison ReferenceЧтение
Why Retrieval Matters for AgentsВидео
What is Retrieval-Augmented Generation (RAG)?Видео
Demonstration: Zero-Shot Image ClassificationВидео
Demonstration: Understanding EmbeddingsВидео
Applying Multimodal RAG ConceptsЗадание
Demonstration: Building a Basic Multimodal Search SystemВидео
Choose the Right Retrieval Strategy for a Multimodal TaskDIALOGUE
Demonstration: Building a Multimodal RAG Pipeline for Document Q&AВидео
Demonstration: Querying PDF Charts with Multimodal RAGВидео
Demonstration: Comparing Text-Only and Multimodal OutputsВидео
Knowledge Check: Vision-Language Foundations and Multimodal RAG Задание
02Agent Architectures and Tool Integration11 материалов
Agent Design PatternsВидеоAgent Frameworks OverviewЧтениеDesigning Reliable Tool SchemasВидеоAnatomy of a Reliable Tool-Calling AgentЧтениеDemonstration: Building a Planner-Executor AgentВидеоApplying Agent Architecture and Tool IntegrationЗаданиеDemonstration: Orchestrating Multiple ToolsВидеоDesign a Reliable Tool-Using AgentDIALOGUEDemonstration: Building a Web Summarization AgentВидеоDemonstration: Creating an Image-to-Image Design AgentВидеоKnowledge Check: Agent Architectures and Tool IntegrationЗадание
03Multi-Agent Collaboration13 материалов
Multi-Agent Handoff PatternsВидеоMulti-Agent Architecture Patterns ReferenceЧтениеDemonstration: Building a Two-Agent Collaboration SystemВидеоDemonstration: Building a Collaborative "Marketing Team" with AI AgentsВидеоPlan a Multi-Agent Collaboration WorkflowDIALOGUEDemonstration: Debugging a Multi-Agent ConversationВидеоApplying Multi-Agent CollaborationЗаданиеManaging Context and Memory Across Multiple AgentsЧтениеDemonstration: Multimodal Multi-Agent PipelineВидеоDesign a Reliable Multimodal Agent ArchitectureDIALOGUEPractice Project: Designing a Multi-Agent AI Research and Recommendation System for InsightSphereЧтениеKnowledge Check: Multimodal Agents with Vision Language ModelsЗаданиеCourse SummaryВидео