Курс от EdurekaThis course covers the fundamentals of multimodal AI agents: what separates an agent from an assistant, and how agents handle text, image, audio, video, and document input. These foundations carry through the Specialization. You start with how large language models work, covering tokens, context windows, and their limitations, then examine the agent lifecycle from perception through reasoning to action. You then process each modality in turn: generating and structuring text, captioning and reasoning over images, transcribing audio, summarizing video, and extracting tables from PDF documents. The course closes by combining prompt design, context handling, memory, and tool calling into your first working multimodal agent. By the end of this course, you will be able to: - Distinguish AI assistants from AI agents and describe the agent lifecycle. - Explain how LLMs process tokens and context, and where they fail. - Process text, image, audio, video, and document inputs with multimodal models. - Design prompts and context handling that keep an agent coherent. - Implement short-term and long-term memory in an agent. - Build an agent that combines multimodal input with a tool call. Intended for Python developers, AI engineers, and data scientists. You should be able to write basic Python and work with APIs. Enroll now to build your first agent that handles more than plain text.
3 модулей · 41 учебных материалов

Преподаватель курса