К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Model Evaluation and Benchmarking · LearnSpace
Назад в каталог
courseraАнализ данных

Model Evaluation and Benchmarking

Курс от Coursera
Средний≈ 7.1 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

The Model Evaluation and Benchmarking course is designed for developers, engineers, and technical product builders who are new to Generative AI but already have intermediate machine learning knowledge, basic Python proficiency, and familiarity with development environments such as VS Code, and who want to engineer, customize, and deploy open generative AI solutions while avoiding vendor lock-in. The course equips learners with the skills to assess and compare the performance of both text and image generative models. Starting with text evaluation, learners apply standard metrics such as perplexity, BLEU (Bilingual Evaluation Understudy), ROUGE (Recall-Oriented Understudy for Gisting Evaluation), and BERTScore, while also designing human evaluation protocols and task-specific methods for applications like summarization or translation. The course then explores image evaluation using technical metrics, including FID (Fréchet Inception Distance), CLIP similarity (Contrastive Language–Image Pretraining similarity), and SSIM (Structural Similarity Index Measure), alongside human perception-based assessment techniques and artifact detection systems. In the final module, learners design comprehensive benchmarking frameworks with reproducible testing environments, version control, and visualization dashboards for continuous monitoring. By the end, learners will be able to implement automated, domain-specific evaluation systems and deliver detailed performance reports that ensure generative models meet rigorous quality standards.

Навыки, которые вы освоите

Model EvaluationData VisualizationContinuous MonitoringGenerative AIEmbeddingsImage QualityDashboard CreationImage AnalysisDashboardLarge Language Modeling

Программа курса

3 модулей · 20 учебных материалов

01Text Generation Metrics and Tools8 материалов
Podcast: The Problems Text Metrics Were Built to SolveВидеоCode Demonstration TranscriptsЧтениеYour Essential Toolkit: Metrics for Text EvaluationЧтениеYour First Evaluation Pipeline with Hugging FaceВидео

Учитесь у экспертов

Professionals from the Industry

Преподаватель курса

Model Evaluation and Benchmarking
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 7.1 ч

3 модулей

Язык: Английский

Субтитры: Арабский, Французский, Итальянский, Бразильский португальский, Корейский, Немецкий, Пушту, Индонезийский, Испанский, Дари, Японский

Часть программы вашего университета
Advanced Evaluation: Human Feedback and Comprehensive ReportingВидео
Why Statistical Testing MattersВидео
Run Your First Text Model EvaluationЛабораторная
Choosing the Best Metric for the TaskЗадание
02Image Quality Assessment Methods6 материалов
Podcast: The Hidden Problems Image Metrics RevealВидеоThe Must-Know Metrics for Image QualityЧтениеEvaluating & Automating Image Quality with TorchMetricsВидеоAdvanced Image Quality: FID, CLIP & Automated GatesВидеоRun Your First Image Model EvaluationЛабораторнаяWhat Do the Metrics Really Tell You?DIALOGUE
03Creating Benchmarking Frameworks6 материалов
Podcast: The Value of Benchmarks in AI WorkflowsВидеоHow to Design Benchmarks That MatterЧтениеTurning Model Outputs into Meaningful ComparisonsВидеоRun a Mini-BenchmarkЛабораторнаяEnd-to-End Benchmarking CheckЗаданиеPodcast: Bringing It All Together: Benchmarking That Builds TrustВидео