К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Building Reliable LLM Systems · LearnSpace
Назад в каталог
courseraАнализ данных

Building Reliable LLM Systems

Курс от Coursera
Средний≈ 20.1 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

Building Reliable LLM Systems is a comprehensive course for AI practitioners looking to move beyond basic models and create production-grade applications. While getting an LLM to generate text is easy, ensuring a consistently accurate, relevant, and trustworthy output is a significant engineering challenge. This course provides a systematic framework for tackling the entire lifecycle of LLM reliability. You will start by learning to quantitatively evaluate model performance using a suite of lexical and semantic metrics, such as BLEU, ROUGE-L, and cosine similarity. You’ll dive deep into debugging, using log analysis and data manipulation to uncover the root causes of critical failures, such as hallucinations, by correlating them with retrieval system performance. The course emphasizes statistical rigor, teaching you to design and analyze A/B tests, apply hypothesis testing, and calculate confidence intervals to prove the significance of your optimizations. Finally, you’ll optimize the foundational data layers, learning to tune SQL queries and vector search parameters to achieve the perfect balance between recall and latency.

Навыки, которые вы освоите

Model EvaluationPerformance TestingData-Driven Decision-MakingStatistical MethodsPerformance TuningSQLStatistical AnalysisRetrieval-Augmented GenerationLLM ApplicationVector DatabasesLarge Language ModelingMLOps (Machine Learning Operations)Python ProgrammingStatistical Hypothesis TestingQuery LanguagesArtificial Intelligence and Machine Learning (AI/ML)Debugging

Программа курса

5 модулей · 70 учебных материалов

01Evaluate and Optimize LLM Performance19 материалов
Welcome: Why Does "Good" AI Content Go Bad?DIALOGUEA Guide to LLM Evaluation: Lexical and Semantic MetricsЧтениеHow to Compute Lexical Metrics: BLEU & ROUGE-L in Python?ВидеоHow to Compute Semantic Similarity with Embeddings?Видео

Учитесь у экспертов

Professionals from the Industry

Преподаватель курса

Building Reliable LLM Systems
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 20.1 ч

5 модулей

Язык: Английский

Субтитры: Арабский, Французский, Итальянский, Бразильский португальский, Корейский, Немецкий, Испанский, Японский

Часть программы вашего университета
Knowledge Check: Choosing Your MetricsЗадание
Building Your First Automated Evaluation ScriptЛабораторная
Why Guess When You Can Know? The Case of the "Better" PromptВидео
The Language of Experimentation: Hypotheses, P-Values, and PowerВидео
Designing a Fair Race: A/B Testing for LLMsЧтение
Running the Numbers: A/B Test Analysis in PythonВидео
Statistical Significance TestingЛабораторная
Knowledge Check: Statistical Testing ConceptsЗадание
Coach Dialogue: Interpreting Your A/B Test ResultsDIALOGUE
From Report to Action: The Optimization LoopВидео
Case Study: Benchmarking a Sentiment AnalyzerВидео
Building a Reproducible Evaluation WorkflowЧтение
Scripting Your First Evaluation ReportВидео
Planning Your Optimization StrategyЛабораторная
Build Your LLM Evaluation ToolkitЗадание
02Analyze Logs: Fix LLM Hallucinations16 материалов
The Rogue ChatbotDIALOGUEWhy Logs Matter: The Air Canada Case?ВидеоAnatomy of a Log FileЧтениеCalculating Retention in PandasВидеоLab 1: Segmenting Users & Finding the DropЛабораторнаяReporting the ProblemDIALOGUEKnowledge Check: Retention MetricsЗаданиеWhy RAG Fails: The Root of Hallucination?ВидеоCorrelating Errors with Retrieval in PandasВидеоLab 2: Proving the Root CauseЛабораторнаяStating Your ConclusionDIALOGUEThe Engineering Brief: From Analysis to ActionЧтениеVisualizing the Proof in MatplotlibВидеоAuthoring the Engineering BriefЧтениеKnowledge Check: Communicating FindingsЗаданиеLLM Diagnostics ReportЗадание
03Evaluate LLMs: Test and Prove Significance16 материалов
Making a High-Stakes DecisionDIALOGUEWhy Single Scores LieВидеоCore Concepts: Confidence and SignificanceЧтениеCalculating Wilson Intervals in PythonВидеоLab 1: Quantifying Model AccuracyЛабораторнаяInterpret Your FindingsDIALOGUEConfidence Intervals QuizЗаданиеWhy Gut Feelings Fail in A/B TestingВидеоRunning a Chi-Square Test in PythonВидеоLab 2: Validating a Model ImprovementЛабораторнаяMaking the CallDIALOGUEStorytelling with Statistical VisualsЧтениеVisualizing Confidence with MatplotlibВидеоLab 3: Create a Comparison ChartЛабораторнаяCommunicating Results QuizЗаданиеLLM Evaluation ReportЗадание
04Optimize SQL and Vector Search Parameters16 материалов
Your First Performance TriageDIALOGUESecure and Efficient Query PatternsЧтениеFrom Inefficient to OptimizedВидеоIdentifying Slowest Queries using Parameterized SQLЛабораторнаяSQL Security and PatternsЗаданиеThe Recall vs. Latency Trade-OffВидеоUnderstanding Vector Search ParametersЧтениеTuning an HNSW IndexВидеоTune HNSW Parameters for Recall and LatencyЛабораторнаяParameter Tuning Scenarios QuizЗаданиеBeyond One-Off Tests: The Need for Continuous BenchmarkingВидеоCore Metrics of a Benchmarking FrameworkЧтениеBuilding a Benchmarking ScriptDIALOGUECreate an Automated Benchmarking SuiteЛабораторнаяInterpreting Benchmark ResultsЗаданиеSubmit Your Performance Optimization ReportЗадание
05End-to-End LLM Performance Audit 3 материалов
The Final Say: From Data to DecisionЧтениеYour Mission: The Performance AuditЧтениеProject: End-to-End LLM Performance Audit Задание