К содержимому
learnspaceYOUR NEXT CHAPTER
ПРОСТРАНСТВО ОБУЧЕНИЯ
ГлавнаяКаталог курсовМоё обучениеCoursera

Знания без границ

Учитесь у лучших университетов и компаний мира.

Открыть Coursera
Интеграция
Пространство университета
Моё пространствоСтраница курса
↵
ЯЛичный кабинетСтудент
© 2026 LearnSpaceКаждый день — возможность узнать больше.Помощь
Optimize AI Inference Speed & Accuracy · LearnSpace
Назад в каталог
courseraАнализ данных

Optimize AI Inference Speed & Accuracy

Курс от Coursera
Средний≈ 5.2 чАнглийский
О курсеНавыкиПрограммаПреподаватели

О курсе

Production ML models failing your latency targets? Learn how to make them run 3-5x faster without losing accuracy. This course helps ML engineers and data scientists optimize neural network inference for real-world deployment—across mobile, edge, and cloud environments. If you face slow model inference, high infrastructure costs, or deployment constraints, this course provides practical solutions. You'll master profiling techniques to identify performance bottlenecks, apply quantization to cut precision requirements, and make smart trade-offs between speed, accuracy, and resource constraints. You'll learn to benchmark optimization techniques and select the right approach for deployment scenarios. You'll explore inference profiling and metrics, pruning strategies, and quantization methods. You'll practice with real-world cases—from streaming platforms to autonomous vehicles—using industry-standard tools like PyTorch Profiler, TensorRT, and pruning utilities. This course is ideal for machine learning engineers, data scientists, and AI practitioners who are deploying or optimizing models in production. It’s also valuable for MLOps professionals and system engineers responsible for performance tuning in resource-constrained environments (e.g., mobile, embedded, or cloud inference systems). Learners should have a good grasp of Python and basic experience with PyTorch or TensorFlow. Familiarity with machine learning concepts, such as model training and evaluation, is expected. Understanding how neural networks work and basic performance metrics like latency and accuracy will help you get the most from this course. By the end of this course, you’ll confidently optimize production models, cut inference costs, meet latency goals, and deploy ML systems that scale efficiently.

Навыки, которые вы освоите

Model OptimizationCloud DeploymentModel DeploymentKeras (Neural Network Library)BenchmarkingModel EvaluationProject PerformanceAI SecurityNetwork ModelModel TrainingProcess Optimization

Программа курса

3 модулей · 23 учебных материалов

01Foundations: Profiling and Understanding Inference Bottlenecks8 материалов

Lesson 1: Foundations: Profiling and Understanding Inference Bottlenecks

Your First Day as ML Performance EngineerDIALOGUEWelcome to the Course: Course OverviewЧтениеCourse Intro: Optimize AI Inference Speed & AccuracyВидеоUnderstanding Inference BottlenecksВидео

Учитесь у экспертов

Starweaver

Global Leaders in Professional & Technology Education

Ritesh Vajariya

Advisor | Leader | Speaker |Author

Optimize AI Inference Speed & Accuracy
В каталоге вашей программы

Инвестируйте в себя

Новые знания — в удобное для вас время.

Начать на Coursera

Обучение откроется на Coursera
в новой вкладке

Обучение на Coursera

≈ 5.2 ч

3 модулей

Язык: Английский

Часть программы вашего университета
Profiling Tools in ActionВидео
Evaluating ML Inference Performance in ProductionВидео
Hands-On-Learning: Profile and Optimize Real-Time Fraud Detection SystemВзаимная проверка
NVIDIA Deep Learning Performance GuideЧтение
02Model Pruning: Reducing Complexity Without Losing Power6 материалов

Lesson 2: Model Pruning: Reducing Complexity Without Losing Power

The Pruning Dilemma at AutoDrive SystemsDIALOGUEPruning Theory and TechniquesВидеоImplementing Pruning in PyTorchВидеоFine-tuning and Recovery StrategiesВидеоHands-On-Learning: Prune and Deploy Mobile Image Classifier Under Size ConstraintsВзаимная проверкаThe Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural NetworksЧтение
03Quantization and Secure Deployment: Speed Meets Security9 материалов

Lesson 3: Quantization and Secure Deployment: Speed Meets Security

HealthAI's Security-Critical DecisionDIALOGUEQuantization FundamentalsВидеоImplementing Quantization WorkflowsВидеоBenchmarking: Pruning vs QuantizationВидеоHands-On-Learning: Optimize and Deploy Real-Time Video Analytics with QuantizationВзаимная проверкаAdversarial Robustness in Model CompressionЧтениеYour Optimization MasteryВидеоProject: Enterprise AI Inference Optimization: Production Deployment Under ConstraintsВзаимная проверкаOptimize AI Inference Speed & AccuracyЗадание