
Model Optimization, Inference & End-to-End Engineering
Ends soon! Get $70+ in savings and build skills with Coursera Plus. Save 40% for 3 months.

Model Optimization, Inference & End-to-End Engineering
This course is part of Microsoft Deep Learning Engineering with Azure Professional Certificate

Instructor: Microsoft
Included with Learn more
Recommended experience
What you'll learn
Apply post-training quantization, pruning, and knowledge distillation to compress models and benchmark accuracy-latency trade-offs.
Configure ONNX Runtime with CUDA and TensorRT execution providers to accelerate inference across hardware targets.
Deploy containerized models to Azure ML online and batch endpoints using autoscaling and blue/green deployment patterns.
Architect and document a complete deep learning engineering lifecycle from distributed training through production deployment.
Details to know

Add to your LinkedIn profile
See how employees at top companies are mastering in-demand skills

Build your Machine Learning expertise
- Learn new concepts from industry experts
- Gain a foundational understanding of a subject or tool
- Develop job-relevant skills with hands-on projects
- Earn a shareable career certificate from Microsoft

Explore more from Machine Learning
Why people choose Coursera for their career

Felipe M.

Jennifer J.

Larry W.








