All drills

Build a Model-Serving Engine

What you'll be able to do

Build the inference layer every ML product runs on — the service that turns a trained model into an API that answers thousands of requests a second without falling over, and the measurements that prove it.

Every company serving ML in production runs this layer — Netflix's recommendation services, Uber's Michelangelo, NVIDIA Triton, and Ray Serve are all the pieces you build here, industrialized.
Start this internship
Create an account to unlock the 10 sections, the workbench, and AskThili.
Begin

Sections

1. Build a Model-Serving Engine
🔒 locked
2. Lesson 1 - The honest baseline
🔒 locked
3. Lesson 2 - Where the time actually goes
🔒 locked
4. Lesson 3 - Adaptive batching
🔒 locked
5. Lesson 4 - The prediction cache
🔒 locked
6. Lesson 5 - When the input is an image
🔒 locked
7. Lesson 6 - Versions, shadows, and canaries
🔒 locked
8. Lesson 7 - Overload
🔒 locked
9. Lesson 8 - The capacity plan
🔒 locked
10. Lesson 9 - Ship it, and the landscape
🔒 locked

Dig deeper

📄Clipper: A Low-Latency Online Prediction Serving System (Crankshaw et al., 2017, NSDI, arXiv:1612.03079)
paper
📄Hidden Technical Debt in Machine Learning Systems (Sculley et al., 2015, NeurIPS)
paper
🔗NVIDIA Triton Inference Server — dynamic batching and model management
docs

Part of these learning paths

I want to become an ML engineer who ships and operates models in production
View path →