← All drillsBuild a Model-Serving Engine
What you'll be able to do- ✓Build a model-serving engine from scratch — the prediction API, batching queue, and cache behind a production inference service
- ✓Measure a server honestly with p50/p99 latency and throughput under real concurrency, not a single-request stopwatch
- ✓Implement adaptive batching that trades tail latency for throughput against an explicit SLO
- ✓Add a prediction cache, measure its hit rate, and recognise the requests that must never be cached
- ✓Serve two model versions side by side and shift traffic with shadow and canary routing
- ✓Keep a server alive under overload with queue limits, timeouts, and load shedding
- ✓Produce a capacity plan — the requests per second one box sustains at a target p99, and the cost per 1,000 predictions
Build the inference layer every ML product runs on — the service that turns a trained model into an API that answers thousands of requests a second without falling over, and the measurements that prove it.
⌁ Every company serving ML in production runs this layer — Netflix's recommendation services, Uber's Michelangelo, NVIDIA Triton, and Ray Serve are all the pieces you build here, industrialized.
Start this internshipCreate an account to unlock the 10 sections, the workbench, and AskThili.
BeginSections
1. Build a Model-Serving Engine
🔒 locked2. Lesson 1 - The honest baseline
🔒 locked3. Lesson 2 - Where the time actually goes
🔒 locked4. Lesson 3 - Adaptive batching
🔒 locked5. Lesson 4 - The prediction cache
🔒 locked6. Lesson 5 - When the input is an image
🔒 locked7. Lesson 6 - Versions, shadows, and canaries
🔒 locked8. Lesson 7 - Overload
🔒 locked9. Lesson 8 - The capacity plan
🔒 locked10. Lesson 9 - Ship it, and the landscape
🔒 lockedDig deeper
📄Clipper: A Low-Latency Online Prediction Serving System (Crankshaw et al., 2017, NSDI, arXiv:1612.03079)
paper📄Hidden Technical Debt in Machine Learning Systems (Sculley et al., 2015, NeurIPS)
paper🔗NVIDIA Triton Inference Server — dynamic batching and model management
docsPart of these learning paths
I want to become an ML engineer who ships and operates models in production
View path →