All drills

Learn Fast Neural Network Quantization

What you'll be able to do

Make models smaller and faster with little accuracy loss — essential for mobile, edge, and cost-efficient serving.

TensorFlow Lite / PyTorch ExecuTorch, NVIDIA TensorRT (INT8/FP8), Qualcomm SNPE, Apple Core ML, and llama.cpp/GGUF.
Start this internship
Create an account to unlock the 7 sections, the workbench, and AskThili.
Begin

Sections

1. Learn Fast Neural Network Quantization
🔒 locked
2. Understanding Neural Network Size and Parameters
🔒 locked
3. Mapping Function
🔒 locked
4. Post Training Dynamic Weights-Only Quantization
🔒 locked
5. Calibration
🔒 locked
6. Post Training Static Quantization (PTQ)
🔒 locked
7. Quantization Aware Training (QAT)
🔒 locked

Dig deeper

📄Quantization for Efficient Integer-Arithmetic-Only Inference (Jacob et al., 2017)
paper
🔗Practical Quantization in PyTorch — PyTorch blog
article

Part of these learning paths

I want to become an ML engineer who ships and operates models in production
View path →
I'm a software engineer and I want to specialize in computer vision
View path →