All drills

Vision Transformers

What you'll be able to do

Apply the transformer architecture to images — the modern alternative to CNNs powering state-of-the-art vision.

Google multimodal models, HuggingFace ViT/CLIP/DINOv2 checkpoints, Meta DINOv2, and vision encoders inside multimodal LLMs.
Start this internship
Create an account to unlock the 5 sections, the workbench, and AskThili.
Begin

Sections

1. Vision Transformers
🔒 locked
2. Introduction to Vision Transformers
🔒 locked
3. Programming ViT Model
🔒 locked
4. Using Pretrained ViTs (with Hugging Face)
🔒 locked
5. Fine Tuning ViT
🔒 locked

Dig deeper

📄An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (Dosovitskiy et al., 2020)
paper
🔗Fine-Tune ViT for Image Classification — Hugging Face blog
article

Part of these learning paths

I'm a software engineer and I want to specialize in computer vision
View path →