What you'll be able to do- ✓Build a Vision Transformer
- ✓Classify images with transformers
- ✓Fine-tune pretrained ViTs
Apply the transformer architecture to images — the modern alternative to CNNs powering state-of-the-art vision.
⌁ Google multimodal models, HuggingFace ViT/CLIP/DINOv2 checkpoints, Meta DINOv2, and vision encoders inside multimodal LLMs.