2021
3 notes from 2021
-
[TokenLearner] What Can 8 Learned Tokens Do for Images and Videos?
Learns a handful of adaptive tokens instead of a dense uniform grid.
-
[DynamicViT] Efficient Vision Transformers with Dynamic Token Sparsification
Dynamically drops redundant tokens per input to speed up ViTs.
-
[CLIP] Learning Transferable Visual Models From Natural Language Supervision
Trains an image encoder and a text encoder to match images with their captions (contrastive) on 400M web pairs — enabling open-vocabulary zero-shot transfer, and becoming the vision encoder most VLMs freeze and reuse.


