Efficient VLM (22)
-
Efficient VLM — Overview
Cutting visual tokens to make VLMs cheaper — mapped on two axes (where: encoder·bridge·LLM, and what criterion: importance·diversity·duplication·sensitivity·spatial), the field's evolution, and a one-glance table of 20 methods.

Key papers
View all 22 Efficient VLM notes →
VLM (24)
-
VLM — Overview
How vision-language / multimodal LLMs are built — the 5-component architecture (encoder · projector · LLM · output projector · generator), the understanding-vs-generation taxonomy, and key models from Flamingo to Qwen2.5-VL.

Key papers
Token Reduction in ViTs (18)
-
Token Reduction in ViTs — Overview
ViT token efficiency — Pruning · Merging · Pooling · Hybrid, with key papers at a glance.

Key papers