Efficient VLM (22)

  • Efficient VLM — Overview

    Cutting visual tokens to make VLMs cheaper — mapped on two axes (where: encoder·bridge·LLM, and what criterion: importance·diversity·duplication·sensitivity·spatial), the field's evolution, and a one-glance table of 20 methods.

    image

Key papers

View all 22 Efficient VLM notes →

VLM (24)

  • VLM — Overview

    How vision-language / multimodal LLMs are built — the 5-component architecture (encoder · projector · LLM · output projector · generator), the understanding-vs-generation taxonomy, and key models from Flamingo to Qwen2.5-VL.

    image

Key papers

View all 24 VLM notes →

Token Reduction in ViTs (18)

Key papers

View all 18 Token Reduction in ViTs notes →