2026

3 notes from 2026

  • Efficient VLM — Overview

    Cutting visual tokens to make VLMs cheaper — mapped on two axes (where: encoder·bridge·LLM, and what criterion: importance·diversity·duplication·sensitivity·spatial), the field's evolution, and a one-glance table of 20 methods.

  • VLM — Overview

    How vision-language / multimodal LLMs are built — the 5-component architecture (encoder · projector · LLM · output projector · generator), the understanding-vs-generation taxonomy, and key models from Flamingo to Qwen2.5-VL.

  • Token Reduction in ViTs — Overview

    ViT token efficiency — Pruning · Merging · Pooling · Hybrid, with key papers at a glance.