Tag
This paper introduces PAYN, a plug-and-play token compression method for MLLM-based referring expression segmentation that relies solely on position information, outperforming existing techniques by preserving spatial relational consistency.
This course note summarizes the architectural evolution from the original Transformer to modern LLMs, focusing on convergent developments such as pre-normalization, RMS normalization, and RoPE, and provides hyperparameter selection recommendations.