Tag
This technical report explores adding lightweight depthwise convolution to the query/key/value projections in Transformer blocks for LLMs, providing local inductive bias that improves downstream accuracy with negligible parameter cost.
A detailed technical explanation of how CPU caches work, covering the principle of locality, cache organization, indexing, and handling writes.
This paper shows that state-of-the-art vision-language models fail at length generalization in visual reasoning due to 'global shortcuts', and demonstrates that combining local foveated perception with recurrence enables robust out-of-distribution generalization.
This paper studies the mismatch between sequence locality and attention-graph reachability in fixed block-sparse causal attention, formalizing boundary artifacts and proposing diagnostic coverage functions and a minimal repair called Boundary Bridge Attention.