Tag
This technical report explores adding lightweight depthwise convolution to the query/key/value projections in Transformer blocks for LLMs, providing local inductive bias that improves downstream accuracy with negligible parameter cost.