Tag
This technical report explores adding lightweight depthwise convolution to the query/key/value projections in Transformer blocks for LLMs, providing local inductive bias that improves downstream accuracy with negligible parameter cost.
Introduces ITNet, a neural architecture based on a learnable integral transform that unifies convolution, attention, and recurrence, achieving strong results across multiple modalities.