Tag
The paper proposes a densing law for user representation learning that quantifies the relationship between data scale and tokenization capacity, and introduces an adaptive tokenization method ALGN to improve efficiency in billion-scale scenarios.