Tag
This paper identifies aggregation-induced reward hacking in multi-reward reinforcement learning for LLMs and proposes an adaptive projection method, AMRP, to dynamically adjust weights for better reward balance and performance.
This paper proposes LA-ReduNet, a lightweight adaptive architecture that uses hyperspherical manifold learning and adaptive step sizes to significantly reduce the number of layers needed for the MCR2 objective in neural networks, achieving comparable classification accuracy with far fewer parameters.
This paper introduces an adaptive on-the-fly multifidelity machine learning algorithm for quantum chemistry that autonomously determines training data composition across fidelities, reducing data generation costs by up to 30x compared to single-fidelity methods and up to 5x compared to standard multifidelity methods.