标签
This paper presents a multiscale reward hedging method for learning from correct demonstrations, extending guarantees to continuous reward classes with a horizon-free bound via metric entropy, and shows polynomial-time cases for specific settings.
本文证明了一个关于无界权重双层神经网络的猜想鲁棒性定律,表明对于连续分段线性激活函数(如ReLU),拟合含噪数据的网络其Lipschitz常数至少为√(n/m)量级(最多差一个对数因子)。