Tag
This paper analyzes the evaluation and design paradigms in deep reinforcement learning, revealing that performance rankings are not monotonic across data regimes and that common low-data regime benchmarks may lead to incorrect conclusions.
A scientist published a peer-reviewed critique in Nature arguing that Microsoft's quantum computing breakthrough claims are invalid due to basic Python errors and omitted data, suggesting the company's topological quantum computer is far from realization.
A new benchmark, FML-Bench, reveals that recent improvements in MLE-Bench scores are largely due to better base models and increased search budget rather than algorithmic advances.
This paper challenges the claim that prediction bottlenecks in models like Mamba recover causal structure, demonstrating through a new benchmark that gains are largely due to confounds and robustness artifacts rather than true causal discovery.
An undergraduate researcher expresses disillusionment with recent mechanistic interpretability research from Anthropic, specifically criticizing their new natural language autoencoder approach as a black-box technique that lacks rigorous metric comparisons against sparse autoencoder baselines.