Tag
This paper presents a large-scale empirical study of offline reinforcement and imitation learning, analyzing over 160,000 training runs to understand the effects of hyperparameters and dataset properties, and introduces JumpStart, a resource suite for reliable policy-learning research.
This article benchmarks various LLMs on their ability to identify mushroom species from images using a large dataset, highlighting accuracy issues and safety concerns for real-world applications.
This paper investigates whether real-world datasets contain natural experiments by using causal discovery and feature selection, finding that they do and can improve model performance.