Tag
This controlled study investigates whether open-weight LLMs can serve as acquisition policies for finite-pool materials optimization, comparing their performance with random selection and conventional Gaussian-process methods across multiple tasks.
The paper argues that detecting an average effect of an acquired LLM-derived signal is not the same as learning per-instance acquisition policies, and establishes a reward-SNR floor (ρ* ≈ 2.8/√N) below which offline routing is impossible. It introduces Structured Hypothesis Embeddings (SHE) and shows across three datasets that learned per-example acquisition collapses below this floor, recommending design-time regime gates instead.