Tag
Strauss Zelnick explains that AI is limited by backward-looking data and can reproduce the known but not create breakthroughs, placing value on human decisions about what to build.
Introduces ForeSci, a temporally controlled benchmark for evaluating whether LLM agents can make forward-looking research judgments from historical evidence. It contains 500 tasks across four AI domains and shows that explicit evidence organization improves traceability but reveals a recurring evidence-decision decoupling.