Tag
WhatWorkedBench is a benchmark designed to measure the experimental understanding of AI agents by evaluating their accuracy in predicting outcomes after budgeted experimentation across various tasks and configurations.
This paper applies graph signal processing to analyze how LLMs internally represent numerical sequences during in-context learning, finding that attention-induced token graphs and hidden-state signals show systematic, context-dependent signatures related to input complexity.