标签
WhatWorkedBench是一个基准测试,旨在通过评估AI代理在各种任务和配置中进行预算实验后预测结果的准确性来衡量其对实验的理解。
This paper applies graph signal processing to analyze how LLMs internally represent numerical sequences during in-context learning, finding that attention-induced token graphs and hidden-state signals show systematic, context-dependent signatures related to input complexity.