Tag
This paper proposes evaluating coding LLMs on their understanding of software execution beyond control flow, including predicting memory usage, runtime, and profiler outputs, finding that all tested models perform poorly, indicating a lack of deep software world model understanding.