Tag
Aphanta is an automated diagnostic framework that evaluates the utility of image-edited intermediates in multimodal reasoning pipelines, showing task-dependent benefits and improving performance on specific tasks.
This paper introduces a framework for interactive task alignment under ambiguity, formalized as a POMDP, and shows that current LLMs recover intended tasks only 22–32% of the time, lagging behind human performance.