Tag
AffordanceVLA introduces a unified framework using structured affordance forecasting as an intermediate representation to improve perception-action mapping in robotic manipulation, leveraging vision-language models and a Mixture-of-Transformer architecture.
AFUN proposes an affordance foundation model that predicts functional masks and 3D motion curves from RGB-D observations and language descriptions, enabling generalizable robot manipulation across diverse environments. The model outperforms baselines on multiple benchmarks and can be deployed for real-world tasks without fine-tuning.
This paper introduces MM-CreativityBench, a benchmark for evaluating creative tool use in large multimodal models under physically constrained environments, and proposes affordance-grounded alignment using Direct Preference Optimization to reduce hallucination and improve grounded reasoning.
This extended paper revisits Semantic Web Services insights for Knowledge Graphs, proposing a four-dimensional formal framework and an Agentic Affordance Profile (AAP) to enable principled KG selection, composition, and failure diagnosis at agent planning time.
The paper introduces CreativityBench, a benchmark for evaluating large language models' ability to creatively repurpose tools based on affordance reasoning. It highlights that current models struggle with creative problem-solving despite strong general reasoning capabilities.