Tag
This paper introduces a dataset of 29,870 walkability ratings from 1,196 respondents and proposes a user-conditioned multimodal deep learning framework that fuses visual features with individual rater attributes to capture subjective variability in walkability perception. The model improves rank agreement by 65% over an image-only baseline, showing that who evaluates an environment matters.
This paper studies how persona prompting influences language generated by multimodal large language models in urban perception, finding that captions converge while justifications vary systematically with persona attributes.