Walkable to Whom? Capturing Subjective Variability in Walkability Perception Using Multimodal Deep Learning
Summary
This paper introduces a dataset of 29,870 walkability ratings from 1,196 respondents and proposes a user-conditioned multimodal deep learning framework that fuses visual features with individual rater attributes to capture subjective variability in walkability perception. The model improves rank agreement by 65% over an image-only baseline, showing that who evaluates an environment matters.
View Cached Full Text
Cached at: 08/10/26, 08:04 AM
# Walkable to Whom? Capturing Subjective Variability in Walkability Perception Using Multimodal Deep Learning
Source: [https://arxiv.org/html/2608.06934](https://arxiv.org/html/2608.06934)
Moloud Damandeh[https://orcid.org/0009-0001-4344-9858](https://orcid.org/0009-0001-4344-9858), , Meead Saberi[https://orcid.org/0000-0002-6526-239X](https://orcid.org/0000-0002-6526-239X)This work was supported by the Australian Research Council under Grant DP220102382\. M\. Saberi was also supported by the Australian Research Council through Future Fellowship FT250100584\.This work involved human subjects\. Approval of all ethical procedures and protocols was granted by the University of New South Wales Human Research Ethics Committee under approval number 8361\.Data and code is available at https://github\.com/Moloudd/user\-conditioned\-walkability\-assessmentConflict of Interest: M\. Saberi is co\-founder and CEO of footpath\.ai, the imagery provider for this study, which had no role in study design or the decision to publish\.The authors are with the School of Civil and Environmental Engineering, University of New South Wales \(UNSW\), Sydney, NSW, Australia, and the Research Centre for Integrated Transport Innovation \(rCITI\)\.Corresponding author: Meead Saberi \(e\-mail: meead\.saberi@unsw\.edu\.au\)\.
###### Abstract
Visual perception of walkability varies substantially across individuals, reflecting differences in personal characteristics, experiences, and preferences\. Existing studies, however, often reduce these diverse judgements to aggregated scores, implicitly assuming uniform perception, and commonly rely on vehicle\-mounted street\-view imagery that does not reflect the pedestrian’s visual experience\. This paper introduces a dataset of 29,870 walkability ratings from 1,196 respondents, linking sidewalk\-view imagery across urban, suburban, and regional Australian environments with individual rater attributes, and proposes the first user\-conditioned multimodal deep learning framework for walkability perception, fusing visual features with respondent\-level representations\. A viewpoint\-comparison study shows that sidewalk\-view images receive significantly higher walkability ratings than matched street\-view images, indicating that imagery source is a substantive design decision in perception surveys\. The user\-conditioned model improves rank agreement with observed ratings by 65% over an image\-only baseline \(quadratic weighted kappa 0\.47 vs\. 0\.29\), demonstrating that who is evaluating an environment carries predictive indication beyond image content alone\. These findings support moving from aggregated, observer\-independent walkability scores toward models that represent diverse users, enabling more inclusive assessment of pedestrian environments\.
###### Index Terms:
Walkability perception, Sidewalk\-view imagery, Data\-based approaches \(learning, deep learning, reinforcement learning\), transportation planning and design, Pedestrian flows and crowds
## IIntroduction
Walkable neighbourhoods support stronger social interaction and economic activity within communities, making walkability an important consideration in urban planning and design\[[1](https://arxiv.org/html/2608.06934#bib.bib1),[2](https://arxiv.org/html/2608.06934#bib.bib2)\]\. Understanding how people perceive the walkability of their environment, and why those perceptions differ, is essential for designing streets that serve diverse communities\[[3](https://arxiv.org/html/2608.06934#bib.bib3)\]\. While visual characteristics of the built environment play a key role in shaping perceived walkability\[[4](https://arxiv.org/html/2608.06934#bib.bib4),[5](https://arxiv.org/html/2608.06934#bib.bib5)\], perception is not determined by environmental characteristics alone\[[3](https://arxiv.org/html/2608.06934#bib.bib3),[6](https://arxiv.org/html/2608.06934#bib.bib6)\]\. It is also shaped by user\-specific factors, including demographic characteristics, geographic context, socio\-cultural background, and mobility\-related needs\. As a result, the same environment may be perceived differently by different people\[[3](https://arxiv.org/html/2608.06934#bib.bib3),[6](https://arxiv.org/html/2608.06934#bib.bib6),[7](https://arxiv.org/html/2608.06934#bib.bib7)\], adding further complexity to the already difficult task of quantifying walkability\[[5](https://arxiv.org/html/2608.06934#bib.bib5),[7](https://arxiv.org/html/2608.06934#bib.bib7)\]\. While prior work has made substantial progress in modeling urban perception from visual features\[[11](https://arxiv.org/html/2608.06934#bib.bib11),[27](https://arxiv.org/html/2608.06934#bib.bib27),[14](https://arxiv.org/html/2608.06934#bib.bib14)\], predictive models that account for this subjective variability remain underexplored\.
Another limitation of existing image\-based methods for predicting walkability perception is their reliance on street\-view imagery captured from vehicle\-mounted cameras rather than from a pedestrian viewpoint\. Such imagery can miss important details relevant to pedestrians\. In contrast, sidewalk\-view imagery provides a more pedestrian\-centred perspective, reducing viewpoint bias and better capturing the elements that matter to active mobility users\[[8](https://arxiv.org/html/2608.06934#bib.bib8)\]\. However, whether this viewpoint mismatch affects perceived walkability ratings remains largely unexamined, even though recent studies have begun to establish protocols for image\-based urban perception surveys\[[9](https://arxiv.org/html/2608.06934#bib.bib9)\]\.
Alongside standardized survey protocols, progress in this area also depends on the availability of open datasets that can support reproducible benchmarking\[[4](https://arxiv.org/html/2608.06934#bib.bib4)\]\. Place Pulse 2\.0 remains one of the most widely used large\-scale public benchmarks for urban visual perception, providing high\-quality perception annotations but limited respondent attributes\[[11](https://arxiv.org/html/2608.06934#bib.bib11)\]\. More recently, SPECS has provided, to our knowledge, the only publicly available urban perception dataset that links perception ratings with richer respondent characteristics and analyses differences across demographic and personality groups\[[6](https://arxiv.org/html/2608.06934#bib.bib6)\]\. However, SPECS primarily supports group\-level analyses of broader urban perception indicators\. Publicly available datasets designed for individual\-level, user\-conditioned modeling of walkability perception remain scarce, particularly those pairing sidewalk\-view imagery with walkability ratings and rater\-specific attributes\.
Figure 1:Overview of the proposed dataset structure\. \(a\) Representative sidewalk\-view panoramic images sampled across urban, suburban, and regional environments\. \(b\) Example image–respondent records showing how each sidewalk\-view image is linked to walkability ratings and respondent\-level attributes\.This paper addresses these gaps through three main contributions\. First, it introduces a new sidewalk\-view walkability perception dataset linked with individual rater attributes\. Second, it demonstrates that image viewpoint affects perceived walkability, showing that the choice of visual data source should be considered when designing surveys for walkability perception studies\. Third, it proposes and evaluates a multimodal user\-conditioned deep learning framework that fuses visual features with rater\-level representations to model individual\-level variability in walkability perception\. By incorporating respondent attributes alongside image content, the framework demonstrates that subjective differences in perceived walkability can be learned and explained, providing a more personalized representation of visual walkability than image\-only approaches\. Beyond urban assessment and design, these user\-conditioned predictions could also support personalized pedestrian routing, where predicted street\-level walkability ratings enter route cost functions alongside travel time\. This would enable different users to receive routes that better reflect their own perceptions and preferences, analogous to multi\-criteria cycling routing approaches that incorporate subjective criteria such as comfort and safety\[[12](https://arxiv.org/html/2608.06934#bib.bib12)\]\.
## IIRELATED WORK
### II\-AVisual Walkability Assessment
Visual walkability assessment has increasingly used street\-level imagery and deep learning to capture physical features as experienced by pedestrians at street level\[[4](https://arxiv.org/html/2608.06934#bib.bib4)\]\. A dominant methodological approach in image\-based urban perception is crowdsourced pairwise comparison, where urban perception is formulated as a ranking problem\. This paradigm was established using the Place Pulse 2\.0 dataset, which contains 1\.17 million pairwise comparisons across 56 cities, and a Siamese Convolutional Neural Network \(CNN\) trained to predict which of two streetscapes was perceived as safer, livelier, or more beautiful\[[11](https://arxiv.org/html/2608.06934#bib.bib11)\]\. The same framework has since been adapted to walkability by collecting pairwise visual walkability ratings and training a multitask CNN to predict both overall walkability and contributing dimensions, including safety, comfort, and accessibility\[[5](https://arxiv.org/html/2608.06934#bib.bib5)\]\. A related comparison\-based approach has also been applied to cycling safety perception, where a Siamese CNN was trained using pairwise comparisons, and deployed for city\-wide assessment\[[27](https://arxiv.org/html/2608.06934#bib.bib27)\]\. Other studies have used direct rating rather than pairwise comparison\. For example, five\-point Likert\-scale ratings of Google Street View images \(SVI\) have been used to train a CNN to classify images into five ordinal walkability categories\[[14](https://arxiv.org/html/2608.06934#bib.bib14)\]\. However, these models characterise the visual environment independently of the observer, treating individual differences in perception as noise to be averaged out rather than as meaningful variation to be modeled\. In this study, walkability ratings were collected using a five\-point Likert scale\. Unlike pairwise comparison methods, which yield image\-level aggregate scores and require additional grouping to analyse demographic differences\[[6](https://arxiv.org/html/2608.06934#bib.bib6)\], direct ratings preserve respondent\-level scores for each image\. This is a structural requirement for the proposed user\-conditioned modelling framework\. The Likert scale also requires fewer ratings per image to obtain robust estimates\[[9](https://arxiv.org/html/2608.06934#bib.bib9)\], making it well\-suited to the scale of this study\.
### II\-BObserver\-Level Variability in Urban Perception
The concept of walkability often assumes that pedestrians form a homogeneous group, but empirical evidence increasingly challenges this assumption\[[7](https://arxiv.org/html/2608.06934#bib.bib7)\]\. Cross\-cultural studies show that although a broad understanding of walkability may be shared across groups, the specific built\-environment attributes that define a walkable place vary systematically with sociocultural background and residential history\[[7](https://arxiv.org/html/2608.06934#bib.bib7)\]\. At a broader scale, a large\-scale survey of participants across five countries found that streetscape perceptions vary significantly by sociodemographic attributes\[[6](https://arxiv.org/html/2608.06934#bib.bib6)\]\. The study further showed that models trained on aggregated responses can introduce systematic bias by masking this variation\. Within pedestrian planning, studies similarly show that the attributes most important to walkability differ substantially across user groups with infrastructure quality and physical accessibility being especially salient for older pedestrians\[[3](https://arxiv.org/html/2608.06934#bib.bib3)\]\. Despite this evidence, current planning approaches that account for user differences typically rely on predefined group typologies derived from expert knowledge, rather than learning observer\-specific differences from individual\-level perceptual data\[[3](https://arxiv.org/html/2608.06934#bib.bib3)\]\. At the same time, computational walkability models generally do not incorporate observer characteristics as inputs to prediction functions trained on empirical ratings\. Consequently, the individual\-level perceptual variability documented across these studies remains largely unrepresented in current predictive architectures\.
### II\-CPersonalized Subjective Image Assessment
The task of predicting subjective image scores for individual observers, rather than a single aggregate score, has recently gained attention in the image aesthetics assessment literature\[[15](https://arxiv.org/html/2608.06934#bib.bib15),[20](https://arxiv.org/html/2608.06934#bib.bib20)\]\. A key distinction in this field is between generic image aesthetic assessment, which predicts a mean opinion score across all raters, and personalized image aesthetic assessment, which predicts observer\-specific scores\[[20](https://arxiv.org/html/2608.06934#bib.bib20)\]\. Early personalized image aesthetic assessment methods primarily used image\-level attributes to adjust generic predictions toward individual preferences\[[28](https://arxiv.org/html/2608.06934#bib.bib28)\]\. More recent work incorporates personal attributes such as demographics and personality traits into prediction models\. The PARA dataset, for example, includes image ratings from 438 subjects along with attributes such as age, gender, education, and Big Five personality scores\[[15](https://arxiv.org/html/2608.06934#bib.bib15)\]\. Models trained on PARA show improved performance when conditioned on subject information\. Similarly, the LAPIS dataset provides rich annotator attributes for artistic images, and ablation studies show that removing certain attributes degrades prediction performance\[[20](https://arxiv.org/html/2608.06934#bib.bib20)\], confirming that observer characteristics carry predictive signal beyond image content alone\. In the transportation planning domain, existing walkability and urban perception models remain closer to the generic formulation\. They typically produce a single aggregate score per image or location, treating all observers as equivalent\. User\-conditioned prediction, in which explicit observer attributes are jointly encoded with image features, has therefore not yet been applied to walkability or pedestrian environment perception\. Learning individual user preferences has also been recognised in applications beyond visual perception, including personalised multimodal route planning systems, where user\-specific preferences are learned to improve route choice modelling\[[37](https://arxiv.org/html/2608.06934#bib.bib37)\]\. More broadly, multimodal fusion of heterogeneous data sources has been identified as an important direction within the intelligent transportation systems \(ITS\) deep learning literature\[[38](https://arxiv.org/html/2608.06934#bib.bib38)\]\. However, prior applications have predominantly focused on transportation forecasting, perception, and control\. The present study extends multimodal learning to pedestrian walkability perception by jointly encoding sidewalk\-view imagery and respondent\-level characteristics\.
### II\-DSidewalk\-View Imagery and View\-Point Bias
Street\-view imagery, also known as SVI, has become an important data source in urban research because it enables scalable assessment of street\-level built\-environment characteristics that are difficult to capture using conventional GIS\-based measures\[[8](https://arxiv.org/html/2608.06934#bib.bib8),[29](https://arxiv.org/html/2608.06934#bib.bib29)\]\. However, most existing SVI\-based studies rely on imagery collected from vehicle\-mounted cameras\[[35](https://arxiv.org/html/2608.06934#bib.bib35)\]\. This introduces a potential viewpoint bias, because vehicle\-based imagery represents the street from the carriageway rather than from the pedestrian path\. Recent studies have begun to quantify this limitation by comparing semantic information extracted from street\-view and sidewalk\-view imagery\. These studies report weak or inconsistent correspondence between the two viewpoints, suggesting that the visual features visible from the vehicle lane can differ substantially from those encountered on the sidewalk\[[8](https://arxiv.org/html/2608.06934#bib.bib8),[22](https://arxiv.org/html/2608.06934#bib.bib22)\]\. Despite growing evidence that viewpoint affects the visual representation of the street environment, existing image\-based walkability studies have largely relied on vehicle\-based SVI or static image views\. More importantly, little is known about whether people evaluate the same location differently when it is presented from a street\-view versus a sidewalk\-view perspective\. This leaves unresolved whether viewpoint differences translate into differences in perceived walkability for the same location\.
## IIIMethodology
### III\-ADataset
#### III\-A1Image Selection
Sidewalk\-view imagery was sourced from footpath\.ai111[https://footpath\.ai](https://footpath.ai/), a large\-scale provider of high\-resolution 360\-degree street\-level imagery\. The source dataset includes image\-level sidewalk and pedestrian infrastructure attributes relevant to walkability assessment, extracted using AI\-based semantic segmentation and object detection methods\. These attributes include greenery, sky, building coverage, shaded areas, sidewalk width, benches, garbage bins, marked crossings, and accessibility curb cuts\. They are key dimensions of the physical pedestrian environment\[[10](https://arxiv.org/html/2608.06934#bib.bib10)\]and were used to guide a representative image sampling strategy\.
To capture variability in walkability perception across settlement types, we selected nine Australian locations spanning urban, suburban, and regional contexts\. Representative examples of sidewalk\-view images across settlement types are shown in Fig\.[1](https://arxiv.org/html/2608.06934#S1.F1)\(a\)\. This sampling design extends beyond the city\-centre focus common in existing visual perception datasets\[[11](https://arxiv.org/html/2608.06934#bib.bib11),[6](https://arxiv.org/html/2608.06934#bib.bib6)\], none of which include suburban or regional settlement types\.
From the image pool across the selected settlement types, a representative subset was constructed using Threshold\-Constrained Stratified Sampling \(TCSS\)\. Unlike semantic clustering approaches that group images by dominant visual composition\[[9](https://arxiv.org/html/2608.06934#bib.bib9)\], TCSS stratifies images based on fine\-grained combinations of walkability\-relevant attributes\. This ensures that the sample preserves the real\-world distribution of pedestrian infrastructure conditions, including sidewalk width, tree presence, and crossing availability, rather than generic visual diversity\. The procedure first partitions the source pool into strata defined by unique combinations of the image\-level attributes described above\. Low\-frequency strata are then aggregated to reduce excessive fragmentation, and a binary search procedure is used to identify the minimum sample size satisfying a predefined distributional constraint, withδ=0\.05\\delta=0\.05\. Under this constraint, the marginal distribution of each walkability\-relevant attribute in the final sample deviates from the corresponding distribution in the source pool by no more than 5%\. This tolerance was selected to balance sample compactness and distributional fidelity\. The procedure yielded a final dataset of 974 images reflecting the diversity of pedestrian infrastructure conditions across the selected locations\.
#### III\-A2Online Survey
An online survey was designed and conducted to collect human perceptions of walkability for sidewalk\-view images\. The survey consisted of two components\. First, participants provided self\-reported information, including age, gender, current place of residence, childhood residential environment, walking frequency, country of upbringing, and disability or health condition affecting walking\. Second, each participant rated a block of 25 sidewalk\-view images, pre\-assigned to ensure a representative mix of urban, suburban, and regional environments\. Images were displayed as interactive 360\-degree panoramic views, allowing respondents to rotate and examine the sidewalk environment from multiple perspectives before submitting a rating on a five\-point Likert scale in response to the prompt, ”On a scale of 1 \(poor walkability\) to 5 \(excellent walkability\), how walkable does this environment appear to you?”
To support consistent interpretation while allowing subjective judgement, an information icon was displayed alongside each image with the following definition: “Walkability refers to how suitable an area is for walking based on its overall environment and conditions\. People may perceive walkability differently depending on their own experiences and expectations\.” This definition was intentionally broad, acknowledging that individuals may weigh environmental attributes differently depending on their residential background, walking habits, and mobility\-related needs\. This design is consistent with the objective of modeling individual\-level variability in walkability perception rather than imposing a fixed expert\-defined criterion\. Fig\.[1](https://arxiv.org/html/2608.06934#S1.F1)\(b\) illustrates the resulting record structure, in which each sidewalk\-view image is linked to an individual respondent’s attributes and their walkability rating\.
To ensure settlement\-type diversity within each participant’s rating block, images were assigned using a round\-robin procedure\. The image pool was first shuffled and grouped by settlement type and sampling stratum\. Images were then sequentially allocated across rating blocks, ensuring that each block included urban, suburban, and regional images while maintaining diversity across the sampling strata\.
Participants were recruited primarily through Pureprofile222[https://www\.pureprofile\.com](https://www.pureprofile.com/), a commercial online research panel provider, and were compensated upon completing the survey\. To broaden participation, additional voluntary responses were collected through professional and social networking channels\. Recruitment aimed to obtain a geographically and socio\-demographically diverse respondent pool across urban, suburban, and regional contexts, corresponding to the settlement\-type categories represented in the image dataset\.
Respondent reliability was assessed using multiple quality\-control criteria\. First, two randomly selected images were repeated within each survey session as consistency checks\. Respondents were excluded if their ratings for both repeated images were inconsistent with their original ratings\. Second, speeders were identified based on completion times below a minimum plausible threshold\. Third, straightliners were identified as respondents who submitted the same rating to all image\-rating items\. Additional survey controls were implemented to reduce automated and duplicate responses\[[13](https://arxiv.org/html/2608.06934#bib.bib13)\]\. All quality\-control filters were applied before analysis\. The final dataset comprised 1,196 valid participants\. The average survey completion time was approximately 10 minutes\.
### III\-BViewpoint Comparison Survey
Figure 2:Overview of the proposed user\-conditioned multimodal framework for walkability perception prediction\. The study evaluates alternative image backbones, tabular encoders, fusion mechanisms, and loss functions; the configuration shown corresponds to the best\-performing model\.To validate the use of sidewalk\-view imagery as a more appropriate data source for walkability perception modeling, a comparison study was conducted using matched sidewalk\-view and street\-view image pairs from 100 locations\. For each location, a sidewalk\-view image from the proposed dataset was paired with a corresponding Google SVI\. The selected locations were distributed evenly across the three settlement types represented in the main survey, ensuring coverage of urban, suburban, and regional environments\. Each participant rated 20 images, consisting of one sidewalk\-view and one street\-view image from each of 10 randomly assigned locations, using the same five\-point Likert scale as the main survey\. Participants were not informed of the viewpoint distinction in order to reduce priming effects\. Because ratings were paired by location, measured on an ordinal scale, and not normally distributed, a Wilcoxon signed\-rank test was used to assess whether viewpoint systematically influenced perceptions of walkability\.
### III\-CUser\-Conditioned Multimodal Deep Learning Framework for Walkability Perception
#### III\-C1Problem Formulation
LetIiI\_\{i\}denote a sidewalk\-view image and let𝐱j∈ℝd\\mathbf\{x\}\_\{j\}\\in\\mathbb\{R\}^\{d\}denote a respondent\-level attribute vector associated with participantjj\. The attribute vector encodes individual characteristics that may influence walkability perception, including demographic information, residential background, walking habits, and mobility\-related conditions\. For each image–respondent pair\(Ii,𝐱j\)\(I\_\{i\},\\mathbf\{x\}\_\{j\}\), the observed target is a perceived walkability ratingyij∈\{1,2,3,4,5\}y\_\{ij\}\\in\\\{1,2,3,4,5\\\}on a five\-point ordinal Likert scale\.
Motivated by evidence that walkability perception varies systematically with individual characteristics\[[6](https://arxiv.org/html/2608.06934#bib.bib6)\], and by the demonstrated effectiveness of conditioning visual assessments on observer\-level attributes to model individual differences in subjective perception\[[15](https://arxiv.org/html/2608.06934#bib.bib15)\], the objective is to learn a user\-conditioned prediction function:
y^ij=f\(Ii,𝐱j\),\\hat\{y\}\_\{ij\}=f\(I\_\{i\},\\mathbf\{x\}\_\{j\}\),\(1\)wherey^ij\\hat\{y\}\_\{ij\}represents the predicted walkability rating assigned by respondentjjto sidewalk imageIiI\_\{i\}\. This formulation differs from conventional image\-based approaches that assign a single aggregate score to each location independent of the observer\[[11](https://arxiv.org/html/2608.06934#bib.bib11),[14](https://arxiv.org/html/2608.06934#bib.bib14)\]\. Instead, it requires the model to jointly encode visual environmental cues and respondent\-level profile information, allowing perceived walkability to be modeled as an interaction between the sidewalk environment and the individual evaluating it\.
#### III\-C2Multimodal Representation Learning
The formulation requires learning from two heterogeneous sources of information that differ in both data format and semantic role\. The image describes the pedestrian environment being rated, whereas the respondent attributes describe the individual performing the rating\. Multimodal learning from heterogeneous sources, particularly the combination of visual and structured tabular data, commonly relies on modality\-specific encoders that project each input into a learned representation space before fusion\[[16](https://arxiv.org/html/2608.06934#bib.bib16)\]\.
Fig\.[2](https://arxiv.org/html/2608.06934#S3.F2)illustrates the proposed framework\. The image encodergI\(⋅\)g\_\{I\}\(\\cdot\)maps each sidewalk\-view image into a latent visual representation:
𝐳iI=gI\(Ii\),\\mathbf\{z\}^\{I\}\_\{i\}=g\_\{I\}\(I\_\{i\}\),\(2\)capturing pedestrian\-environment cues relevant to perceived walkability\. In parallel, the respondent encodergX\(⋅\)g\_\{X\}\(\\cdot\)maps the attribute vector into a latent user representation:
𝐳jX=gX\(𝐱j\),\\mathbf\{z\}^\{X\}\_\{j\}=g\_\{X\}\(\\mathbf\{x\}\_\{j\}\),\(3\)encoding individual\-level characteristics that condition the perception of the same environment\.
The two representations are combined through a learned fusion module:
𝐳ij=h\(𝐳iI,𝐳jX\),\\mathbf\{z\}\_\{ij\}=h\\left\(\\mathbf\{z\}^\{I\}\_\{i\},\\,\\mathbf\{z\}^\{X\}\_\{j\}\\right\),\(4\)whereh\(⋅,⋅\)h\(\\cdot,\\cdot\)is a learned function that combines the two modality\-specific representations, enabling the model to learn interactions between visual environmental features and respondent characteristics in a shared representation space before prediction\.
For the image encodergI\(⋅\)g\_\{I\}\(\\cdot\), we adopt swin\-tiny\[[31](https://arxiv.org/html/2608.06934#bib.bib31)\], a hierarchical vision Transformer based on shifted\-window self\-attention, pretrained on ImageNet\-1K\. Each 360° image is stored as an 2:1 equirectangular panorama, and processed according to the pretrained Swin\-Tiny input configuration: the shorter side is resized to 224 pixels, after which a centred 224×224 crop is used as the network input, retaining the central 180° horizontal field of view\. This follows prior work applying standard 2D vision architectures directly to panoramic imagery without geometric correction\[[5](https://arxiv.org/html/2608.06934#bib.bib5)\]\. For the respondent encodergX\(⋅\)g\_\{X\}\(\\cdot\), we employ FT\-Transformer\[[18](https://arxiv.org/html/2608.06934#bib.bib18)\], a transformer\-based model designed for tabular data\. The fusion moduleh\(⋅,⋅\)h\(\\cdot,\\cdot\)is implemented as a Transformer\-based module\[[34](https://arxiv.org/html/2608.06934#bib.bib34)\]projecting each modality’s representation to a common 768\-dimensional space, and processing the result through 3 Transformer blocks to obtain a 768\-dimensional joint embedding\. The justification for these specific choices is provided through ablation experiments in Section[IV\-C4](https://arxiv.org/html/2608.06934#S4.SS3.SSS4)\.
#### III\-C3Ordinal Prediction
The target variableyijy\_\{ij\}is an ordered categorical rating rather than a nominal class\. Treating this task as standard multiclass classification ignores the ordinal structure of the scale and treats all incorrect predictions as equally distinct, regardless of their distance on the rating scale\. This is undesirable for perceived walkability ratings, where adjacent categories may reflect minor differences in judgment, while larger ordinal deviations indicate more substantial disagreement between predicted and observed perception\.
To preserve the ordered structure of the response variable, the framework follows the ordinal regression formulation of CORAL\[[19](https://arxiv.org/html/2608.06934#bib.bib19)\], which decomposes aKK\-class ordinal prediction problem intoK−1K\-1binary threshold tasks while enforcing rank\-consistent predictions\. Rank consistency ensures that if the model predicts the rating exceeds thresholdkk, it must also predict that it exceeds all lower thresholdsk′<kk^\{\\prime\}<k\. For the five\-point walkability scale used in this study,K=5K=5, and each observed ratingyijy\_\{ij\}is transformed into a set of binary ordinal labels:
rij,k=𝟙\(yij\>k\),k∈\{1,2,3,4\}\.r\_\{ij,k\}=\\mathbbm\{1\}\(y\_\{ij\}\>k\),\\quad k\\in\\\{1,2,3,4\\\}\.\(5\)The model estimates the probability that the perceived walkability rating exceeds each ordinal threshold:
pij,k=P\(yij\>k∣Ii,𝐱j\)\.p\_\{ij,k\}=P\(y\_\{ij\}\>k\\mid I\_\{i\},\\mathbf\{x\}\_\{j\}\)\.\(6\)
The final predicted rating is then recovered by summing the predicted threshold outcomes:
y^ij=1\+∑k=1K−1𝟙\(pij,k\>0\.5\)\.\\hat\{y\}\_\{ij\}=1\+\\sum\_\{k=1\}^\{K\-1\}\\mathbbm\{1\}\(p\_\{ij,k\}\>0\.5\)\.\(7\)
The framework is trained end\-to\-end by minimising the sum of binary cross\-entropy losses across theK−1K\-1ordinal thresholds:
ℒord=−∑i,j∑k=1K−1\[rij,klogpij,k\+\(1−rij,k\)log\(1−pij,k\)\]\.\\mathcal\{L\}\_\{\\text\{ord\}\}=\-\\sum\_\{i,j\}\\sum\_\{k=1\}^\{K\-1\}\\left\[r\_\{ij,k\}\\log p\_\{ij,k\}\+\(1\-r\_\{ij,k\}\)\\log\(1\-p\_\{ij,k\}\)\\right\]\.\(8\)
## IVResults
### IV\-AAnalysis of Dataset
The final dataset contains 29,870 valid image–respondent ratings for 974 sidewalk\-view images\. Each image was rated by 25 to 35 distinct respondents, with an average of 31 ratings per image\. This exceeds the minimum number of ratings per image recommended for reliable Likert\-scale image\-based perception surveys\[[9](https://arxiv.org/html/2608.06934#bib.bib9)\]\. The respondent pool includes variation in age, gender, current residential context, childhood residential context, walking frequency, country of upbringing, and disability or health condition affecting walking\. This diversity supports the modeling of individual\-level variation in walkability perception, while recognising that some groups are represented less frequently than others\.
#### IV\-A1Rating distribution
Figure[3](https://arxiv.org/html/2608.06934#S4.F3)\(a\) presents the distribution of perceived walkability ratings overall and by settlement type, with a skew\-normal distribution fitted to each group using a common parametric family to enable direct comparison\. Ratings were concentrated in the middle categories, with scores 3 and 4 accounting for 61\.1% of all responses, while score 1 represented only 5\.8% of ratings\. This distribution reflects a moderate class imbalance and is consistent with patterns observed in subjective image assessment datasets\[[20](https://arxiv.org/html/2608.06934#bib.bib20)\], partially attributable to respondents’ tendency to avoid the extremes of rating scales\[[21](https://arxiv.org/html/2608.06934#bib.bib21)\]\. The fitted distributions captured this structure well\. All four groups exhibited a negative shape parameter \(α<0\\alpha<0\), indicating left\-skewed distributions in which responses concentrated toward the middle\-to\-moderately\-high part of the scale with only a thin tail extending toward low ratings\.
Differences were also observed across settlement types\. Urban images showed a higher concentration of ratings in the upper categories, peaking at rating 4, with only 4\.4% assigned to category 1\. In contrast, suburban and regional images peaked at rating 3, and contained a higher proportion of lower ratings\. These differences were reflected in the fitted parameters\. The urban distribution had the highest location \(ξ\\xi\) and the strongest negative skew, whereas the suburban and regional distributions were near\-identical to one another and shifted toward lower ratings\. Notably, the scale parameter \(ω\\omega\) was comparable across all groups, indicating that the distributions differed primarily in central location rather than in dispersion or overall shape\. This trend suggests that perceived walkability varies systematically across settlement contexts, potentially reflecting differences in pedestrian infrastructure quality and sidewalk conditions across urban, suburban, and regional environments\.
Figure 3:Perceived walkability ratings by image settlement type\. \(a\) Distribution of ratings overall and for urban, suburban, and regional imagery\. Symbols denote the observed proportion of responses at each rating; Solid \(dashed for Overall\) curves are skew\-normal distributions fitted to each group using a common parametric family, parameterized by location \(ξ\\xi\), scale \(ω\\omega\), and shape \(α\\alpha\)\. \(b\) Interaction between image settlement type and respondent residential context; points represent respondent\-level mean walkability ratings and error bars indicate 95% confidence intervals\.
#### IV\-A2Inter\-Rater Variability
Inter\-rater agreement was quantified using pairwise Cohen’s quadratic weighted kappa\. This metric is appropriate for ordinal Likert\-scale ratings because it accounts for chance agreement while assigning larger penalties to disagreements that are farther apart on the rating scale\[[9](https://arxiv.org/html/2608.06934#bib.bib9)\]\. Across all eligible respondent pairs \(pairs who both rated at least five common images\), the mean quadratic weighted kappa was 0\.17, indicating limited agreement among respondents when evaluating the same sidewalk environments\. This value is lower than kappa values reported in studies using more homogeneous participant pools and smaller image sets\[[9](https://arxiv.org/html/2608.06934#bib.bib9)\], reflecting the greater socio\-demographic diversity of respondents and environmental variability of images in the present study\.
To assess whether this agreement exceeded chance levels, we compared the observed kappa against a permutation\-based null distribution obtained by randomly shuffling image assignments across ratings\. The observed agreement was higher than expected under the null model \(mean nullκ≈0\.000\\kappa\\approx 0\.000,SD=0\.026SD=0\.026,p<0\.001p<0\.001\), suggesting that respondents share some common visual interpretation of walkability while still exhibiting substantial individual\-level variation\. This finding supports the use of a user\-conditioned modeling framework rather than relying solely on aggregate image\-level scores\.
#### IV\-A3Respondent–Environment Interactions in Walkability Perception
We examined the relationship between respondent\-level attributes and perceived walkability ratings using Spearman correlation for ordinal attributes and Kruskal–Wallis tests for categorical attributes\. No strong main effects were observed for age, gender, residential context, or disability affecting walking\. Walking frequency showed statistically significant but negligible association with mean rating \(r=0\.06r=0\.06,p=0\.034p=0\.034\), suggesting that respondent attributes do not uniformly shift overall rating levels\.
In contrast, significant interaction effects were observed between image settlement type and several respondent attributes, based on mixed\-effects likelihood\-ratio tests with random intercepts for respondent and image\. Image settlement type interacted significantly with current residential context \(LRχ2\(4\)=58\.58\\text\{LR \}\\chi^\{2\}\(4\)=58\.58,p<0\.001p<0\.001\), age \(LRχ2\(10\)=51\.83\\text\{LR \}\\chi^\{2\}\(10\)=51\.83,p<0\.001p<0\.001\), and walking frequency \(LRχ2\(8\)=63\.94\\text\{LR \}\\chi^\{2\}\(8\)=63\.94,p<0\.001p<0\.001\)\. As shown in Fig\.[3](https://arxiv.org/html/2608.06934#S4.F3)\(b\), respondents from different residential contexts rated urban images similarly, but diverged for suburban and regional images, with urban\-dwelling respondents assigning lower ratings to non\-urban environments\. Comparable patterns were observed for age and walking frequency, indicating that respondent attributes affect how individuals differentiate between sidewalk environments rather than simply shifting their overall rating tendency\. These patterns are consistent with prior work on subjective image assessment datasets\[[20](https://arxiv.org/html/2608.06934#bib.bib20)\], where observer\-level attributes may influence ratings through interactions with image content rather than through uniform shifts in mean scores\. The observed interactions provide empirical support for conditioning walkability prediction on both sidewalk\-view image content and respondent\-level characteristics\.
Figure 4:Viewpoint comparison across 100 matched locations\. \(a–c\) Sidewalk and \(d–f\) tree pixel coverage difference \(sidewalk minus street, %\) across Sydney CBD, Maroubra, and Tamworth, shown as representative examples of consistent and heterogeneous viewpoint effects respectively\. \(g\) Distribution of mean walkability ratings per image\. \(h\) Paired rating combination frequencies across 1,076 individual rating pairs\.
### IV\-BView point comparative analysis
Figure[4](https://arxiv.org/html/2608.06934#S4.F4)\(a\)–\(f\) presents representative spatial distributions of differences in sidewalk and tree pixel coverage between matched sidewalk\-view and street\-view images across the three study areas\. Sidewalk coverage was generally higher in sidewalk\-view imagery, particularly in urban and regional areas, while a more spatially variable pattern was observed in the suburban areas\. In contrast, tree coverage did not exhibit a consistent directional pattern across settlement types\. These results indicate that viewpoint differences do not uniformly increase or decrease the visibility of all visual features; rather, their effects vary with the local street environment, consistent with evidence that perspective differences have varying geographic effects on urban perceptions\[[22](https://arxiv.org/html/2608.06934#bib.bib22)\]\.
Figure[4](https://arxiv.org/html/2608.06934#S4.F4)\(g\) presents the distribution of mean walkability ratings per image for each viewpoint\. On average, sidewalk\-view images received higher walkability ratings than their matched street\-view counterparts \(mean = 3\.48 vs\. 3\.16\)\. A Wilcoxon signed\-rank test conducted on location\-aggregated median pairs confirmed that this difference was statistically significant \(W=207\.5W=207\.5,Z=4\.385Z=4\.385,p<0\.001p<0\.001\), with a large effect size \(rank\-biserial r=0\.61, location\-level Mdn = 3\.50 vs\. 3\.00\)\. Among the 100 matched locations, 52 exhibited non\-zero median differences, while 48 had identical median ratings across the two viewpoints\. At the observation level, 40\.1% of individual rating pairs were higher for the sidewalk\-view image, 41\.4% were identical, and 18\.6% were higher for the street\-view image, as shown in Fig\.[4](https://arxiv.org/html/2608.06934#S4.F4)\(h\)\. When ratings differed, the most common shift was a one\-point increase in favour of the sidewalk\-view perspective, suggesting that the effect is consistent in direction but modest in magnitude\. The higher ratings observed for sidewalk\-view imagery, together with the feature\-level differences identified in the matched image pairs, suggest that pedestrian\-level imagery better captures aspects of the walking environment that are relevant to perceived walkability\. These findings support the use of sidewalk\-view imagery as the primary visual input for modeling walkability perception\.
### IV\-CModel results
#### IV\-C1Experimental Setup
The dataset was divided into training, validation, and test subsets using a 70/10/20 split\. Stratified sampling was applied based on walkability score and settlement type to preserve comparable distributions across the three subsets\. This ensured that the validation and test sets closely followed the distribution of the training set with respect to both perceived walkability ratings and settlement categories\. A representative test set is essential for obtaining a reliable estimate of model performance\. Since the dataset is imbalanced toward middle\-range walkability scores, a model biased toward predicting average ratings may still achieve seemingly reasonable performance\. Therefore, an unrepresentative test set could lead to misleading conclusions about the model’s predictive ability\. By applying stratified sampling, we ensured that the test set covers the full range of walkability scores and avoids disproportionate representation of images that may be easier for the model to predict\.
The model was evaluated under a rating\-completion protocol: raters and images could recur across the training and test partitions, but no individual image–respondent rating was shared between them\. This setting was used to assess whether conditioning on respondent attributes improves the prediction of individual walkability judgments over a non\-user\-conditioned, image\-only model, while keeping the image distribution comparable across partitions\. Because the model was conditioned on respondent attributes rather than respondent identity, individual raters could not be directly memorized; the observed improvement was therefore interpreted as evidence of learnable structure in how respondent characteristics shape walkability perception\. Accordingly, the results were framed as evidence for user\-conditioned preference modelling within the sampled image pool, with broader generalisation left to future work with larger\-scale data\.
All models were implemented within the AutoGluon\-Multimodal \(AutoMM\) framework\[[39](https://arxiv.org/html/2608.06934#bib.bib39)\], extended to support ordinal prediction\. All models were trained using AdamW with a learning rate of1×10−41\\times 10^\{\-4\}, weight decay of1×10−31\\times 10^\{\-3\}, layer\-wise learning\-rate decay with a factor of 0\.9, gradient\-norm clipping at 1\.0, and a cosine decay schedule with linear warm\-up over the first 10% of training steps\[[23](https://arxiv.org/html/2608.06934#bib.bib23)\]\. Training was run for up to 50 epochs with early stopping using a patience of 10 validation checks based on quadratic weighted kappa \(QWK\)\. To account for training variability, all models were trained over five independent runs with different random seeds, and results are reported as mean ± standard deviation\. All runs used mixed\-precision \(bf16\) training with an effective batch size of 128\.
#### IV\-C2Evaluation Criteria
Because the prediction target was ordinal, evaluation metrics needed to reflect the ordered structure of the rating scale\. Exact\-match accuracy \(ACC\) was reported as a secondary indicator, as it treated all misclassifications equally, regardless of their distance from the true rating\. We therefore reported mean absolute error \(MAE\) to quantify the average rating\-level deviation between predicted and observed scores, and within\-one accuracy to measure the proportion of predictions that fell within one rating level of the ground truth\. However, within\-one accuracy could be overly permissive for imbalanced five\-point rating distributions, particularly when observations were concentrated in the middle categories\. A model biased toward central ratings could still achieve high within\-one accuracy without capturing the full ordinal structure of the data\. For this reason, QWK was used as the primary model agreement metric between predicted and observed walkability ratings\. This use follows common practice in ordinal prediction tasks, where model outputs are evaluated against reference human\-assigned scores\[[24](https://arxiv.org/html/2608.06934#bib.bib24),[25](https://arxiv.org/html/2608.06934#bib.bib25)\]\.
#### IV\-C3Effect of User Conditioning
Table[I](https://arxiv.org/html/2608.06934#S4.T1)and Figure[5](https://arxiv.org/html/2608.06934#S4.F5)present the main comparison between the image\-only baseline and the proposed user\-conditioned walkability perception framework\. Incorporating respondent attributes yields substantial improvements across all metrics\. Compared with the image\-only baseline, QWK increases from 0\.285±\\pm0\.023 to 0\.469±\\pm0\.026, corresponding to an approximately 65% relative improvement in rank agreement\. MAE decreases from 0\.885±\\pm0\.014 to 0\.785±\\pm0\.029, while accuracy increases from 0\.344±\\pm0\.005 to 0\.410±\\pm0\.018\. The improvement in QWK is particularly meaningful because this metric penalizes large ordinal errors quadratically, indicating that user conditioning not only improves classification performance but also reduces the severity of ordinal prediction errors\.
Per\-class results in table[II](https://arxiv.org/html/2608.06934#S4.T2)reveal that the aggregate improvement is driven primarily by gains at the rating extremes\. For classes 1 and 2, representing environments perceived as least walkable, recall increases from0\.106±0\.0350\.106\\pm 0\.035to0\.271±0\.0650\.271\\pm 0\.065and from0\.124±0\.0850\.124\\pm 0\.085to0\.324±0\.0190\.324\\pm 0\.019, respectively, with corresponding improvements in F1\. Class 5 shows a similar pattern, with recall increasing from0\.224±0\.0300\.224\\pm 0\.030to0\.402±0\.0160\.402\\pm 0\.016and F1 from0\.278±0\.0170\.278\\pm 0\.017to0\.422±0\.0220\.422\\pm 0\.022\. These gains suggest that respondent attributes contribute discriminative information, particularly at the extremes of the walkability scale, where the image\-only model shows the weakest recall\. The exception is class 3, where the image\-only model achieves marginally higher recall \(0\.478±±0\.0970\.478\\pm\\textpm 0\.097vs\.0\.437±±0\.0120\.437\\pm\\textpm 0\.012\); given that mid\-range classes account for 61\.1% of ratings, this likely reflects a central tendency bias rather than stronger visual discriminability\. This interpretation is supported by the confusion matrices in Fig\.[5](https://arxiv.org/html/2608.06934#S4.F5), where the image\-only model concentrates predictions toward classes 3 and 4, whereas the user\-conditioned model produces a more distributed prediction profile across the rating scale\. Overall, these results suggest that respondent\-level attributes provide meaningful information for modelling individual differences in perceived walkability and help mitigate the central tendency bias observed when predictions rely only on visual features\.
TABLE I:Test\-set comparison between the image\-only baseline and the proposed user\-conditioned model\. W\-1 denotes within\-one accuracy\.TABLE II:Per\-class test\-set performance\.Figure 5:Row\-normalized test confusion matrices averaged over seeds for \(a\) the image\-only baseline and \(b\) the proposed user\-conditioned image\-tabular model\.
#### IV\-C4Component ablations
We conduct a sequential component ablation to validate the architectural choices underlying the proposed user\-conditioned walkability perception framework\. At each stage, all previously determined components are fixed at their best\-performing configuration while a single component is varied, isolating the contribution of each design decision\.
Image Backbone\.We first evaluate four image backbones: CaFormer\-B36\[[17](https://arxiv.org/html/2608.06934#bib.bib17)\], ConvNeXt\-Tiny\[[32](https://arxiv.org/html/2608.06934#bib.bib32)\], Swin\-Tiny, and ResNet50\[[30](https://arxiv.org/html/2608.06934#bib.bib30)\]\. These models span convolutional, hierarchical, and MetaFormer\-based architectures, with the tabular encoder fixed at FT\-Transformer, Transformer\-based fusion, and loss at CORAL\. As shown in Table[III](https://arxiv.org/html/2608.06934#S4.T3), Swin\-Tiny achieves the highest mean validation QWK and the lowest MAE, while ResNet50 obtains comparable performance and slightly higher Within\-1 accuracy\. We select Swin\-Tiny as the image encoder for the subsequent experiments based on its best overall validation performance\.
Tabular Encoder\.With Swin\-Tiny fixed, we compare a standard MLP with FT\-Transformer\[[18](https://arxiv.org/html/2608.06934#bib.bib18)\]as the respondent attribute encoder, while keeping the fusion module and loss function fixed as Transformer\-based and CORAL, respectively\. As shown in Table[III](https://arxiv.org/html/2608.06934#S4.T3), FT\-Transformer outperforms the MLP encoder across all metrics, likely due to its per\-feature tokenisation mechanism that projects each categorical respondent attribute into a dedicated embedding and learns pairwise inter\-attribute interactions through self\-attention, rather than relying on implicit feature interactions in a flat concatenated input\[[18](https://arxiv.org/html/2608.06934#bib.bib18)\]\.
Fusion Module\.With Swin\-tiny and FT\-Transformer fixed, we compare an MLP fusion head with a Transformer\-based fusion module, while keeping the loss function fixed as CORAL\. As shown in Table[III](https://arxiv.org/html/2608.06934#S4.T3), the Transformer\-based fusion head achieved superior performance\.
Loss Function\.Finally, with all architectural components fixed, we compare three loss functions: standard cross\-entropy, CORN\[[33](https://arxiv.org/html/2608.06934#bib.bib33)\], and CORAL\. CORN \(Conditional Ordinal Regression for Neural Networks\) models conditional probabilities,\(P\(y\>k∣y\>k−1\)\)\(P\(y\>k\\mid y\>k\-1\)\), using independent output neurons and eligible training subsets for each ordinal threshold\. As shown in Table[III](https://arxiv.org/html/2608.06934#S4.T3), both ordinal loss functions outperformed cross\-entropy, confirming that the five\-point walkability scale contains meaningful rank\-order information that is discarded by nominal classification losses\. CORAL further outperformed CORN across all metrics\. This result is notable because existing evaluations of CORAL and CORN have primarily focused on unimodal ordinal regression settings\[[19](https://arxiv.org/html/2608.06934#bib.bib19),[33](https://arxiv.org/html/2608.06934#bib.bib33)\], whereas the present study evaluates them within a multimodal image–tabular ordinal classification framework\. CORAL was therefore selected as the loss function for all reported experiments\.
TABLE III:Component ablation on the validation set\. The selected model uses Swin\-Tiny, FT\-Transformer, fusion Transformer, and CORAL\. W\-1 denotes Within\-1 accuracy\.TABLE IV:Permutation\-based importance of respondent\-level features\.
### IV\-DAttribute importance
To examine the sensitivity of model performance to individual respondent attributes, permutation\-based feature importance\[[26](https://arxiv.org/html/2608.06934#bib.bib26)\]was estimated by randomly shuffling each attribute’s values in the test set while holding all others fixed\. The resulting change in performance was measured using QWK, Within\-1 accuracy, and MAE\. For QWK and Within\-1 accuracy, importance was defined as the decrease from the baseline score, whereas for MAE it was defined as the increase in prediction error\. Each attribute was permuted 30 times per seed; importance was first summarised as the median change across permutations for each seed, and then reported as the median with interquartile range \(IQR\) across all seeds\.
Table[IV](https://arxiv.org/html/2608.06934#S4.T4)presents the results, which we interpret as a relative ranking of attribute influence rather than as independent, additive marginal contributions to model performance\. Age produced the largest drop in QWK \(Δ\\DeltaQWK = 0\.204, IQR = 0\.034\), followed by residence type and walking frequency with comparable importance scores \(Δ\\DeltaQWK = 0\.166 and 0\.164, IQR = 0\.020 and 0\.023, respectively\), indicating that lived environmental context and the habitual walking behaviour are both informative to the model’s predictions\. Notably, theΔ\\DeltaQWK for age alone exceeds the total QWK improvement of the user\-conditioned model over the image\-only baseline \(0\.184; Table[I](https://arxiv.org/html/2608.06934#S4.T1)\), which would not be expected if these values represented independent, additive marginal contributions\. This is consistent with a documented limitation of permutation\-based importance under correlated predictors, whereby shuffling one attribute independently of others \(plausibly including age, residence type, childhood residential environment, and walking frequency\) can generate unrealistic respondent profiles and inflate apparent importance beyond an attribute’s true marginal contribution\[[36](https://arxiv.org/html/2608.06934#bib.bib36)\]\. Accordingly, we do not interpret the magnitudes in Table[IV](https://arxiv.org/html/2608.06934#S4.T4)as a decomposition of the total attribute contribution to model performance; instead, they should be viewed as indicating the relative influence of each attribute under this perturbation procedure\. Childhood area and gender showed moderate importance, while disability status and childhood country had the lowest importance scores, which may partly reflect their limited variability in the sample: 83\.8% of respondents reported no walking\-related disability, and 80\.0% were raised in Australia\. Overall, the relative ranking of attributes was broadly consistent across QWK, Within\-1 accuracy, and MAE, supporting the robustness of the feature\-importance pattern\.
## VConclusion
This paper introduced a new walkability perception dataset spanning urban, suburban, and regional areas in Australia, linking sidewalk\-view imagery with individual rater attributes\. A supplementary viewpoint comparison showed that sidewalk\-view imagery received significantly higher walkability ratings than matched street\-view imagery, indicating that the choice of visual data source is a substantive design decision in walkability perception surveys rather than an interchangeable convenience\. The proposed multimodal user\-conditioned framework, which conditions walkability perception predictions on both sidewalk\-view imagery and rater characteristics, substantially outperformed an image\-only baseline across all evaluation metrics\. To the best of our knowledge, this is the first study to formulate visual walkability perception as a user\-conditioned prediction task, demonstrating that respondent\-level attributes provide a predictive signal beyond image content alone\. Complementary interaction analysis showed that respondent attributes shape how individuals differentiate between environments rather than simply shifting overall rating levels, a pattern reinforced by permutation\-based feature importance showing that predictive contribution is distributed across the respondent profile rather than driven by a single dominant attribute\. These findings suggest that individual\-level variation in perceived walkability is systematic and should not be treated merely as noise to be averaged out\. Overall, the results support moving beyond aggregated, observer\-independent walkability scores toward models that explicitly represent who is evaluating an environment, providing a step toward more inclusive approaches to walkability assessment\.
This study has several limitations\. The dataset is drawn from Australian locations and a predominantly Australian\-raised respondent pool, and the transferability of both the perception patterns and the trained model to other cultural and infrastructural contexts remains to be established\. The respondent attributes capture systematic, group\-level sources of perceptual variation; fully distinctive preferences, which the modest inter\-rater agreement suggests are substantial, are beyond the reach of our current dataset and study\. The viewpoint comparison was a supplementary analysis intended to motivate the choice of sidewalk\-view imagery rather than a full investigation of viewpoint effects; a systematic examination with broader imagery data and participant pools is left for future work\.
## References
- \[1\]D\. T\. Duncan, J\. Aldstadt, J\. Whalen, S\. J\. Melly, and S\. L\. Gortmaker, “Validation of Walk Score for estimating neighborhood walkability: An analysis of four US metropolitan areas,”*International Journal of Environmental Research and Public Health*, vol\. 8, no\. 11, pp\. 4160–4179, 2011\.
- \[2\]H\. Zhou, S\. He, Y\. Cai, M\. Wang, and S\. Su, “Social inequalities in neighborhood visual walkability: Using street view imagery and deep learning technologies to facilitate healthy city planning,”*Sustainable Cities and Society*, vol\. 50, Art\. no\. 101605, 2019\.
- \[3\]U\. Jehle, M\. T\. Baquero Larriva, M\. BaghaiePoor, and B\. Büttner, “How does pedestrian accessibility vary for different people? Development of a perceived user\-specific accessibility measure for walking \(PAW\),”*Transportation Research Part A: Policy and Practice*, vol\. 189, p\. 104203, 2024\.
- \[4\]K\. Ito, Y\. Kang, Y\. Zhang, F\. Zhang, and F\. Biljecki, “Understanding urban perception with visual data: A systematic review,”*Cities*, vol\. 152, Art\. no\. 105169, 2024\.
- \[5\]Y\. Li, N\. Yabuki, and T\. Fukuda, “Measuring visual walkability perception using panoramic street view images, virtual reality, and deep learning,”*Sustainable Cities and Society*, vol\. 86, p\. 104140, 2022\.
- \[6\]M\. Quintana, Y\. Gu, X\. Liang, Y\. Hou, K\. Ito, Y\. Zhu, M\. Abdelrahman, and F\. Biljecki, “Global urban visual perception varies across demographics and personalities,”*Nature Cities*, pp\. 1–15, 2025\.
- \[7\]C\. Dickinson, K\. Manaugh, P\. Pathak, and R\. Sengupta, “Geographic identity and perceptions of walkable space,”*Travel Behaviour and Society*, vol\. 34, p\. 100703, 2024\.
- \[8\]K\. Ito, M\. Quintana, X\. Han, R\. Zimmermann, and F\. Biljecki, “Translating street view imagery to correct perspectives to enhance bikeability and walkability studies,”*International Journal of Geographical Information Science*, 2024\.
- \[9\]Y\. Gu, M\. Quintana, X\. Liang, K\. Ito, W\. Yap, and F\. Biljecki, “Designing effective image\-based surveys for urban visual perception,”*Landscape and Urban Planning*, vol\. 260, p\. 105368, 2025\.
- \[10\]R\. Ewing and S\. Handy, “Measuring the unmeasurable: Urban design qualities related to walkability,”*Journal of Urban Design*, vol\. 14, no\. 1, pp\. 65–84, 2009\.
- \[11\]A\. Dubey, N\. Naik, D\. Parikh, R\. Raskar, and C\. A\. Hidalgo, “Deep learning the city: Quantifying urban perception at a global scale,” in*Proc\. European Conf\. Computer Vision \(ECCV\)*, 2016, pp\. 196–212\.
- \[12\]J\. Hrnčíř, P\. Žilecký, Q\. Song, and M\. Jakob, “Practical multicriteria urban bicycle routing,”*IEEE Transactions on Intelligent Transportation Systems*, vol\. 18, no\. 3, pp\. 493–504, 2017\.
- \[13\]Qualtrics, “Response quality checks,” 2024\. \[Online\]\. Available:[https://www\.qualtrics\.com/support/survey\-platform/survey\-module/survey\-checker/response\-quality/](https://www.qualtrics.com/support/survey-platform/survey-module/survey-checker/response-quality/)\. Accessed: Aug\. 2, 2025\.
- \[14\]I\. Blečić, A\. Cecchini, and G\. A\. Trunfio, “Towards automatic assessment of perceived walkability,” in*Proc\. Int\. Conf\. Computational Science and Its Applications \(ICCSA\)*, 2018, pp\. 351–365\.
- \[15\]Y\. Yang, L\. Xu, L\. Li, N\. Qie, Y\. Li, P\. Zhang, and Y\. Guo, “Personalized image aesthetics assessment with rich attributes,” in*Proc\. IEEE/CVF Conf\. Computer Vision and Pattern Recognition \(CVPR\)*, 2022, pp\. 19861–19869\.
- \[16\]S\. Ebrahimi, S\. O\. Arik, Y\. Dong, and T\. Pfister, “Lanistr: Multimodal learning from structured and unstructured data,”*arXiv preprint arXiv:2305\.16556*, 2023\.
- \[17\]W\. Yu, C\. Si, P\. Zhou, M\. Luo, Y\. Zhou, J\. Feng, S\. Yan, and X\. Wang, “MetaFormer baselines for vision,”*IEEE Transactions on Pattern Analysis and Machine Intelligence*, vol\. 46, no\. 2, pp\. 896–912, 2023\.
- \[18\]Y\. Gorishniy, I\. Rubachev, V\. Khrulkov, and A\. Babenko, “Revisiting deep learning models for tabular data,” in*Advances in Neural Information Processing Systems*, vol\. 34, 2021, pp\. 18932–18943\.
- \[19\]W\. Cao, V\. Mirjalili, and S\. Raschka, “Rank consistent ordinal regression for neural networks with application to age estimation,”*Pattern Recognition Letters*, vol\. 140, pp\. 325–331, 2020, doi: 10\.1016/j\.patrec\.2020\.11\.008\.
- \[20\]A\.\-S\. Maerten, L\.\-W\. Chen, S\. De Winter, C\. Bossens, and J\. Wagemans, “LAPIS: A novel dataset for personalized image aesthetic assessment,” in*Proc\. IEEE/CVF Conf\. Computer Vision and Pattern Recognition \(CVPR\)*, 2025, pp\. 6302–6311\.
- \[21\]N\. A\. de Rezende and D\. D\. de Medeiros, “How rating scales influence responses’ reliability, extreme points, middle point and respondent’s preferences,”*Journal of Business Research*, vol\. 138, pp\. 266–274, 2022\.
- \[22\]J\. Rui, “Measuring streetscape perceptions from driveways and sidewalks to inform pedestrian\-oriented street renewal in Düsseldorf,”*Cities*, vol\. 141, p\. 104472, 2023\.
- \[23\]Z\. Tang, Z\. Zhong, T\. He, and G\. Friedland, “Bag of tricks for multimodal AutoML with image, text, and tabular data,”*arXiv preprint arXiv:2412\.16243*, 2024\.
- \[24\]K\. Taghipour and H\. T\. Ng, “A neural approach to automated essay scoring,” in*Proc\. Conf\. Empirical Methods in Natural Language Processing \(EMNLP\)*, 2016, pp\. 1882–1891\.
- \[25\]J\. Sahlsten, J\. Jaskari, J\. Kivinen, L\. Turunen, E\. Jaanio, K\. Hietala, and K\. Kaski, “Deep learning fundus image analysis for diabetic retinopathy and macular edema grading,”*Scientific Reports*, vol\. 9, no\. 1, p\. 10750, 2019\.
- \[26\]L\. Breiman, “Random forests,”*Machine Learning*, vol\. 45, no\. 1, pp\. 5–32, 2001\.
- \[27\]M\. Costa, M\. Marques, C\. L\. Azevedo, F\. W\. Siebert, and F\. Moura, “Which cycling environment appears safer? Learning cycling safety perceptions from pairwise image comparisons,”*IEEE Transactions on Intelligent Transportation Systems*, vol\. 26, no\. 2, pp\. 1689–1700, 2025\.
- \[28\]J\. Ren, X\. Shen, Z\. Lin, R\. Mech, and D\. J\. Foran, “Personalized image aesthetics,” in*Proc\. IEEE Int\. Conf\. Computer Vision \(ICCV\)*, 2017, pp\. 638–647\.
- \[29\]F\. Biljecki and K\. Ito, “Street view imagery in urban analytics and GIS: A review,”*Landscape and Urban Planning*, vol\. 215, p\. 104217, 2021\.
- \[30\]K\. He, X\. Zhang, S\. Ren, and J\. Sun, “Deep residual learning for image recognition,” in*Proc\. IEEE Conf\. Computer Vision and Pattern Recognition \(CVPR\)*, 2016, pp\. 770–778\.
- \[31\]Z\. Liu, Y\. Lin, Y\. Cao, H\. Hu, Y\. Wei, Z\. Zhang, S\. Lin, and B\. Guo, “Swin Transformer: Hierarchical vision transformer using shifted windows,” in*Proc\. IEEE/CVF Int\. Conf\. Computer Vision \(ICCV\)*, 2021, pp\. 10012–10022\.
- \[32\]Z\. Liu, H\. Mao, C\.\-Y\. Wu, C\. Feichtenhofer, T\. Darrell, and S\. Xie, “A ConvNet for the 2020s,” in*Proc\. IEEE/CVF Conf\. Computer Vision and Pattern Recognition \(CVPR\)*, 2022, pp\. 11976–11986\.
- \[33\]X\. Shi, W\. Cao, and S\. Raschka, “Deep neural networks for rank\-consistent ordinal regression based on conditional probabilities,”*Pattern Analysis and Applications*, vol\. 26, no\. 3, pp\. 941–955, 2023\.
- \[34\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, L\. Kaiser, and I\. Polosukhin, “Attention is all you need,” in*Proc\. Advances in Neural Information Processing Systems \(NeurIPS\)*, 2017\.
- \[35\]T\. Qin, C\. Li, H\. Ye, S\. Wan, M\. Li, H\. Liu, and M\. Yang, “Crowd\-sourced NeRF: Collecting data from production vehicles for 3D street view reconstruction,”*IEEE Transactions on Intelligent Transportation Systems*, vol\. 25, no\. 11, pp\. 16145–16156, 2024\.
- \[36\]G\. Hooker and L\. Mentch, “Please stop permuting features: An explanation and alternatives,”*arXiv preprint arXiv:1905\.03151*, 2019\.
- \[37\]T\. A\. Arentze, “Adaptive personalized travel information systems: A Bayesian method to learn users’ personal preferences in multimodal transport networks,”*IEEE Transactions on Intelligent Transportation Systems*, vol\. 14, no\. 4, pp\. 1957–1966, 2013\.
- \[38\]M\. Veres and M\. Moussa, “Deep learning for intelligent transportation systems: A survey of emerging trends,”*IEEE Transactions on Intelligent Transportation Systems*, vol\. 21, no\. 8, pp\. 3152–3168, 2020\.
- \[39\]Z\. Tang, H\. Fang, S\. Zhou, T\. Yang, Z\. Zhong, T\. Hu, K\. Kirchhoff, and G\. Karypis, “AutoGluon\-Multimodal \(AutoMM\): Supercharging multimodal AutoML with foundation models,”*arXiv preprint arXiv:2404\.16233*, 2024\.
Moloud Damandehreceived the B\.Sc\. and M\.Sc\. degrees in industrial engineering from Sharif University of Technology, Tehran, Iran\. She is currently pursuing the Ph\.D\. degree with the School of Civil and Environmental Engineering, University of New South Wales \(UNSW\), Sydney, Australia, where she is a member of the Research Centre for Integrated Transport Innovation \(rCITI\)\. Her research focuses on measuring perceived walkability of urban environments using pedestrian\-perspective imagery, computer vision, and multimodal deep learning\.
Meead Saberiis a Professor in the School of Civil and Environmental Engineering at the University of New South Wales \(UNSW\), Sydney, Australia, a member of the Research Centre for Integrated Transport Innovation \(rCITI\), and an Australian Research Council Future Fellow \(2026\-2031\)\. His work has advanced both theoretical and applied understanding of transport networks, urban mobility, and sustainable transport systems\.Similar Articles
In-Context Learning to Assess Built Environment Impacts on Perceived Neighborhood Walkability Among Mobility-impaired Older Adults
This paper investigates using in-context learning with the TabPFN model to predict perceived neighborhood walkability based on built environment features for mobility-impaired older adults, achieving higher performance than baseline models and employing SHAP-IQ for interaction analysis.
Learn to Quantify Social Interaction with Constraints for Pedestrian Walking
This paper introduces a method called 'Learn to Cluster' to quantify and interpret social interactions among pedestrians for better trajectory prediction. It uses probabilistic latent variable generative learning to cluster social interactions without labels, improving robustness for autonomous driving and social robots.
Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling
This paper identifies perceptual judgment bias in multimodal LLM judges, where they over-reward fluent but visually wrong responses, and proposes a dataset PPJD and a trained model Perception-Judge using GRPO with batch-ranking reward to mitigate this bias and improve perception-grounded evaluation.
Mapping the City Through the Lens of Language Models
This paper measures the implicit assumptions language models make about 'a city' by scoring anonymized urban profiles across 40 indicators, finding a shared preference for larger, faster-growing, and more infrastructure-rich cities. It uses open-weight checkpoints and replication data to make the default portrait of cities in LLMs empirically traceable.
Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception
Urban-ImageNet is a large-scale multi-modal dataset and evaluation benchmark for urban space perception from social media imagery, supporting scene classification, cross-modal retrieval, and instance segmentation tasks across 61 urban sites in 24 Chinese cities.