Adaptive Multi-Agent Feature Selection for Personalized Fall Risk Prevention
Summary
The paper introduces PAFIR, an adaptive feature selection framework using reinforcement learning for personalized fall risk prevention from longitudinal multimodal health data, demonstrating improved effectiveness over baseline methods.
View Cached Full Text
Cached at: 08/20/26, 10:26 AM
# Adaptive Multi-Agent Feature Selection for Personalized Fall Risk Prevention Source: [https://arxiv.org/html/2608.18450](https://arxiv.org/html/2608.18450) \\jmlrpages Chang LiuEmail:[chang\.liu@ucf\.edu](mailto:[email protected])Affiliation:School of Data, Mathematical, and Statistical Sciences University of Central Florida Orlando, FL, USA and College of Nursing University of Central Florida Orlando, FL, USA and School of Computing and Augmented Intelligence Arizona State University Tempe, AZ, USA and School of Data, Mathematical, and Statistical Sciences, College of Nursing University of Central Florida Orlando, FL, USAYanjie FuEmail:[Yanjie\.Fu@asu\.edu](mailto:[email protected])Affiliation:Rui Xie†\\daggerEmail:[Rui\.Xie@ucf\.edu](mailto:[email protected])Affiliation: ###### Abstract Falls among older adults represent a major public health challenge driven by complex, time\-varying interactions across multiple risk domains\. Effective fall risk factor identification requires learning from heterogeneous longitudinal data while accounting for sparse and delayed fall\-related outcome events\. However, existing approaches are largely static and fail to adaptively model evolving, individualized risk factors across modalities and time\. We proposePAFIR, aPersonalized andAdaptiveFeature selection framework for fall riskIdentification and pRevention, which formulates adaptive feature selection as a reinforcement learning problem over longitudinal multimodal health data\. PAFIR jointly models structural dependencies among correlated assessment variables and temporal dynamics in wearable\-derived physical activity data, and learns adaptive selection policies across repeated study visits using reward signals derived from sparse fall incidence outcomes\. We apply PAFIR to data from the Physio fEedback Exercise pRogram \(PEER\) cluster\-randomized trial\. Experimental results demonstrate that PAFIR more effectively captures longitudinal and structural patterns of feature relevance than state\-of\-the\-art baselines, and enables dynamic, subject\-specific feature selection\. By adapting selected features over time, PAFIR supports more timely and personalized fall prevention strategies\. ††proceedings:PMLR: Proceedings of Machine Learning Research††volume:340††year:2026††workshop:Machine Learning for Healthcare$\\dagger$$\\dagger$footnotetext:Corresponding author\.## 1Introduction Fall prevention among older adults remains an urgent and complex public health challenge\. In the United States, falls are the leading cause of fatal and non\-fatal injuries among adults aged 65 and older, resulting in over38,00038\{,\}000deaths annually and substantial healthcare costs\([4](https://arxiv.org/html/2608.18450#bib.bib4);[19](https://arxiv.org/html/2608.18450#bib.bib27)\)\. Beyond acute injury, falls often trigger long\-term functional decline, loss of independence, and diminished quality of life\([54](https://arxiv.org/html/2608.18450#bib.bib22)\)\. As populations age and care resources become increasingly constrained, there is a critical need for timely, scalable, and personalized strategies to identify fall risk and intervene before adverse events occur\. Recent advances in mobile and smart health technologies enable continuous monitoring of gait, activity, balance, and contextual behaviors, creating new opportunities for early fall risk detection\([56](https://arxiv.org/html/2608.18450#bib.bib23);[22](https://arxiv.org/html/2608.18450#bib.bib24);[50](https://arxiv.org/html/2608.18450#bib.bib21)\)\. However, translating these heterogeneous, high\-dimensional data sources, which include continuous sensor streams, multi\-time\-scale signals, and structured survey\-based measurements, into effective prevention remains challenging\. Figure 1:Overview of thePAFIRframework\. Heterogeneouslongitudinal datafrom multiple modalities are integrated to construct ahierarchical state representationof the individual state\. This state is iteratively updated across clinical visits and in response to sparsely observed events, i\.e\., fall incidence\. Amulti\-agent reinforcement learningmodule then operates on the evolving state to perform adaptivefeature selectionat both group and individual levels\. The resulting learned policies supportpersonalized and community\-level recommendationsfor fall risk identification and prevention\.A critical insight, however, is that fall risk is inherently dynamic and highly individualized, shaped by both objective physical capacity and subjective risk perception\. Older adults may underestimate their physiological vulnerability or, conversely, restrict activity due to fear of falling despite preserved function\([15](https://arxiv.org/html/2608.18450#bib.bib51);[12](https://arxiv.org/html/2608.18450#bib.bib49);[10](https://arxiv.org/html/2608.18450#bib.bib50);[42](https://arxiv.org/html/2608.18450#bib.bib59)\)\. Such misalignment between theBody\(mobility, balance\) and theMind\(confidence, cognition, and risk perception\) can lead to maladaptive behaviors, delayed risk recognition, and reduced engagement with preventive care\([26](https://arxiv.org/html/2608.18450#bib.bib53);[80](https://arxiv.org/html/2608.18450#bib.bib52);[46](https://arxiv.org/html/2608.18450#bib.bib54)\)\. These complexities make effective prevention not only a problem of detecting risk signals, but also of adaptively determining*which*features are most informative and*when*to intervene for each individual\([11](https://arxiv.org/html/2608.18450#bib.bib25);[58](https://arxiv.org/html/2608.18450#bib.bib26)\)\. Although many individual fall risk factors have been studied, how physiological, psychological, and behavioral factors interact and co\-evolve over time remains poorly understood\. Most existing approaches model these factors in isolation or assume static relationships across visits\([23](https://arxiv.org/html/2608.18450#bib.bib55);[87](https://arxiv.org/html/2608.18450#bib.bib56);[70](https://arxiv.org/html/2608.18450#bib.bib57)\), limiting their ability to capture dynamic longitudinal interactions between physical capacity, risk perception, and behavior\. The Physio fEedback Exercise pRogram \(PEER\) study\([71](https://arxiv.org/html/2608.18450#bib.bib20)\)is a technology\-based intervention, specifically, designed to address this mismatch, integrating real\-time physio\-feedback, cognitive reframing, and peer\-led exercise to jointly engage physical and psychological fall risk domains\. This study offers rich, longitudinal multimodal data spanning structured survey\-based assessments, clinical measurements, fall incidence, and wearable sensor\-derived time\-series data, as visualized in Fig\.[2](https://arxiv.org/html/2608.18450#S1.F2), providing a comprehensive data foundation for data\-driven fall risk identification and prevention\. This gap motivates data\-driven methods that integrate multimodal signals and adaptively model evolving risk factor interactions, framing personalized fall risk identification as an adaptive feature selection problem over heterogeneous physiological, behavioral, and contextual data, as illustrated in the overview of the proposed PAFIR framework in Fig\.[1](https://arxiv.org/html/2608.18450#S1.F1)\. Figure 2:Illustration of the multimodal longitudinal PEER dataset across visits \(T1\-T4\), including \(1\) survey\-based fall risk assessments, \(2\) high\-frequency wearable time\-series data capturing physical activity, and \(3\) instrument\-based measurements such as body composition\.Fall incidence, the primary outcome, is recorded as sparse, time\-stamped events\.Emerging developments in reinforcement learning \(RL\) for healthcare have shown strong promise for adaptive feature selection as a sequential decision process, enabling models to update feature relevance based on observations and clinical feedback\([90](https://arxiv.org/html/2608.18450#bib.bib41);[32](https://arxiv.org/html/2608.18450#bib.bib40)\)\. Despite these advances, existing feature selection systems continue to face fundamental limitations when deployed in real\-world healthcare settings, particularly for fall prevention\. A primary challenge in fall risk factor identification arises from the integration of highly diverse data sources, including self\-reported surveys\([60](https://arxiv.org/html/2608.18450#bib.bib30)\), body composition measurements\([52](https://arxiv.org/html/2608.18450#bib.bib33)\), muscle strength and balance evaluations\([47](https://arxiv.org/html/2608.18450#bib.bib29)\), daily physical activity captured via wearable devices\([28](https://arxiv.org/html/2608.18450#bib.bib31);[40](https://arxiv.org/html/2608.18450#bib.bib58)\), and gait and posture assessments\([39](https://arxiv.org/html/2608.18450#bib.bib28)\)\. Most existing studies focus on only a subset of these data types or rely on simplified fusion strategies\([24](https://arxiv.org/html/2608.18450#bib.bib34)\), limiting their ability to capture cross\-modal interactions and evolving risk patterns\. Additionally, fall risk modeling is further complicated by the dynamic and heterogeneous nature of activity and risk signals in older adults\. Wearable sensors produce continuous, minute\-level physical activity data\([1](https://arxiv.org/html/2608.18450#bib.bib46)\), while fall outcomes are sparse, delayed, and irregular\([76](https://arxiv.org/html/2608.18450#bib.bib47);[37](https://arxiv.org/html/2608.18450#bib.bib45)\), complicating temporal alignment and longitudinal analysis\. Gradual changes in mobility, cognition, and psychological state may develop over extended periods, whereas acute events can induce abrupt shifts in risk\. However, learning from sparse outcome signals for fall incidence remains a major challenge\. For example, in the PEER study\([71](https://arxiv.org/html/2608.18450#bib.bib20)\), fall events are inherently rare, with only 69 events observed across 1364 visit\-level observations and many participants experiencing no falls throughout the entire trial\. When the reward is defined solely based on the binary occurrence of a fall, the RL agent receives informative feedback in only a small fraction of training steps\([79](https://arxiv.org/html/2608.18450#bib.bib81);[16](https://arxiv.org/html/2608.18450#bib.bib82)\), corresponding to a setting with limited and delayed feedback where informative signals are difficult to propagate across time\([81](https://arxiv.org/html/2608.18450#bib.bib83)\)\. Fall\-risk research has employed a broad range of statistical and machine\-learning methods for identifying relevant risk factors\. Marginal filtering and importance\-ranking approaches typically evaluate variables separately, whereas multivariable methods, including LASSO\([75](https://arxiv.org/html/2608.18450#bib.bib77)\)and Group LASSO\([91](https://arxiv.org/html/2608.18450#bib.bib78)\), perform joint selection under sparsity constraints\. Established longitudinal and survival\-analysis approaches can also accommodate repeated measurements and event outcomes\. For example, joint models\([78](https://arxiv.org/html/2608.18450#bib.bib61);[17](https://arxiv.org/html/2608.18450#bib.bib60);[9](https://arxiv.org/html/2608.18450#bib.bib62)\)link longitudinal biomarker trajectories with time\-to\-event outcomes, while frailty models\([27](https://arxiv.org/html/2608.18450#bib.bib63)\)account for unobserved heterogeneity\. These methods are particularly appropriate when the primary objective is to estimate time to fall, covariate effects, or the probability of falling within a prespecified prediction horizon\([55](https://arxiv.org/html/2608.18450#bib.bib84);[68](https://arxiv.org/html/2608.18450#bib.bib64)\)\. To address the challenges above mentioned, we introducePAFIR, aPersonalized andAdaptiveFeature Selection Fall RiskIdentification and pRevention framework \(Fig\.[1](https://arxiv.org/html/2608.18450#S1.F1)\) that formulates fall risk factor identification as a multi\-agent reinforcement learning feature selection problem over longitudinal multimodal health data\. PAFIR models heterogeneous clinical assessments and high\-resolution activity data while accounting for temporal dynamics and individual variability\. By jointly capturing structural relationships among risk factors and their evolution over time, PAFIR learns adaptive feature relevance from longitudinal data\. An RL policy guided by sparse fall outcomes and proxy fall risk appraisal measures enables personalized and group\-aware identification of evolving fall risk factors to support timely prevention\. #### Contributions Overview - •Hierarchical structure for fall risk feature selection in the clinical trial: We embedded a hierarchical structure in a topology\-aware reinforcement learning framework to effectively capture and represent multi\-level dependencies among features, which is crucial for fall risk screening, where factors are naturally organized hierarchically\. For example, within fear\-of\-falling assessments, questionnaire items measuring activity avoidance and balance confidence belong to the same higher\-level domain while reflecting distinct behavioral and perceptual contributors to fall risk\. - •Adaptive comparison\-driven feature selection mechanism: We propose an adaptive comparison\-driven feature selection mechanism that decomposes the selection process into candidate generation, reference\-based comparison, and iterative refinement\. Instead of evaluating features independently, the method assesses their incremental contribution by comparing candidate feature groups against reference groups, enabling more reliable fall risk factor identification under complex healthcare settings\. - •Sparse\-aware rewards design with clinical proxies: We introduce a sparse\-aware reward design that augments sparse fall\-event signals with clinically informed proxy outcomes \(i\.e\., FES\-I and BBS\), allowing the PAFIR framework to capture broader fall\-risk patterns and improve the robustness of feature selection\. #### Generalizable Insights about Machine Learning in the Context of Healthcare The proposed framework can generalize to settings that require identifying clinically relevant risk factors from heterogeneous and multimodal data, including structured clinical assessments, behavioral measures, and high\-frequency sensor signals\. Such settings are common in healthcare, where patient information is collected across diverse modalities and scales, and effective modeling requires integrating these sources while preserving their structural relationships\. Moreover, the approach accommodates both discrete and continuous outcomes, ranging from binary events \(e\.g\., disease onset or adverse events\) to continuous measures \(e\.g\., functional decline or physiological indicators\), enabling a unified and flexible framework for modeling feature relevance across diverse clinical settings\. In addition, the framework naturally extends to sequential decision\-making settings, where feature relevance evolves over time and depends on prior observations and selections\. This is particularly important in longitudinal healthcare data, where risk factors may change gradually or abruptly, and adaptive modeling is required to capture temporally evolving patterns\. Finally, the use of informative feedback mechanisms, such as incorporating proxy or auxiliary signals, enables effective learning in scenarios with sparse or delayed outcomes\. Such conditions are prevalent in healthcare applications, where clinically meaningful events are often rare, and leveraging additional signals can improve stability and guide learning toward meaningful trajectories\. Together, these insights suggest that integrating multimodal data, supporting diverse outcome types, and enabling adaptive sequential learning are key to developing generalizable machine learning methods for healthcare\. ## 2Related Work Fall\-risk assessment traditionally relies on clinically interpretable screening and functional measures, including the Timed Up and Go test, balance assessments, and STEADI\-related indicators\. With the increasing availability of wearable sensors and longitudinal cohorts\([57](https://arxiv.org/html/2608.18450#bib.bib36)\), machine\-learning studies have incorporated physical activity, gait, cognitive, psychological, and functional features for broader risk characterization\([18](https://arxiv.org/html/2608.18450#bib.bib37)\)\. However, many existing applications still use fixed feature sets or collapse longitudinal observations into visit\-level summaries\([5](https://arxiv.org/html/2608.18450#bib.bib35);[62](https://arxiv.org/html/2608.18450#bib.bib38)\)\. Statistical methods provide an established framework for longitudinal fall\-outcome analysis\. Time\-varying and recurrent\-event survival models accommodate changing covariates and repeated falls, frailty models account for unobserved heterogeneity\([27](https://arxiv.org/html/2608.18450#bib.bib63)\), joint models link longitudinal biomarker trajectories with event outcomes\([78](https://arxiv.org/html/2608.18450#bib.bib61);[17](https://arxiv.org/html/2608.18450#bib.bib60);[9](https://arxiv.org/html/2608.18450#bib.bib62)\), and competing\-risk models address terminal events such as death\([36](https://arxiv.org/html/2608.18450#bib.bib71)\)\. Feature selection can be incorporated through LASSO, Group LASSO, longitudinal penalization, or boosting\([75](https://arxiv.org/html/2608.18450#bib.bib77);[91](https://arxiv.org/html/2608.18450#bib.bib78);[85](https://arxiv.org/html/2608.18450#bib.bib79);[14](https://arxiv.org/html/2608.18450#bib.bib65)\)\. These approaches are well suited to estimating covariate effects, event risk, or survival probabilities under a specified outcome model\([29](https://arxiv.org/html/2608.18450#bib.bib66);[68](https://arxiv.org/html/2608.18450#bib.bib64);[66](https://arxiv.org/html/2608.18450#bib.bib67)\)\. PAFIR instead focuses on adaptively updating feature sets from visit\-level assessments and wearable sequences as new fall or proxy feedback becomes available\. Healthcare data are inherently multimodal, irregular, and hierarchically structured, combining sparse events, wearable sensor streams, and structured assessments\([21](https://arxiv.org/html/2608.18450#bib.bib44);[20](https://arxiv.org/html/2608.18450#bib.bib39)\)\. Many clinical surveys exhibit hierarchical organization, such as cognitive domains composed of multiple subscales\([61](https://arxiv.org/html/2608.18450#bib.bib42)\), yet most models ignore these dependencies, leading to information loss\([51](https://arxiv.org/html/2608.18450#bib.bib43)\), particularly in longitudinal settings\. Recent work on hierarchical structural encoding\([44](https://arxiv.org/html/2608.18450#bib.bib14)\)motivates the need for personalized, temporally adaptive modeling, which our framework aims to address\. To address these limitations, recent work has explored dynamic and adaptive feature selection strategies\. Reinforcement learning has emerged as a promising approach for dynamic feature selection and sequential decision\-making in healthcare, with applications such as adaptive screening\([90](https://arxiv.org/html/2608.18450#bib.bib41)\)and treatment optimization in critical care\([32](https://arxiv.org/html/2608.18450#bib.bib40)\)\. For feature selection, TTG\([30](https://arxiv.org/html/2608.18450#bib.bib6)\)formulates feature transformation as a graph construction problem and applies RL to search for optimal subsets, but lacks temporal modeling and assumes static relationships\. Structure\-Aware Transformer \(SAT\)\([6](https://arxiv.org/html/2608.18450#bib.bib7)\)improves graph\-based modeling by extracting node\-centric subgraphs before attention computation, yet it does not capture evolving longitudinal or personalized dynamics\. Topology\-Aware Reinforcement Learning \(TAR\)\([89](https://arxiv.org/html/2608.18450#bib.bib8)\)introduces dynamic graph adaptation through reinforcement\-based feature space reconstruction\. However, existing adaptive feature\-selection methods provide important foundations for graph\-based or sequential feature\-space optimization, but they are not designed specifically for the combination considered here: clinically defined hierarchical feature groups, visit\-level structured assessments, minute\-level wearable sequences, asynchronous information updates, and sparse fall outcomes supplemented by proxy feedback\. PAFIR integrates these elements within a unified adaptive feature\-selection framework for longitudinal multimodal health data\. ## 3Clinical Trial for Fall Prevention:Physio\-fEedbackExercise pRogram We are motivated by the heterogeneous and longitudinal data structure \(Fig\.[2](https://arxiv.org/html/2608.18450#S1.F2)\) from thePhysio\-fEedbackExercise pRogram \(PEER\) study\([71](https://arxiv.org/html/2608.18450#bib.bib20)\), a two\-arm clustered randomized controlled clinical trial designed to evaluate a technology\-assisted, body and mind intervention for fall prevention and physical activity promotion among older adults in a free\-living environment \(clinicaltrials\.gov ID: NCT05778604\)\. The analyzed cohort includes 341 participants, aged 61 to 89 years, from 15 senior living sites in Central Florida\. In this two\-arm cluster randomized trial, senior living centers were randomly assigned to either the PEER intervention arm or the control arm, with all participants within each center receiving the same allocation\. Each participant completed four clinical assessments: baseline at week 0 \(T1\), post\-intervention at week 9 \(T2\), 2 months \(60 days\) post\-intervention \(T3\), and 6 months post\-intervention \(T4\)\. All assessments and wearable\-monitoring periods followed the scheduled visit protocol and were not triggered by fall events\. Following each scheduled visit, participants were asked to wear the accelerometer for seven consecutive days, regardless of whether a fall had occurred\. At each visit, participants contributed three complementary categories of data\. First,structured survey\-based assessmentscaptured physical, cognitive, and functional fall risk factors using validated instruments \(e\.g\., the 7\-item Short Fall Efficacy Scale\), resulting in low\-frequency, visit\-level tabular measurements\. Second,clinical measurements, including body composition, grip strength, and standardized evaluations of muscle strength and balance, provided objective physiological measures at each visit\. Third, to characterize physical activity in naturalistic settings, participants wore ActiGraph triaxial accelerometers for seven consecutive days following each visit, generating high\-frequency, minute\-leveltime\-series physical activity datasuch as step counts, vector magnitude \(VM\), and posture states \(e\.g\., sitting, standing, and lying\)\.Fall incidenceserved as the primary outcome and was recorded as self\-reported, time\-stamped events, resulting in sparse and irregular event sequences for most participants\. No participant in the analyzed cohort died before the end of follow\-up; therefore, death\-related competing\-risk censoring did not arise in the current experiments\. Together, these multimodal data spanning survey\-based, clinical, and wearable\-derived measurements across multiple temporal resolutions enable the identification of personalized and dynamically evolving fall risk factors, supporting adaptive fall risk identification and prevention\. ## 4PAFIR Framework ThePersonalized andAdaptiveFeature selection framework for fall riskIdentification and pRevention \(PAFIR; Fig\.[1](https://arxiv.org/html/2608.18450#S1.F1)\) is designed to tackle the challenges of selecting effective fall risk factors from longitudinal, multimodal health data and developing personalized prevention strategies for healthcare applications\. PAFIR integrates three key innovations: \(1\) hierarchical structure linking feature groups to individual features; \(2\) a comparison\-driven feature selection mechanism that iteratively contrasts candidate and reference feature groups to refine the identification of clinically relevant features; and \(3\) a sparse\-aware reward design with clinical proxies\.PAFIR optimizes RL policies for outcome\-relevant feature selection, enhancing the interpretability and effectiveness of longitudinal health modeling in real\-world, dynamic care environments\. Formally, letiiindex participants and let𝒳\\mathcal\{X\}denote the complete feature space, including structured clinical variables and wearable\-derived temporal features\. PAFIR learns a feature\-selection policy shared across the training population and applies it to each participant\-specific state to produce an individualized selected feature set\. The indexttdenotes an asynchronous information update\. Let𝐬i,t\\mathbf\{s\}\_\{i,t\}denote the multimodal state available for participantiiat updatett, and letFi,t−1⊆𝒳F\_\{i,t\-1\}\\subseteq\\mathcal\{X\}denote the active feature set carried into that update\. The three coordinated agent policies define a distribution over the composite feature\-selection action,πθ\(ai,t∣𝐬i,t,Fi,t−1\)\\pi\_\{\\theta\}\(a\_\{i,t\}\\mid\\mathbf\{s\}\_\{i,t\},F\_\{i,t\-1\}\), whereai,ta\_\{i,t\}consists of candidate selection, reference selection, and an add, remove, or retain operation\. Applyingai,ta\_\{i,t\}produces an updated feature setFi,tF\_\{i,t\}, which may contain more, fewer, or the same number of features asFi,t−1F\_\{i,t\-1\}\. At each update,yi,ty\_\{i,t\}denotes the available outcome signal: the observed fall outcome for a fall\-related update, or a visit\-level clinical proxy, such as FES\-I or BBS, for a scheduled assessment update\. ### 4\.1Hierarchical Multimodal Structure for Fall Risk Feature Selection To effectively model both structural clinical assessments and temporal dynamics in time\-series wearable sensor data, we construct a state representation𝐬i,t\\mathbf\{s\}\_\{i,t\}that integrates structured clinical features and time\-series physical activity signals\. This design captures both the organization of clinical variables and the evolution of patient status across visits\. #### Hierarchical Structural Representation Longitudinal clinical data and time\-series physical activity signals form a structured and multimodal representation of patient health\. Features are organized into clinically meaningful feature groups based on their functional roles\. For example, the body composition group includes variables such as BMI and weight, the physical functional group includes the measurements of grip strength, and the psychological group includes assessments such as the FES, which captures fear of falling\. Features within each group are often correlated and reflect related aspects of participant health, while relationships across groups capture interactions among broader fall\-risk domains\. Deterioration in physical performance often co\-occurs with elevated fear of falling, and cognitive decline frequently accompanies functional limitation\. Capturing both levels is essential for identifying which risk factors are clinically relevant and how they jointly evolve over time\. We partition the feature space𝒳\\mathcal\{X\}intoMMclinically meaningful feature sets,\{𝒱1,𝒱2,…,𝒱M\}\\\{\\mathcal\{V\}\_\{1\},\\mathcal\{V\}\_\{2\},\\dots,\\mathcal\{V\}\_\{M\}\\\}, where each𝒱m⊆𝒳\\mathcal\{V\}\_\{m\}\\subseteq\\mathcal\{X\}consists of a subset of related features, and⋃m=1M𝒱m=𝒳\\bigcup\_\{m=1\}^\{M\}\\mathcal\{V\}\_\{m\}=\\mathcal\{X\}\. Within each clinically defined group, features are represented as nodes in an intra\-group graph𝒢m=\(𝒱m,ℰm\)\\mathcal\{G\}\_\{m\}=\(\\mathcal\{V\}\_\{m\},\\mathcal\{E\}\_\{m\}\)\. The feature groups are predefined according to the PEER assessment domains, and every pair of features within the same domain is connected, forming a fully connected intra\-group graph\. This construction encodes shared clinical\-domain membership and enables related features to be evaluated jointly\. Across groups, we further define an inter\-group topology𝒯=\(𝒰,𝒲\)\\mathcal\{T\}=\(\\mathcal\{U\},\\mathcal\{W\}\), where each nodeum∈𝒰u\_\{m\}\\in\\mathcal\{U\}represents a feature group𝒱m\\mathcal\{V\}\_\{m\}\. For participantiiat updatett, letΦi,t\(m\)=∑v∈𝒱mϕi,t,v\\Phi\_\{i,t\}^\{\(m\)\}=\\sum\_\{v\\in\\mathcal\{V\}\_\{m\}\}\\phi\_\{i,t,v\}denote the initial importance score of groupmm, whereϕi,t,v\\phi\_\{i,t,v\}denotes the importance of featurevv\. These initial importance scores are used to construct the initial structural representation and to initialize the reward reference and feature\-selection process\. Each weighted edgewmn∈𝒲w\_\{mn\}\\in\\mathcal\{W\}reflects the co\-variation in initial importance between groupsmmandnn, that is, the extent to which their aggregated initial importance scores increase or decrease together across participants and updates,wmn=Corr\(i,t\)∈𝒟train\(Φi,t\(m\),Φi,t\(n\)\)w\_\{mn\}=\\operatorname\{Corr\}\_\{\(i,t\)\\in\\mathcal\{D\}\_\{\\mathrm\{train\}\}\}\\\!\\left\(\\Phi\_\{i,t\}^\{\(m\)\},\\Phi\_\{i,t\}^\{\(n\)\}\\right\), where the correlation is computed across available participant\-update observations in the training data\. For each featurev∈𝒱mv\\in\\mathcal\{V\}\_\{m\}, we construct a structural representation𝐡i,t,v\\mathbf\{h\}\_\{i,t,v\}that captures its relationships with other features in the same clinical domain for participantiiat updatett\. Let𝐇i,t=\{𝐡i,t,v:v∈𝒱m,m=1,…,M\}\\mathbf\{H\}\_\{i,t\}=\\left\\\{\\mathbf\{h\}\_\{i,t,v\}:v\\in\\mathcal\{V\}\_\{m\},\\;m=1,\\ldots,M\\right\\\}denote the collection of structural feature representations at updatett\. These representations summarize within\-group dependencies and group\-level organization, providing a compact encoding of the structured feature space\. Detailed descriptions of the encoding procedure are provided in Appendix[C\.1](https://arxiv.org/html/2608.18450#A3.SS1)\. #### Temporal Representation In addition to structured clinical features, physical activity data are inherently time\-series observations, capturing fine\-grained behavioral dynamics across time\. While clinical variables are typically recorded at discrete visits, physical activity signals provide continuous measurements between visits, offering complementary information about patient behavior\. Let𝐗i,:,b\(τ\)∈ℝL×1\\mathbf\{X\}\_\{i,:,b\}^\{\(\\tau\)\}\\in\\mathbb\{R\}^\{L\\times 1\}denote the time\-series data collected for participantiifollowing visitτ\\tau, whereLLis the sequence length andbbdenotes the corresponding temporal feature\. At the corresponding updatett, each time series is mapped to a latent representation𝐳i,t,b\\mathbf\{z\}\_\{i,t,b\}that captures temporal patterns such as step counts and physical activity vector magnitude\. A detailed description of the construction of the temporal representation𝐳i,t,b\\mathbf\{z\}\_\{i,t,b\}is provided in Appendix[C\.2](https://arxiv.org/html/2608.18450#A3.SS2)\. #### State Construction Let𝐙i,t=\{𝐳i,t,b:b∈ℬ\}\\mathbf\{Z\}\_\{i,t\}=\\left\\\{\\mathbf\{z\}\_\{i,t,b\}:b\\in\\mathcal\{B\}\\right\\\}denote the set of temporal embeddings available for participantiiat updatett, whereℬ\\mathcal\{B\}denotes the set of time\-series physical activity features\. The final state representation is constructed by integrating structural and temporal information, 𝐬i,t=f\(𝐇i,t,𝐙i,t\),\\displaystyle\\mathbf\{s\}\_\{i,t\}=f\\left\(\\mathbf\{H\}\_\{i,t\},\\mathbf\{Z\}\_\{i,t\}\\right\),\(1\)wheref\(⋅\)f\(\\cdot\)represents a fusion mechanism that aligns structured features with temporal dynamics\. This formulation allows𝐬i,t\\mathbf\{s\}\_\{i,t\}to capture both the topology of the feature space and the temporal evolution of patient behavior across visits, supporting more stable and context\-aware feature selection\. ### 4\.2Adaptive Comparison\-Driven Feature Selection Mechanism We design an adaptive comparison\-driven feature selection mechanism to iteratively identify outcome\-relevant features under complex and correlated clinical settings\. The key idea is to decompose feature selection into a sequence of structured decisions that balance candidate discovery, comparative evaluation, and refinement\([92](https://arxiv.org/html/2608.18450#bib.bib80)\)\. To implement this process, we adopt a three\-agent feature selection framework\([89](https://arxiv.org/html/2608.18450#bib.bib8)\), where each agent is responsible for a distinct role in the stepwise selection procedure\. Specifically, the candidate agent proposes a feature group for evaluation, the reference agent selects a competing group for comparison, and the operation agent determines whether the active feature set should be updated based on their relative outcome relevance \(in Fig\.[1](https://arxiv.org/html/2608.18450#S1.F1)\)\. A single agent would need to select the candidate group, reference group, and operation jointly\. The three\-agent formulation instead factorizes this action into smaller conditional decisions\. The agents are executed sequentially, share the same participant\-specific state and reward, and condition each downstream decision on the preceding actions\. They therefore form a coordinated feature\-selection policy rather than independent decision makers\. LetFi,t−1⊆𝒳F\_\{i,t\-1\}\\subseteq\\mathcal\{X\}denote the active feature set for participantiibefore updatett, and let𝒞=\{𝒱1,𝒱2,…,𝒱M\}\\mathcal\{C\}=\\\{\\mathcal\{V\}\_\{1\},\\mathcal\{V\}\_\{2\},\\ldots,\\mathcal\{V\}\_\{M\}\\\}denote the collection of feature groups derived from the hierarchical structure\. At each update, the candidate agent selects a feature group based on its potential relevance to the outcome,Ci,tcand∼πθ1\(C∣𝐬i,t,Fi,t−1\),Ci,tcand∈𝒞C\_\{i,t\}^\{\\mathrm\{cand\}\}\\sim\\pi\_\{\\theta\_\{1\}\}\\left\(C\\mid\\mathbf\{s\}\_\{i,t\},F\_\{i,t\-1\}\\right\),\\quad C\_\{i,t\}^\{\\mathrm\{cand\}\}\\in\\mathcal\{C\}\. Rather than evaluating the candidate group in isolation, the reference agent selects a comparison group from the remaining feature space,Ci,tref∼πθ2\(C∣𝐬i,t,Fi,t−1,Ci,tcand\)C\_\{i,t\}^\{\\mathrm\{ref\}\}\\sim\\pi\_\{\\theta\_\{2\}\}\\left\(C\\mid\\mathbf\{s\}\_\{i,t\},F\_\{i,t\-1\},C\_\{i,t\}^\{\\mathrm\{cand\}\}\\right\),Ci,tref∈𝒞∖\{Ci,tcand\}C\_\{i,t\}^\{\\mathrm\{ref\}\}\\in\\mathcal\{C\}\\setminus\\\{C\_\{i,t\}^\{\\mathrm\{cand\}\}\\\}\. The operation agent then selects an add, remove, or retain operation,oi,t∼πθ3\(o∣𝐬i,t,Fi,t−1,Ci,tcand,Ci,tref\)o\_\{i,t\}\\sim\\pi\_\{\\theta\_\{3\}\}\\left\(o\\mid\\mathbf\{s\}\_\{i,t\},F\_\{i,t\-1\},C\_\{i,t\}^\{\\mathrm\{cand\}\},C\_\{i,t\}^\{\\mathrm\{ref\}\}\\right\), whereoi,t∈\{add,remove,retain\}o\_\{i,t\}\\in\\\{\\mathrm\{add\},\\mathrm\{remove\},\\mathrm\{retain\}\\\}\. The resulting composite action isai,t=\(Ci,tcand,Ci,tref,oi,t\)a\_\{i,t\}=\\left\(C\_\{i,t\}^\{\\mathrm\{cand\}\},C\_\{i,t\}^\{\\mathrm\{ref\}\},o\_\{i,t\}\\right\)\. The candidate and reference groups are evaluated relative to the current active feature set using the available outcome feedback\. The same comparison procedure is subsequently applied within each retained group to refine the selection at the individual\-feature level while preserving the hierarchical structure\. ### 4\.3Sparse\-Aware Reward Design With Clinical Proxies Learning an effective feature selection policy from fall incidence is challenging due to the sparsity and temporal irregularity of observed events\. In many cases, participants do not experience a fall across multiple visits, resulting in limited and delayed feedback if rewards are defined solely based on fall occurrence\. To address this, we design a dynamic reward mechanism that provides continuous and adaptive feedback throughout the selection process\. At the beginning of training, baseline feature\-importance scores are used to provide an informed initial reward reference and selection prior\. During training, the reward is computed using the outcome information available at each update\. In addition to fall incidence, the reward signal incorporates clinically validated proxy measures that are available at every visit and are known to be associated with fall risk, such as the Falls Efficacy Scale\-International \(FES\-I,[86](https://arxiv.org/html/2608.18450#bib.bib3)\) and the BTrackS Balance Tracking System score \(BBS,[35](https://arxiv.org/html/2608.18450#bib.bib2)\)\. These proxy signals, as fall risk appraisal, provide continuous supervision even in the absence of observed falls, allowing the model to receive meaningful feedback at every step\([72](https://arxiv.org/html/2608.18450#bib.bib1)\)\. Specifically, when a fall incident occurs between visits \(e\.g\., between Visit 2 and Visit 3\), the reward is immediately adjusted to reflect the contribution of the selected features\. Feature groups that are more strongly associated with fall risk receive higher rewards, while those that do not contribute are penalized\. At scheduled assessment updates without fall feedback, the reward is evaluated using an available clinical proxy, such as FES\-I or BBS\. This design ensures that rare but clinically important fall events provide strong learning signals when they occur, while proxy\-based signals maintain stable feedback throughout the remaining visits\. The reward function is defined as ri,t=Perf\(Fi,t,yi,t\)−pi,t−1,r\_\{i,t\}=\\operatorname\{Perf\}\\left\(F\_\{i,t\},y\_\{i,t\}\\right\)\-p\_\{i,t\-1\},\(2\)wherePerf\(Fi,t,yi,t\)\\operatorname\{Perf\}\(F\_\{i,t\},y\_\{i,t\}\)denotes the evaluation measure used for the fall outcome when fall feedback is available and for the clinical proxy outcome otherwise\. The termpi,t−1p\_\{i,t\-1\}denotes the best previously observed performance for the corresponding type of outcome feedback, following[92](https://arxiv.org/html/2608.18450#bib.bib80)\. Each agent is trained using a separate deep Q\-network\. For agentj∈\{cand,ref,op\}j\\in\\\{\\mathrm\{cand\},\\mathrm\{ref\},\\mathrm\{op\}\\\}, the temporal\-difference loss at updatettis, ℒi,t\(j\)=\(Qj\(𝐬i,t,ai,t\(j\)\)−\[ri,t\+γmaxa′Qj\(𝐬i,t\+1,a′\)\]\)2,\\displaystyle\\mathcal\{L\}\_\{i,t\}^\{\(j\)\}=\\Bigg\(Q\_\{j\}\\left\(\\mathbf\{s\}\_\{i,t\},a\_\{i,t\}^\{\(j\)\}\\right\)\-\\bigg\[r\_\{i,t\}\+\\gamma\\max\_\{a^\{\\prime\}\}Q\_\{j\}\\left\(\\mathbf\{s\}\_\{i,t\+1\},a^\{\\prime\}\\right\)\\bigg\]\\Bigg\)^\{2\},\(3\)whereQjQ\_\{j\}denotes the action\-value function of agentjj,ai,t\(j\)a\_\{i,t\}^\{\(j\)\}denotes the corresponding agent action,ri,tr\_\{i,t\}is the shared reward, andγ\\gammais the discount factor\. All Q\-networks are randomly initialized and trained using the temporal\-difference loss\. The baseline feature importance scores are used to initialize the reward and the initial feature selection\. This reward design allows the model to learn from both sparse fall events and continuous clinical signals, making the learning process more stable while maintaining alignment with clinically meaningful patterns\. ## 5In\-Field Experiment Evaluations PEER Fall Prevention Data\.We evaluate PAFIR using longitudinal, multimodal data from the PEER study collected in real\-world settings, emphasizing adaptive feature selection across repeated clinical visits\. The dataset includes 341 community\-dwelling participants assessed at four time points, yielding 1,364 participant–visit observations\. Following each scheduled visit, participants wore sensors continuously for seven consecutive days, producing minute\-level wearable data and totaling over 2\.5 million time\-stamped observations across all participants and visits\. At each scheduled visit, participants contribute high\-dimensional structured clinical assessments, spanning physical, cognitive, psychological, and functional domains\. After preprocessing, the structured modality comprises 587 fall\-related features organized into 35 clinically meaningful feature groups, including body composition, physical activity, cognitive function, balance confidence, and psychological status\. Each participant–visit instance is modeled as a graph\-structured input, which is encoded by a graph encoder to capture hierarchical and correlational dependencies among risk factors\. In parallel, minute\-level wearable sensor data, including step counts, vector magnitude \(VM\), and posture states \(sitting, standing, and lying\), are processed by a time\-series encoder to model temporal dynamics within each seven\-day monitoring period\. The resulting temporal representations are aligned with visit\-level structured embeddings and fused via a cross\-attention Transformer module, yielding a hierarchical latent state that integrates structural and temporal information and serves as the input to the reinforcement learning component\. Fall incidence events are recorded as self\-reported, time\-stamped outcomes, providing sparse and delayed supervision signals that guide adaptive learning of feature relevance across visits\. The proxy outcomes, namely the FES\-I and BBS scores, are assessed at each clinical visit\. PAFIR Implementation\.Reinforcement learning feature selection operates at the visit level or when a fall incidence occurs, enabling the model to update feature importance dynamically as new temporal and structural information becomes available\. The structured feature space is encoded using a six\-layer graph encoder with a hidden dimension of 128, residual connections, and a dropout rate of 0\.5\. Mean pooling and Laplacian positional encoding with up to ten eigenvectors are applied to aggregate node\-level representations\. The temporal encoder uses four stacked Transformer encoder blocks, each with four attention heads and a feed\-forward dimension of 256, with wearable inputs projected through a shared linear embedding layer\. The model is trained for 100 epochs using the Adam optimizer with a learning rate of 0\.001\. The Q\-network of each agent is optimized using the temporal\-difference loss defined in the previous section\. Binary cross\-entropy loss is used for the fall\-outcome classification model employed in the fall\-based reward calculation\. All experiments are conducted on NVIDIA Quadro RTX 6000 GPUs using CUDA 13\.0\. Each random\-seed run requires approximately 101 seconds, and the complete 20\-seed longitudinal stability analysis requires approximately 34 minutes\. The peak allocated GPU memory is approximately 240 MB\. Computational resource requirements are reported in Appendix[B\.5](https://arxiv.org/html/2608.18450#A2.SS5)and Table[6](https://arxiv.org/html/2608.18450#A2.T6)\. ### 5\.1Results: Recommendations for Fall Prevention in PEER Study #### Community\-Level Recommendations for Fall Prevention\. \\subfigure\[PEER intervention site:Kinneret\]\\subfigure\[Control site:LCA\] Figure 3:Top fall risk factors identified across four visits at the community level\.We first describe how PAFIR produces the longitudinal community\-level results shown in Fig\.[3](https://arxiv.org/html/2608.18450#S5.F3)\. PAFIR is applied to multimodal data collected at four visits \(T1–T4\), including clinical assessments, survey\-based measures, and wearable\-derived physical activity features\. The model is trained using data from all participants at the community level\. At each visit, PAFIR selects important fall risk factors by considering both changes over time and relationships between feature groups\. This procedure produces visit\-level feature selection results for each site, allowing us to examine how selected fall risk factors evolve over time in both the intervention and control communities\. Across both sites, PAFIR consistently identifiesphysical activity,body composition, anddemographic informationacross visits, which aligns with established evidence that mobility\-related measures and body composition are among the most persistent predictors of falls in community\-dwelling older adults\([74](https://arxiv.org/html/2608.18450#bib.bib32);[31](https://arxiv.org/html/2608.18450#bib.bib85)\)\. At the Kinneret intervention site \(Fig\.[3](https://arxiv.org/html/2608.18450#S5.F3)[3](https://arxiv.org/html/2608.18450#S5.F3)\), PAFIR identifies a broader set of fall risk features beyond physical performance, including psychological and behavioral measures such asanxiety screening,mindful attention awareness,regulatory focus, andpsychological inhibition sensitivity\(BIS\), which are selected across multiple visits\. This pattern is consistent with prior evidence that psychological factors, including fear of falling, attentional control, and behavioral regulation, are independently associated with fall risk in community\-dwelling older adults\([88](https://arxiv.org/html/2608.18450#bib.bib86);[67](https://arxiv.org/html/2608.18450#bib.bib87)\)\. The consistent selection of psychological and behavioral features at Kinneret across all four visits suggests that PEER’s cognitive and psychological components actively engage these domains as modifiable risk factors\([71](https://arxiv.org/html/2608.18450#bib.bib20);[73](https://arxiv.org/html/2608.18450#bib.bib90)\), making them detectable and trackable by PAFIR throughout the intervention period\. This is in line with evidence that multicomponent interventions combining physical exercise with cognitive\-behavioral elements produce greater reductions in fall risk than single\-component physical programs alone\([42](https://arxiv.org/html/2608.18450#bib.bib59)\)\. In contrast, the LCA control site \(Fig\.[3](https://arxiv.org/html/2608.18450#S5.F3)[3](https://arxiv.org/html/2608.18450#S5.F3)\) predominantly selects traditional physical and functional risk factors, includingfrailty,physical activity,body composition,age\-related medical conditions, andfunctional performance measures\(Sit\-to\-Stand, RAPA\), across visits\. This pattern reflects the well\-documented role of physical frailty, muscle function, and chronic health conditions as core fall risk determinants in older adult populations not receiving structured behavioral interventions\([7](https://arxiv.org/html/2608.18450#bib.bib88);[64](https://arxiv.org/html/2608.18450#bib.bib89)\)\. The absence of psychological and behavioral features at LCA reflects the profile expected in a population not receiving structured multidimensional intervention, where conventional physical and functional factors dominate the fall risk landscape\([40](https://arxiv.org/html/2608.18450#bib.bib58);[52](https://arxiv.org/html/2608.18450#bib.bib33)\)\. The divergence in feature profiles between the two sites provides indirect evidence of PEER’s effectiveness, demonstrating that an effective intervention not only addresses physical fall risk but also renders psychological and behavioral risk factors consistently detectable over time\([42](https://arxiv.org/html/2608.18450#bib.bib59)\)\. These findings reinforce the clinical value of integrating psychological engagement and cognitive reframing into community\-level fall prevention, complementing physical activity and body composition monitoring as a more comprehensive risk management strategy\. #### Temporal Evolving of Fall Risk Factors Figure 4:Population\-level fall risk contributions across study visits relative to a fall event\. Visits are aligned by their order before and after the fall, and the vertical dashed line indicates the fall event\. Markers at the dashed line denote the fall\-related update, whereas the remaining markers denote scheduled visit\-level contributions\.We next examine how model\-derived feature contributions vary across visits relative to a recorded fall event\. Participant trajectories are aligned according to the visit order before and after the fall, with V−2\-2and V−1\-1representing pre\-fall visits and V\+1\+1and V\+2\+2representing post\-fall visits\. As shown in Fig\.[4](https://arxiv.org/html/2608.18450#S5.F4), several features exhibit distinct temporal patterns\. Activity\-related features, includingSteps,Vector Magnitude,Lying Time,Off\-Body Time, andSitting Time, receive relatively high contributions before or around the fall, consistent with the established importance of mobility and activity patterns in fall\-risk characterization\([71](https://arxiv.org/html/2608.18450#bib.bib20);[69](https://arxiv.org/html/2608.18450#bib.bib93)\)\. In contrast, the contributions ofRace/Ethnicity,Sleeping Pill Use,AgeIAT, andSmokingincrease across later visits, suggesting that demographic, behavioral, and functional factors remain relevant during post\-fall follow\-up\([84](https://arxiv.org/html/2608.18450#bib.bib92);[8](https://arxiv.org/html/2608.18450#bib.bib91);[74](https://arxiv.org/html/2608.18450#bib.bib32)\)\. The increasing contribution ofAgeIATmay also reflect the broader role of cognitive and psychological factors in fall\-related outcomes\([67](https://arxiv.org/html/2608.18450#bib.bib87)\)\. These patterns represent changes in the relevance assigned to each feature by PAFIR rather than changes in the underlying raw feature values, illustrating how the framework updates feature priorities across longitudinal assessments\. #### Personalized Recommendation for Fall Prevention Figure 5:Personalized Temporal Evolution of Feature Importance Before and After Fall forSubject 1017\. The vertical dashed line indicates the recorded fall event\. Circles denote scheduled visit\-level contributions, whereas squares denote contributions at the fall\-related update\.We next present a personalized analysis to illustrate how PAFIR supports individual\-level longitudinal monitoring through temporally evolving and adaptive feature selection\. Figure[5](https://arxiv.org/html/2608.18450#S5.F5)visualizes the trajectories of selected fall\-related features for Subject 1017 across four study visits, with a recorded fall occurring shortly after the T2 visit\. Figure[5](https://arxiv.org/html/2608.18450#S5.F5)reveals a clear transition in feature contributions across the pre\- and post\-fall periods\. At T1,Lying Time\(Physical Performance Assessments\) exhibits the highest contribution, indicating that sedentary behavior is a prominent fall\-related feature at baseline\([84](https://arxiv.org/html/2608.18450#bib.bib92);[71](https://arxiv.org/html/2608.18450#bib.bib20)\)\. At T2, shortly before the fall,Hand Grip Strength \(R\)\(Physical Performance Assessments\) receives the highest visit\-level contribution, whereasLying Timeis not selected at the scheduled visit\. This pre\-fall pattern, centered on physical activity and muscle function, is consistent with established fall\-risk factors among older adults\([74](https://arxiv.org/html/2608.18450#bib.bib32)\)\. Following the fall, the selected feature profile broadens across multiple domains\. At the fall\-related update shortly after T2,Lying Timereceives the largest contribution, whileHand Grip Strengthremains prominent\. At T3,Lying Timereaches its highest observed contribution, alongsideMindfulness \(MAAS\-9\)\(Questionnaires–Psychological\), whileSTEADI Item 11\(Clinical\) is also selected\. At T4,Chronotype \(MEQ\)\(Questionnaires–Behavior\) becomes the most prominent selected feature\([88](https://arxiv.org/html/2608.18450#bib.bib86);[69](https://arxiv.org/html/2608.18450#bib.bib93)\)\. The increased contribution ofLying Timeafter the fall may be consistent with reduced mobility or fear\-related activity restriction, while the emergence of psychological, behavioral, and clinical screening features illustrates the multidimensional nature of post\-fall follow\-up\. From a personalized monitoring perspective, these temporal patterns demonstrate PAFIR’s capacity to generate adaptive, individual\-level feature priorities\. The pre\-fall selections emphasize physical activity and muscle function, whereas the post\-fall profile expands to include physical, psychological, behavioral, and clinical screening domains, suggesting that these areas may warrant broader review during subsequent follow\-up\([74](https://arxiv.org/html/2608.18450#bib.bib32)\)\. ### 5\.2Benchmark and Ablation Analysis Evaluation Metrics\.We compare PAFIR with several state\-of\-the\-art feature selection baselines, including TAR\([89](https://arxiv.org/html/2608.18450#bib.bib8)\), SAT\([6](https://arxiv.org/html/2608.18450#bib.bib7)\), TTG\([30](https://arxiv.org/html/2608.18450#bib.bib6)\), PCA\([45](https://arxiv.org/html/2608.18450#bib.bib76)\), Group LASSO\([91](https://arxiv.org/html/2608.18450#bib.bib78)\), and LASSO\([75](https://arxiv.org/html/2608.18450#bib.bib77)\)\. Participants are randomly divided into 80% training and 20% testing sets, and all longitudinal observations from the same participant are retained in the same partition\. All data\-dependent quantities, including feature\-importance scores and inter\-group correlations, are estimated using the training partition only\. All experiments are repeated 20 times with different random seeds, and the results are reported as mean±\\pmstandard deviation\. We assess the quality of feature selection directly using four complementary metrics\.Feature Recovery Rate \(FRR\)measures the proportion of fall\-relevant features identified\.False Discovery Rate \(FDR\)quantifies the proportion of irrelevant features selected\([3](https://arxiv.org/html/2608.18450#bib.bib72)\)\.Selection Stabilityevaluates consistency across runs\([33](https://arxiv.org/html/2608.18450#bib.bib73)\)\.Correlation Recoveryassesses how well feature dependencies are preserved\([49](https://arxiv.org/html/2608.18450#bib.bib74);[48](https://arxiv.org/html/2608.18450#bib.bib75)\)\. The complete feature\- and group\-level initial reference sets used for evaluation are provided in Appendix[B\.2](https://arxiv.org/html/2608.18450#A2.SS2)\. Table 1:Performance results for fall risk factor selection on the PEER study\.MethodFRR↑\\uparrowFDR↓\\downarrowSelection Stability↑\\uparrowCorrelation Recovery↑\\uparrowPAFIR \(Ours\)0\.874±\\pm0\.0070\.207±\\pm0\.0090\.990±\\pm0\.0030\.974±\\pm0\.014TAR \([89](https://arxiv.org/html/2608.18450#bib.bib8)\)0\.795±\\pm0\.0090\.209±\\pm0\.0080\.989±\\pm0\.0010\.973±\\pm0\.014SAT \([6](https://arxiv.org/html/2608.18450#bib.bib7)\)0\.731±\\pm0\.0100\.311±\\pm0\.0110\.985±\\pm0\.0030\.973±\\pm0\.013TTG \([30](https://arxiv.org/html/2608.18450#bib.bib6)\)0\.664±\\pm0\.0090\.379±\\pm0\.0120\.985±\\pm0\.0040\.972±\\pm0\.013PCA \([45](https://arxiv.org/html/2608.18450#bib.bib76)\)0\.624±\\pm0\.0150\.434±\\pm0\.0120\.978±\\pm0\.0020\.971±\\pm0\.016Group LASSO \([91](https://arxiv.org/html/2608.18450#bib.bib78)\)0\.573±\\pm0\.1540\.667±\\pm0\.0610\.879±\\pm0\.1340\.970±\\pm0\.005LASSO \([75](https://arxiv.org/html/2608.18450#bib.bib77)\)0\.278±\\pm0\.0050\.500±\\pm0\.0000\.990±\\pm0\.0090\.902±\\pm0\.011 #### Performance and Ablation Study Table[1](https://arxiv.org/html/2608.18450#S5.T1)summarizes the overall feature selection performance on the PEER study\. PAFIR achieves the highest Feature Recovery Rate \(FRR\) and Selection Stability, while maintaining a competitive False Discovery Rate \(FDR\)\. This indicates that PAFIR is able to consistently identify clinically relevant fall\-risk features\. Its advantage is especially clear in Correlation Recovery, where PAFIR better preserves the relationships among features compared to other methods\. This is important in practice, as it allows clinicians to interpret how different risk factors interact, rather than treating them as independent variables\. Table 2:Ablation Study for Fall Risk Factors Selection on the PEER Study\.MethodFRR↑\\uparrowFDR↓\\downarrowSelection Stability↑\\uparrowCorrelation Recovery↑\\uparrowPAFIR \(Full\)0\.874±\\pm0\.0070\.207±\\pm0\.0090\.990±\\pm0\.0030\.974±\\pm0\.014w/ Random Feature\-Importance Prior0\.870±\\pm0\.0150\.211±\\pm0\.0080\.980±\\pm0\.0110\.970±\\pm0\.014w/o Hierarchical Structure0\.768±\\pm0\.0100\.229±\\pm0\.0090\.769±\\pm0\.0100\.752±\\pm0\.008w/o Three\-Agent Comparison\-Driven Mechanism0\.678±\\pm0\.0050\.327±\\pm0\.0130\.675±\\pm0\.0070\.765±\\pm0\.033w/o Sparse\-Aware Rewards0\.655±\\pm0\.0110\.325±\\pm0\.0080\.665±\\pm0\.0070\.701±\\pm0\.013 Table[2](https://arxiv.org/html/2608.18450#S5.T2)shows that removing any component leads to consistent performance degradation\. Without the hierarchical structure, FRR and selection stability decrease\. Without the comparison\-driven selection mechanism, FRR decreases and FDR increases\. Without the sparse\-aware reward, FRR and correlation recovery are the lowest\. These results highlight the complementary roles of structural representation, comparison\-driven selection, and reward design\. Replacing the baseline feature\-importance prior with a random prior yields only modest performance changes, indicating that the prior provides useful guidance without determining the learned policy\. Additional results are reported in Appendix[B\.4](https://arxiv.org/html/2608.18450#A2.SS4), Table[5](https://arxiv.org/html/2608.18450#A2.T5), and Appendices[D](https://arxiv.org/html/2608.18450#A4)–[F](https://arxiv.org/html/2608.18450#A6)\. The public benchmark experiments support methodological applicability but do not constitute external clinical validation\. #### Longitudinal Stability in the No\-Fall Subgroup Among participants with no recorded falls, PAFIR shows high group\-level stability across visits, with an overall Jaccard similarity of0\.890±0\.0090\.890\\pm 0\.009\. In contrast, individual\-feature selections vary more substantially\. Detailed group\- and feature\-level results are provided in Appendix[B\.3](https://arxiv.org/html/2608.18450#A2.SS3)and Table[4](https://arxiv.org/html/2608.18450#A2.T4)\. ## 6Conclusion, Future Work, and Limitations This study presents PAFIR, a multi\-level RL\-based feature selection framework designed to support personalized and adaptive fall prevention for older adults\. By modeling the hierarchical structural assessments and aligning evolving physical activities from wearable sensors, PAFIR adaptively identifies temporally dynamic and fall risk factors, which can inform personalized recommendations\. Particularly, PAFIR demonstrates its potential to identify the actionable early warning signal from the in\-field PEER study and to inform personalized monitoring and fall\-prevention planning as risk factors evolve\. #### Future Work and Limitations Future directions for PAFIR include integration with digital health and mHealth platforms, such as Ecological Momentary Assessment \(EMA\) systems, to enable real\-time monitoring and adaptive feedback\. By incorporating high\-frequency self\-reported data on symptoms and behaviors, PAFIR may support just\-in\-time adaptive interventions \(JITAIs\) that deliver personalized prompts when relevant changes in mobility or functional capacity are identified\. This integration could bridge passive sensing with active behavioral support\. Future studies will evaluate PAFIR\-powered digital interventions in community\-dwelling older adults\. External validation is currently constrained by the lack of comparable longitudinal multimodal fall\-prevention cohorts\. ###### acknowledgments\-disclosure\-of\-funding\. We thank the reviewers and the Area Chair for their constructive feedback\. This research was supported by NIH the National Institute on Minority Health and Health Disparities \(R01MD018025\) and the NIH Office of the Director, Chief Officer for Scientific Workforce Diversity \(COSWD\) \(3R01MD018025\-02S1\), and the Learning Institute for Elders at University of Central Florida Richard Tucker Gerontology Applied Research Grant\. ## References - Alsadoonet al\.\(2024\)A\. Alsadoon, G\. Al\-Naymat, and O\. D\. JerewAn architectural framework of elderly healthcare monitoring and tracking through wearable sensor technologies\.Multimedia Tools and Applications83\(26\),pp\. 67825–67870\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p7.1)\. - Barrick and Zimmerman \(2005\)M\. R\. Barrick and R\. D\. ZimmermanReducing voluntary, avoidable turnover through selection\.\.Journal of applied psychology90\(1\),pp\. 159\.Cited by:[§B\.3](https://arxiv.org/html/2608.18450#A2.SS3.p1.1)\. - Benjamini and Hochberg \(1995\)Y\. Benjamini and Y\. HochbergControlling the false discovery rate: a practical and powerful approach to multiple testing\.Journal of the Royal statistical society: series B \(Methodological\)57\(1\),pp\. 289–300\.Cited by:[§5\.2](https://arxiv.org/html/2608.18450#S5.SS2.p1.1)\. - Centers for Disease Control and Prevention \(2026\)Centers for Disease Control and PreventionAbout older adult fall prevention\.Note:[https://www\.cdc\.gov/falls/about/index\.html](https://www.cdc.gov/falls/about/index.html)Accessed: 2026\-01\-27Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p1.1)\. - Chantanachaiet al\.\(2021\)T\. Chantanachai, D\. L\. Sturnieks, S\. R\. Lord, N\. Payne, L\. Webster, and M\. E\. TaylorRisk factors for falls in older people with cognitive impairment living in the community: systematic review and meta\-analysis\.Ageing research reviews71,pp\. 101452\.Cited by:[§2](https://arxiv.org/html/2608.18450#S2.p1.1)\. - Chenet al\.\(2022\)D\. Chen, L\. O’Bray, and K\. BorgwardtStructure\-aware transformer for graph representation learning\.InInternational conference on machine learning,pp\. 3469–3489\.Cited by:[Table 5](https://arxiv.org/html/2608.18450#A2.T5.2.1.6.1),[§D\.1](https://arxiv.org/html/2608.18450#A4.SS1.p1.1),[Table 8](https://arxiv.org/html/2608.18450#A4.T8.2.1.12.2),[Table 8](https://arxiv.org/html/2608.18450#A4.T8.2.1.18.2),[Table 8](https://arxiv.org/html/2608.18450#A4.T8.2.1.6.2),[§2](https://arxiv.org/html/2608.18450#S2.p4.1),[§5\.2](https://arxiv.org/html/2608.18450#S5.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.18450#S5.T1.2.1.4.1.1)\. - Chittrakulet al\.\(2020\)J\. Chittrakul, P\. Siviroj, S\. Sungkarat, and R\. SapbamrerPhysical frailty and fall risk in community\-dwelling older adults: a cross\-sectional study\.Journal of aging research2020\(1\),pp\. 3964973\.Cited by:[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px1.p2.1)\. - Colón\-Emericet al\.\(2024\)C\. S\. Colón\-Emeric, C\. L\. McDermott, D\. S\. Lee, and S\. D\. BerryRisk assessment and prevention of falls in older community\-dwelling adults: a review\.Jama331\(16\),pp\. 1397–1406\.Cited by:[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px2.p1.1)\. - Crowtheret al\.\(2016\)M\. J\. Crowther, T\. M\. Andersson, P\. C\. Lambert, K\. R\. Abrams, and K\. HumphreysJoint modelling of longitudinal and survival data: incorporating delayed entry and an assessment of model misspecification\.Statistics in medicine35\(7\),pp\. 1193–1209\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p8.1),[§2](https://arxiv.org/html/2608.18450#S2.p2.1)\. - Cuevas\-Trisan \(2017\)R\. Cuevas\-TrisanBalance problems and fall risks in the elderly\.Physical Medicine and Rehabilitation Clinics28\(4\),pp\. 727–737\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p3.1)\. - Dautzenberget al\.\(2021\)L\. Dautzenberg, S\. Beglinger, S\. Tsokani, S\. Zevgiti, R\. C\. Raijmann, N\. Rodondi, R\. J\. Scholten, A\. W\. Rutjes, M\. Di Nisio, M\. Emmelot\-Vonk,et al\.Interventions for preventing falls and fall\-related fractures in community\-dwelling older adults: a systematic review and network meta\-analysis\.Journal of the American Geriatrics Society69\(10\),pp\. 2973–2984\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p3.1)\. - Delbaereet al\.\(2010\)K\. Delbaere, J\. C\. Close, H\. Brodaty, P\. Sachdev, and S\. R\. LordDeterminants of disparities between perceived and physiological risk of falling among elderly people: cohort study\.Bmj341\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p3.1)\. - Dobson and Doig \(2003\)P\. D\. Dobson and A\. J\. DoigDistinguishing enzyme structures from non\-enzymes without alignments\.Journal of molecular biology330\(4\),pp\. 771–783\.Cited by:[§D\.2](https://arxiv.org/html/2608.18450#A4.SS2.p1.1)\. - Donget al\.\(2022\)J\. Dong, Y\. Chen, B\. Yao, X\. Zhang, and N\. ZengA neural network boosting regression model based on xgboost\.Applied Soft Computing125,pp\. 109067\.Cited by:[§2](https://arxiv.org/html/2608.18450#S2.p2.1)\. - Duncanet al\.\(1993\)P\. W\. Duncan, J\. Chandler, S\. Studenski, M\. Hughes, and B\. PrescottHow do physiological components of balance affect mobility in elderly men?\.Archives of physical medicine and rehabilitation74\(12\),pp\. 1343–1349\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p3.1)\. - Ecoffetet al\.\(2021\)A\. Ecoffet, J\. Huizinga, J\. Lehman, K\. O\. Stanley, and J\. CluneFirst return, then explore\.Nature590\(7847\),pp\. 580–586\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p7.1)\. - Elashoffet al\.\(2016\)R\. Elashoff N\. Liet al\.Joint modeling of longitudinal and time\-to\-event data\.Chapman and Hall/CRC\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p8.1),[§2](https://arxiv.org/html/2608.18450#S2.p2.1)\. - Fanet al\.\(2024\)L\. Fan, J\. Zhao, Y\. Hu, J\. Zhang, X\. Wang, F\. Wang, M\. Wu, and T\. LinPredicting physical functioning status in older adults: insights from wrist accelerometer sensors and derived digital biomarkers of physical activity\.Journal of the American Medical Informatics Association31\(11\),pp\. 2571–2582\.Cited by:[§2](https://arxiv.org/html/2608.18450#S2.p1.1)\. - Florenceet al\.\(2018\)C\. S\. Florence, G\. Bergen, A\. Atherly, E\. Burns, J\. Stevens, and C\. DrakeMedical costs of fatal and nonfatal falls in older adults\.Journal of the American Geriatrics Society66\(4\),pp\. 693–698\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p1.1)\. - Franklinet al\.\(2024\)J\. B\. Franklin, C\. Marra, K\. Z\. Abebe, A\. J\. Butte, D\. J\. Cook, L\. Esserman, L\. A\. Fleisher, C\. I\. Grossman, N\. E\. Kass, H\. M\. Krumholz,et al\.Modernizing the data infrastructure for clinical research to meet evolving demands for evidence\.JAMA332\(16\),pp\. 1378–1385\.Cited by:[§2](https://arxiv.org/html/2608.18450#S2.p3.1)\. - Ghat \(2023\)A\. B\. R\. GhatFuture evolution of telemedicine: enhancing healthcare accessibility and reliability through the integration of machine learning techniques\.Ph\.D\. Thesis,Dublin, National College of Ireland\.Cited by:[§2](https://arxiv.org/html/2608.18450#S2.p3.1)\. - Giovanniniet al\.\(2022\)S\. Giovannini, F\. Brau, V\. Galluzzo, D\. A\. Santagada, C\. Loreti, L\. Biscotti, A\. Laudisio, G\. Zuccala, and R\. BernabeiFalls among older adults: screening, identification, rehabilitation, and management\.Applied Sciences12\(15\),pp\. 7934\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p2.1)\. - Gonget al\.\(2024\)K\. Gong, X\. Song, W\. Li, and S\. WangHN\-gccf: high\-order neighbor\-enhanced graph convolutional collaborative filtering\.Knowledge\-Based Systems283,pp\. 111122\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p4.1)\. - González\-Castroet al\.\(2024\)A\. González\-Castro, R\. Leirós\-Rodríguez, C\. Prada\-García, and J\. A\. Benítez\-AndradesThe applications of artificial intelligence for assessing fall risk: systematic review\.Journal of medical internet research26,pp\. e54934\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p6.1)\. - Hastieet al\.\(2009\)T\. Hastie, R\. Tibshirani, J\. Friedman,et al\.The elements of statistical learning\.Springer series in statistics New\-York\.Cited by:[§B\.3](https://arxiv.org/html/2608.18450#A2.SS3.p1.1)\. - Heet al\.\(2022\)Y\. He, B\. Brouwers, H\. Liu, H\. Liu, K\. Lawler, E\. Mendes de Oliveira, D\. Lee, Y\. Yang, A\. R\. Cox, J\. M\. Keogh,et al\.Human loss\-of\-function variants in the serotonin 2c receptor associated with obesity and maladaptive behavior\.Nature Medicine28\(12\),pp\. 2537–2546\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p3.1)\. - Hougaard \(1995\)P\. HougaardFrailty models for survival data\.Lifetime data analysis1\(3\),pp\. 255–273\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p8.1),[§2](https://arxiv.org/html/2608.18450#S2.p2.1)\. - Howcroftet al\.\(2017\)J\. Howcroft, J\. Kofman, and E\. D\. LemaireProspective fall\-risk prediction models for older adults based on wearable sensors\.IEEE transactions on neural systems and rehabilitation engineering25\(10\),pp\. 1812–1820\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p6.1)\. - Kalbfleisch and Schaubel \(2023\)J\. D\. Kalbfleisch and D\. E\. SchaubelFifty years of the cox model\.Annual Review of Statistics and Its Application10\(1\),pp\. 1–23\.Cited by:[§2](https://arxiv.org/html/2608.18450#S2.p2.1)\. - Khuranaet al\.\(2018\)U\. Khurana, H\. Samulowitz, and D\. TuragaFeature engineering for predictive modeling using reinforcement learning\.InProceedings of the AAAI conference on artificial intelligence,Vol\.32\.Cited by:[Table 5](https://arxiv.org/html/2608.18450#A2.T5.2.1.8.1),[§D\.1](https://arxiv.org/html/2608.18450#A4.SS1.p1.1),[Table 8](https://arxiv.org/html/2608.18450#A4.T8.2.1.14.2),[Table 8](https://arxiv.org/html/2608.18450#A4.T8.2.1.20.2),[Table 8](https://arxiv.org/html/2608.18450#A4.T8.2.1.8.2),[§2](https://arxiv.org/html/2608.18450#S2.p4.1),[§5\.2](https://arxiv.org/html/2608.18450#S5.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.18450#S5.T1.2.1.5.1.1)\. - Kohler\-Voinovet al\.\(2025\)L\. C\. Kohler\-Voinov, Z\. N\. Sayyid, and K\. E\. CullenMultifactorial predictors of falls in older adults: a decade of data from the national health and aging trends study\.BMC geriatrics25\(1\),pp\. 950\.Cited by:[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px1.p2.1)\. - Komorowskiet al\.\(2018\)M\. Komorowski, L\. A\. Celi, O\. Badawi, A\. C\. Gordon, and A\. A\. FaisalThe artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care\.Nature medicine24\(11\),pp\. 1716–1720\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p5.1),[§2](https://arxiv.org/html/2608.18450#S2.p4.1)\. - Kuncheva \(2007\)L\. I\. KunchevaA stability index for feature selection\.\.InArtificial intelligence and applications,pp\. 421–427\.Cited by:[§5\.2](https://arxiv.org/html/2608.18450#S5.SS2.p1.1)\. - Laiet al\.\(2018\)G\. Lai, W\. Chang, Y\. Yang, and H\. LiuModeling long\-and short\-term temporal patterns with deep neural networks\.InThe 41st international ACM SIGIR conference on research & development in information retrieval,pp\. 95–104\.Cited by:[§D\.2](https://arxiv.org/html/2608.18450#A4.SS2.p1.1)\. - Levyet al\.\(2018\)S\. S\. Levy, K\. J\. Thralls, and S\. A\. KviatkovskyValidity and reliability of a portable balance tracking system, btracks, in older adults\.Journal of geriatric physical therapy41\(2\),pp\. 102–107\.Cited by:[§4\.3](https://arxiv.org/html/2608.18450#S4.SS3.p1.1)\. - Liet al\.\(2022\)F\. Li, W\. Lu, Y\. Wang, Z\. Pan, E\. J\. Greene, G\. Meng, C\. Meng, O\. Blaha, Y\. Zhao, P\. Peduzzi,et al\.A comparison of analytical strategies for cluster randomized trials with survival outcomes in the presence of competing risks\.Statistical Methods in Medical Research31\(7\),pp\. 1224–1241\.Cited by:[§2](https://arxiv.org/html/2608.18450#S2.p2.1)\. - Li and Surineni \(2025\)S\. Li and K\. SurineniFalls in hospitalized patients and preventive strategies: a narrative review\.The American Journal of Geriatric Psychiatry: Open Science, Education, and Practice5,pp\. 1–9\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p7.1)\. - Liet al\.\(2023\)Z\. Li, S\. Qi, Y\. Li, and Z\. XuRevisiting long\-term time series forecasting: an investigation on linear mapping\.arXiv preprint arXiv:2305\.10721\.Cited by:[§D\.1](https://arxiv.org/html/2608.18450#A4.SS1.p1.1),[Table 9](https://arxiv.org/html/2608.18450#A4.T9.2.1.1.5)\. - Limet al\.\(2024\)Z\. K\. Lim, T\. Connie, M\. K\. O\. Goh, and N\. ‘\. B\. SaedonFall risk prediction using temporal gait features and machine learning approaches\.Frontiers in Artificial Intelligence7,pp\. 1425713\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p6.1)\. - Liuet al\.\(2025a\)C\. Liu, L\. Thiamwong, Y\. Fu, and R\. XieDiffusion policies with offline and inverse reinforcement learning for promoting physical activity in older adults using wearable sensors\.In2025 International Conference on Machine Learning and Applications \(ICMLA\),Vol\.,pp\. 275–282\.External Links:[Document](https://dx.doi.org/10.1109/ICMLA66185.2025.00043)Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p6.1),[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px1.p2.1)\. - Liuet al\.\(2022\)M\. Liu, A\. Zeng, M\. Chen, Z\. Xu, Q\. Lai, L\. Ma, and Q\. XuScinet: time series modeling and forecasting with sample convolution and interaction\.Advances in Neural Information Processing Systems35,pp\. 5816–5828\.Cited by:[§D\.1](https://arxiv.org/html/2608.18450#A4.SS1.p1.1),[Table 9](https://arxiv.org/html/2608.18450#A4.T9.2.1.1.8)\. - Liuet al\.\(2025b\)Y\. Liu, C\. Liu, L\. Ni, W\. Zhang, C\. Chen, J\. Lopez, H\. Zheng, L\. Thiamwong, and R\. XieEffectiveness of peer intervention on older adults’ physical activity time series using smoothing spline anova\.Mathematics13\(3\),pp\. 516\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p3.1),[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px1.p2.1)\. - Liuet al\.\(2023\)Y\. Liu, T\. Hu, H\. Zhang, H\. Wu, S\. Wang, L\. Ma, and M\. LongItransformer: inverted transformers are effective for time series forecasting\.arXiv preprint arXiv:2310\.06625\.Cited by:[§C\.2](https://arxiv.org/html/2608.18450#A3.SS2.p1.1),[§D\.1](https://arxiv.org/html/2608.18450#A4.SS1.p1.1),[§D\.2](https://arxiv.org/html/2608.18450#A4.SS2.p2.1),[§D\.3](https://arxiv.org/html/2608.18450#A4.SS3.p1.1),[Table 9](https://arxiv.org/html/2608.18450#A4.T9),[Table 9](https://arxiv.org/html/2608.18450#A4.T9.2.1.1.4)\. - Luoet al\.\(2024\)Y\. Luo, H\. Li, L\. Shi, and X\. WuEnhancing graph transformers with hierarchical distance structural encoding\.Advances in Neural Information Processing Systems37,pp\. 57150–57182\.Cited by:[Table 5](https://arxiv.org/html/2608.18450#A2.T5.2.1.4.1),[§C\.1](https://arxiv.org/html/2608.18450#A3.SS1.SSS0.Px1.p1.1),[§D\.1](https://arxiv.org/html/2608.18450#A4.SS1.p1.1),[Table 8](https://arxiv.org/html/2608.18450#A4.T8.2.1.10.2),[Table 8](https://arxiv.org/html/2608.18450#A4.T8.2.1.16.2),[Table 8](https://arxiv.org/html/2608.18450#A4.T8.2.1.4.2),[§2](https://arxiv.org/html/2608.18450#S2.p3.1)\. - Maćkiewicz and Ratajczak \(1993\)A\. Maćkiewicz and W\. RatajczakPrincipal components analysis \(pca\)\.Computers & Geosciences19\(3\),pp\. 303–342\.Cited by:[§5\.2](https://arxiv.org/html/2608.18450#S5.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.18450#S5.T1.2.1.6.1.1)\. - Maruszewskaet al\.\(2025\)A\. Maruszewska, T\. Ambroży, and Ł\. RydzikRisk factors and socioeconomic determinants of falls among older adults\.Frontiers in Public Health13,pp\. 1571312\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p3.1)\. - McManuset al\.\(2022\)K\. McManus, B\. R\. Greene, L\. G\. M\. Ader, and B\. CaulfieldDevelopment of data\-driven metrics for balance impairment and fall risk assessment in older adults\.IEEE Transactions on Biomedical Engineering69\(7\),pp\. 2324–2332\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p6.1)\. - Meinshausen and Bühlmann \(2006\)N\. Meinshausen and P\. BühlmannHigh\-dimensional graphs and variable selection with the lasso\.Cited by:[§5\.2](https://arxiv.org/html/2608.18450#S5.SS2.p1.1)\. - Meinshausen and Bühlmann \(2010\)N\. Meinshausen and P\. BühlmannStability selection\.Journal of the Royal Statistical Society Series B: Statistical Methodology72\(4\),pp\. 417–473\.Cited by:[§5\.2](https://arxiv.org/html/2608.18450#S5.SS2.p1.1)\. - Mortazaviet al\.\(2023\)S\. Mortazavi, A\. Delbari, M\. Vahedi, R\. Fadayevatan, M\. Moodi, H\. Fakhrzadeh, M\. Khorashadizadeh, A\. Sobhani, M\. Payab, M\. Ebrahimpur,et al\.Low physical activity and depression are the prominent predictive factors for falling in older adults: the birjand longitudinal aging study \(blas\)\.BMC geriatrics23\(1\),pp\. 758\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p2.1)\. - Mukherjeeet al\.\(2023\)A\. Mukherjee, I\. Garg, and K\. RoyEncoding hierarchical information in neural networks helps in subpopulation shift\.IEEE Transactions on Artificial Intelligence5\(2\),pp\. 827–838\.Cited by:[§2](https://arxiv.org/html/2608.18450#S2.p3.1)\. - Nguyenet al\.\(2024\)T\. Nguyen, L\. Thiamwong, Q\. Lou, and R\. XieUnveiling fall triggers in older adults: a machine learning graphical model analysis\.Mathematics12\(9\),pp\. 1271\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p6.1),[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px1.p2.1)\. - Nieet al\.\(2022\)Y\. Nie, N\. H\. Nguyen, P\. Sinthong, and J\. KalagnanamA time series is worth 64 words: long\-term forecasting with transformers\.arXiv preprint arXiv:2211\.14730\.Cited by:[§D\.1](https://arxiv.org/html/2608.18450#A4.SS1.p1.1),[Table 9](https://arxiv.org/html/2608.18450#A4.T9.2.1.1.6)\. - Paliwalet al\.\(2017\)Y\. Paliwal, P\. W\. Slattum, and S\. M\. RatliffChronic health conditions as a risk factor for falls among the community\-dwelling us older adults: a zero\-inflated regression modeling approach\.BioMed research international2017\(1\),pp\. 5146378\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p1.1)\. - Pepe \(2003\)M\. S\. PepeThe statistical evaluation of medical tests for classification and prediction\.Oxford university press\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p8.1)\. - Pfortmuelleret al\.\(2014\)C\. Pfortmueller, G\. Lindner, and A\. ExadaktylosReducing fall risk in the elderly: risk factors and fall prevention, a systematic review\.Minerva Med105\(4\),pp\. 275–81\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p2.1)\. - Picernoet al\.\(2021\)P\. Picerno, M\. Iosa, C\. D’Souza, M\. G\. Benedetti, S\. Paolucci, and G\. MoroneWearable inertial sensors for human movement analysis: a five\-year update\.Expert review of medical devices18\(sup1\),pp\. 79–94\.Cited by:[§2](https://arxiv.org/html/2608.18450#S2.p1.1)\. - Pillayet al\.\(2024\)J\. Pillay, L\. A\. Gaudet, S\. Saba, B\. Vandermeer, A\. R\. Ashiq, A\. Wingert, and L\. HartlingFalls prevention interventions for community\-dwelling older adults: systematic review and meta\-analysis of benefits, harms, and patient values and preferences\.Systematic Reviews13\(1\),pp\. 289\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p3.1)\. - Riesen and Bunke \(2008\)K\. Riesen and H\. BunkeIAM graph database repository for graph based pattern recognition and machine learning\.InJoint IAPR international workshops on statistical techniques in pattern recognition \(SPR\) and structural and syntactic pattern recognition \(SSPR\),pp\. 287–297\.Cited by:[§D\.2](https://arxiv.org/html/2608.18450#A4.SS2.p1.1)\. - Ritcheyet al\.\(2022\)K\. Ritchey, A\. Olney, S\. Chen, and E\. A\. PhelanSTEADI self\-report measures independently predict fall risk\.Gerontology and geriatric medicine8,pp\. 23337214221079222\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p6.1)\. - Rosellini and Brown \(2021\)A\. J\. Rosellini and T\. A\. BrownDeveloping and validating clinical questionnaires\.Annual review of clinical psychology17\(1\),pp\. 55–81\.Cited by:[§2](https://arxiv.org/html/2608.18450#S2.p3.1)\. - Rykovet al\.\(2021\)Y\. Rykov, T\. Thach, I\. Bojic, G\. Christopoulos, and J\. CarDigital biomarkers for depression screening with wearable devices: cross\-sectional study with machine learning modeling\.JMIR mHealth and uHealth9\(10\),pp\. e24872\.Cited by:[§2](https://arxiv.org/html/2608.18450#S2.p1.1)\. - Salop and Salop \(1976\)J\. Salop and S\. SalopSelf\-selection and turnover in the labor market\.The Quarterly Journal of Economics90\(4\),pp\. 619–627\.Cited by:[§B\.3](https://arxiv.org/html/2608.18450#A2.SS3.p1.1)\. - Saunderset al\.\(2025\)S\. Saunders, C\. D’Amore, Q\. Hao, N\. Abd El\-Moneim, J\. Richardson, A\. Kuspinar, and M\. BeauchampRisk factors for falls in community\-dwelling older adults: an umbrella review\.Journal of the American Medical Directors Association26\(9\),pp\. 105765\.Cited by:[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px1.p2.1)\. - Schomburget al\.\(2004\)I\. Schomburg, A\. Chang, C\. Ebeling, M\. Gremse, C\. Heldt, G\. Huhn, and D\. SchomburgBRENDA, the enzyme database: updates and major new developments\.Nucleic acids research32\(suppl\_1\),pp\. D431–D433\.Cited by:[§D\.2](https://arxiv.org/html/2608.18450#A4.SS2.p1.1)\. - Smithet al\.\(2022\)H\. Smith, M\. Sweeting, T\. Morris, and M\. J\. CrowtherA scoping methodological review of simulation studies comparing statistical and machine learning approaches to risk prediction for time\-to\-event data\.Diagnostic and Prognostic Research6\(1\),pp\. 10\.Cited by:[§2](https://arxiv.org/html/2608.18450#S2.p2.1)\. - Sturniekset al\.\(2025\)D\. L\. Sturnieks, L\. Lloyd, M\. T\. E\. Cerda, C\. H\. Arbona, B\. H\. Pinilla, P\. S\. Martinez, N\. W\. Seng, N\. Smith, J\. C\. Menant, and R\. L\. StephenCognitive functioning and falls in older people: a systematic review and meta\-analysis\.Archives of Gerontology and Geriatrics128,pp\. 105638\.Cited by:[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px1.p2.1),[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px2.p1.1)\. - Sureshet al\.\(2022\)K\. Suresh, C\. Severn, and D\. GhoshSurvival prediction models: an introduction to discrete\-time modeling\.BMC medical research methodology22\(1\),pp\. 207\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p8.1),[§2](https://arxiv.org/html/2608.18450#S2.p2.1)\. - Taheriet al\.\(2025\)N\. Taheri, L\. Becker, L\. Fleig, K\. Kolodziejczak, L\. Cordes, B\. U\. Hoehl, U\. Grittner, L\. Mödl, H\. Schmidt, and M\. PumbergerFear\-avoidance beliefs are associated with changes of back shape and function\.Pain Reports10\(2\),pp\. e1249\.Cited by:[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px2.p1.1),[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px3.p1.1)\. - Theng and Bhoyar \(2024\)D\. Theng and K\. K\. BhoyarFeature selection techniques for machine learning: a survey of more than two decades of research\.Knowledge and Information Systems66\(3\),pp\. 1575–1637\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p4.1)\. - Thiamwonget al\.\(2023a\)L\. Thiamwong, R\. Xie, J\. H\. Park, N\. Lighthall, V\. Loerzel, and J\. StoutOptimizing a technology\-based body and mind intervention to prevent falls and reduce health disparities in low\-income populations: protocol for a clustered randomized controlled trial\.JMIR research protocols12,pp\. e51899\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p4.1),[§1](https://arxiv.org/html/2608.18450#S1.p7.1),[§3](https://arxiv.org/html/2608.18450#S3.p1.1),[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px1.p2.1),[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px2.p1.1),[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px3.p1.1)\. - Thiamwonget al\.\(2020a\)L\. Thiamwong, M\. L\. Sole, B\. P\. Ng, G\. F\. Welch, H\. J\. Huang, and J\. R\. StoutAssessing fall risk appraisal through combined physiological and perceived fall risk measures using innovative technology\.Journal of gerontological nursing46\(4\),pp\. 41–47\.Cited by:[§4\.3](https://arxiv.org/html/2608.18450#S4.SS3.p1.1)\. - Thiamwonget al\.\(2020b\)L\. Thiamwong, J\. R\. Stout, M\. L\. Sole, B\. P\. Ng, X\. Yan, and S\. TalbertPhysio\-feedback and exercise program \(peer\) improves balance, muscle strength, and fall risk in older adults\.Research in gerontological nursing13\(6\),pp\. 289–296\.Cited by:[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px1.p2.1)\. - Thiamwonget al\.\(2023b\)L\. Thiamwong, R\. Xie, N\. E\. Conner, J\. M\. Renziehausen, E\. O\. Ojo, and J\. R\. StoutBody composition, fear of falling and balance performance in community\-dwelling older adults\.Translational medicine of aging7,pp\. 80–86\.Cited by:[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px1.p2.1),[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px2.p1.1),[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px3.p1.1)\. - Tibshirani \(1996\)R\. TibshiraniRegression shrinkage and selection via the lasso\.Journal of the Royal Statistical Society Series B: Statistical Methodology58\(1\),pp\. 267–288\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p8.1),[§2](https://arxiv.org/html/2608.18450#S2.p2.1),[§5\.2](https://arxiv.org/html/2608.18450#S5.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.18450#S5.T1.2.1.8.1.1)\. - Tonchoyet al\.\(2024\)P\. Tonchoy, K\. Seangpraw, P\. Ong\-Artborirak, S\. Kantow, N\. Auttama, M\. Choowanthanapakorn, and S\. BoonyatheeMental health, fall prevention behaviors, and home environments related to fall experiences among older adults from ethnic groups in rural northern thailand\.Heliyon10\(17\)\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p7.1)\. - Uddinet al\.\(2021\)M\. P\. Uddin, M\. A\. Mamun, and M\. A\. HossainPCA\-based feature reduction for hyperspectral remote sensing image classification\.IETE Technical Review38\(4\),pp\. 377–396\.Cited by:[Table 5](https://arxiv.org/html/2608.18450#A2.T5.2.1.7.1),[§D\.1](https://arxiv.org/html/2608.18450#A4.SS1.p1.1),[Table 8](https://arxiv.org/html/2608.18450#A4.T8.2.1.13.2),[Table 8](https://arxiv.org/html/2608.18450#A4.T8.2.1.19.2),[Table 8](https://arxiv.org/html/2608.18450#A4.T8.2.1.7.2)\. - Wang and Zhong \(2025\)J\. Wang and Q\. ZhongJoint modeling of longitudinal and survival data\.Annual Review of Statistics and Its Application12\(1\),pp\. 449–476\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p8.1),[§2](https://arxiv.org/html/2608.18450#S2.p2.1)\. - Wanget al\.\(2020\)J\. Wang, Y\. Liu, and B\. LiReinforcement learning with perturbed rewards\.InProceedings of the AAAI conference on artificial intelligence,Vol\.34,pp\. 6202–6209\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p7.1)\. - Wanget al\.\(2024\)J\. Wang, Y\. Li, G\. Yang, and K\. JinAge\-related dysfunction in balance: a comprehensive review of causes, consequences, and interventions\.Aging and disease16\(2\),pp\. 714\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p3.1)\. - Wanget al\.\(2023\)W\. Wang, D\. Han, X\. Luo, and D\. LiAddressing signal delay in deep reinforcement learning\.InThe Twelfth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p7.1)\. - Wuet al\.\(2022\)H\. Wu, T\. Hu, Y\. Liu, H\. Zhou, J\. Wang, and M\. LongTimesnet: temporal 2d\-variation modeling for general time series analysis\.arXiv preprint arXiv:2210\.02186\.Cited by:[§D\.1](https://arxiv.org/html/2608.18450#A4.SS1.p1.1),[Table 9](https://arxiv.org/html/2608.18450#A4.T9.2.1.1.7)\. - Wuet al\.\(2021\)H\. Wu, J\. Xu, J\. Wang, and M\. LongAutoformer: decomposition transformers with auto\-correlation for long\-term series forecasting\.Advances in neural information processing systems34,pp\. 22419–22430\.Cited by:[§D\.2](https://arxiv.org/html/2608.18450#A4.SS2.p1.1)\. - Xuet al\.\(2019\)C\. Xu, P\. R\. Ebeling, and D\. ScottBody composition and falls risk in older adults\.Current Geriatrics Reports8\(3\),pp\. 210–222\.Cited by:[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px2.p1.1),[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px3.p1.1)\. - Xuet al\.\(2015\)T\. Xu, J\. Sun, and J\. BiLongitudinal lasso: jointly learning features and temporal contingency for outcome prediction\.InProceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,pp\. 1345–1354\.Cited by:[§2](https://arxiv.org/html/2608.18450#S2.p2.1)\. - Yardleyet al\.\(2005\)L\. Yardley, N\. Beyer, K\. Hauer, G\. Kempen, C\. Piot\-Ziegler, and C\. ToddDevelopment and initial validation of the falls efficacy scale\-international \(fes\-i\)\.Age and ageing34\(6\),pp\. 614–619\.Cited by:[§4\.3](https://arxiv.org/html/2608.18450#S4.SS3.p1.1)\. - Yassineet al\.\(2021\)A\. Yassine, L\. Mohamed, and M\. Al AchhabIntelligent recommender system based on unsupervised machine learning and demographic attributes\.Simulation Modelling Practice and Theory107,pp\. 102198\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p4.1)\. - Yiet al\.\(2022\)D\. Yi, S\. Jang, and J\. YimRelationship between associated neuropsychological factors and fall risk factors in community\-dwelling elderly\.InHealthcare,Vol\.10,pp\. 728\.Cited by:[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px1.p2.1),[§5\.1](https://arxiv.org/html/2608.18450#S5.SS1.SSS0.Px3.p1.1)\. - Yinget al\.\(2024\)W\. Ying, H\. Bai, K\. Liu, and Y\. FuTopology\-aware reinforcement feature space reconstruction for graph data\.ACM Transactions on Knowledge Discovery from Data\.Cited by:[Table 5](https://arxiv.org/html/2608.18450#A2.T5.2.1.5.1),[§C\.3](https://arxiv.org/html/2608.18450#A3.SS3.p1.1),[§D\.1](https://arxiv.org/html/2608.18450#A4.SS1.p1.1),[§D\.2](https://arxiv.org/html/2608.18450#A4.SS2.p2.1),[§D\.3](https://arxiv.org/html/2608.18450#A4.SS3.p1.1),[Table 8](https://arxiv.org/html/2608.18450#A4.T8),[Table 8](https://arxiv.org/html/2608.18450#A4.T8.2.1.11.2),[Table 8](https://arxiv.org/html/2608.18450#A4.T8.2.1.17.2),[Table 8](https://arxiv.org/html/2608.18450#A4.T8.2.1.5.2),[§2](https://arxiv.org/html/2608.18450#S2.p4.1),[§4\.2](https://arxiv.org/html/2608.18450#S4.SS2.p1.1),[§5\.2](https://arxiv.org/html/2608.18450#S5.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.18450#S5.T1.2.1.3.1.1)\. - Yuet al\.\(2021\)C\. Yu, J\. Liu, S\. Nemati, and G\. YinReinforcement learning in healthcare: a survey\.ACM Computing Surveys \(CSUR\)55\(1\),pp\. 1–36\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p5.1),[§2](https://arxiv.org/html/2608.18450#S2.p4.1)\. - Yuan and Lin \(2006\)M\. Yuan and Y\. LinModel selection and estimation in regression with grouped variables\.Journal of the Royal Statistical Society Series B: Statistical Methodology68\(1\),pp\. 49–67\.Cited by:[§1](https://arxiv.org/html/2608.18450#S1.p8.1),[§2](https://arxiv.org/html/2608.18450#S2.p2.1),[§5\.2](https://arxiv.org/html/2608.18450#S5.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.18450#S5.T1.2.1.7.1.1)\. - Zhonget al\.\(2012\)W\. Zhong, T\. Zhang, Y\. Zhu, and J\. S\. LiuCorrelation pursuit: forward stepwise variable selection for index models\.Journal of the Royal Statistical Society Series B: Statistical Methodology74\(5\),pp\. 849–870\.Cited by:[§4\.2](https://arxiv.org/html/2608.18450#S4.SS2.p1.1),[§4\.3](https://arxiv.org/html/2608.18450#S4.SS3.p2.2)\. - Zhouet al\.\(2021\)H\. Zhou, S\. Zhang, J\. Peng, S\. Zhang, J\. Li, H\. Xiong, and W\. ZhangInformer: beyond efficient transformer for long sequence time\-series forecasting\.InProceedings of the AAAI conference on artificial intelligence,Vol\.35,pp\. 11106–11115\.Cited by:[§D\.2](https://arxiv.org/html/2608.18450#A4.SS2.p1.1)\. ## Appendix ASupplement Overview This supplementary document presents additional information to support the main manuscript\. It includes detailed descriptions of the datasets used, the full configuration of the PAFIR model, the training and evaluation procedures, and extended experimental results\. These materials aim to enhance the reproducibility and transparency of our study and provide deeper insights into the implementation and performance of PAFIR across different settings\. The code implementing the proposed PAFIR framework and all experimental pipelines is available at:[https://github\.com/changliu1993\-cl/PAFIR](https://github.com/changliu1993-cl/PAFIR) ## Appendix BIn\-Field Evaluation: PEER Study on Fall Prevention in Older Adults ### B\.1PEER Intervention and Control Across Visits The top 10 most important fall risk factor groups and their evolving interconnections across four visits are visualized for both the PEER Intervention and Control in Fig\.[6](https://arxiv.org/html/2608.18450#A2.F6)\. In the PEER cluster \(Fig\.[6](https://arxiv.org/html/2608.18450#A2.F6)[6](https://arxiv.org/html/2608.18450#A2.F6)\), which received peer\-led interventions, we observe greater temporal continuity in key domains such as cognitive function and physical activity\. Fall risk factor groups likeMCST,Mindful Attention Awareness Scale, andbalance confidenceappear repeatedly across multiple visits, suggesting a consistent focus on cognitive engagement and behavioral monitoring throughout the intervention period\. In contrast, the Control cluster shows more fragmented transitions, with higher variability in feature composition across visits in Fig\.[6](https://arxiv.org/html/2608.18450#A2.F6)[6](https://arxiv.org/html/2608.18450#A2.F6)\. Although core factors such asphysical activityandbody compositionpersist, other fall risk groups, particularly those related to cognition and behavior, tend to fluctuate more\. This divergence highlights how the presence of a structured, peer\-led program may help stabilize attention toward critical fall risk factors over time\. Despite these differences, both clusters consistently highlight domains likephysical activityandbody composition, reinforcing their foundational importance in fall risk factors identification\. However, the PEER cluster captures a broader range of psychological and cognitive factors, which may support more holistic and individualized monitoring\. These patterns underscore the role of community\-based interventions in fostering consistent and interpretable health data trajectories across longitudinal assessments\. \\subfigure\[PEER Intervention\]\\subfigure\[Control\] Figure 6:Top 10 Fall Risk Factor Identification across 4 visits\. \(a\) PEER Intervention, \(b\) Control\. ### B\.2Initial Reference Sets for Evaluation To evaluate feature\- and group\-level selection performance, we use visit\-specific initial reference sets defined from the graph\- and node\-embedding results\. Letℛτfeat\\mathcal\{R\}\_\{\\tau\}^\{\\mathrm\{feat\}\}andℛτgrp\\mathcal\{R\}\_\{\\tau\}^\{\\mathrm\{grp\}\}denote the feature\- and group\-level reference sets at visitτ\\tau, respectively\. The reference sets are embedding\-derived evaluation proxies constructed using the training partition only and independently of the evaluated methods\. They provide a common reference for method comparison and should not be interpreted as exhaustive clinical ground truth\. Table 3:Initial feature\- and group\-level reference sets used for evaluation on the PEER study\.Reference GroupReference FeaturesVisitsDemographic Surveyage, education, gender, health, living, number\_of\_falls, race, sleeping\_pills, smokingT1–T4Fall Risk Screening: STEADIste8T1–T4Depression Screening: PHQ\-9phqT1–T4Anxiety Screening: GAI\-SFgai\_sf\_scoreT1, T4Fear of Falling: Short FES\-IfesT1–T4TUG, STS, and BTrackSbbs, tugT1–T4Short Physical Performance Batterybalance, gait, speed\_gait, sppb, sts\_sppb\_3T1–T4Hand Grip Strengthavg\_hgs\_kg\_both\_handsT1–T3Brief Aging Perceptions Questionnaireb\_apq\_chronic, b\_apq\_conseqeunce\_positive, b\_apq\_control\_negative, b\_apq\_control\_positive, b\_apq\_scoreT2–T4Behavioral Inhibition/Activation Systembas\_drive, bas\_fun, bas\_reward, bis\_scoreT4Regulatory Focus Questionnairerfq\_prevention, rfq\_promotionT1, T4FRAIL QuestionnairefraT4Memory Impairment ScreenmisT4RUDASrudas\_scoreT4 ### B\.3Longitudinal Stability in the No\-Fall Subgroup To evaluate the longitudinal stability of PAFIR, we compare the selected feature groups across adjacent visits among participants with no recorded falls during follow\-up\. For each participant, Jaccard similarity\([25](https://arxiv.org/html/2608.18450#bib.bib68)\)is calculated as the number of feature groups selected at both visits divided by the total number of distinct feature groups selected across the two visits\. Selection turnover\([63](https://arxiv.org/html/2608.18450#bib.bib69);[2](https://arxiv.org/html/2608.18450#bib.bib70)\)is defined as one minus the Jaccard similarity, with lower values indicating fewer changes between visits\. For each random seed, Jaccard similarity and turnover are calculated for all eligible participants with valid selections at both visits and then averaged within each adjacent visit pair\. The overall result is obtained by first averaging across the available adjacent visit pairs for each participant and then averaging across eligible participants\. Table[4](https://arxiv.org/html/2608.18450#A2.T4)reports the mean and standard deviation of these results across 20 random seeds\. Visit pairs with missing selections are excluded, andNNdenotes the number of eligible no\-fall participants included in each comparison\. Table 4:Longitudinal stability of PAFIR group\- and feature\-level selections among participants with no recorded falls\. Results are reported as mean±\\pmstandard deviation over 20 random seeds\.Group\-Level SelectionFeature\-Level SelectionVisit PairNNJaccard↑\\uparrowTurnover↓\\downarrowJaccard↑\\uparrowTurnover↓\\downarrowT1–T22530\.744±0\.0070\.744\\pm 0\.0070\.256±0\.0070\.256\\pm 0\.0070\.095±0\.0070\.095\\pm 0\.0070\.905±0\.0070\.905\\pm 0\.007T2–T32270\.956±0\.0140\.956\\pm 0\.0140\.044±0\.0140\.044\\pm 0\.0140\.088±0\.0060\.088\\pm 0\.0060\.912±0\.0060\.912\\pm 0\.006T3–T41580\.969±0\.0150\.969\\pm 0\.0150\.031±0\.0150\.031\\pm 0\.0150\.064±0\.0070\.064\\pm 0\.0070\.936±0\.0070\.936\\pm 0\.007Overall2530\.890±0\.0090\.890\\pm 0\.0090\.110±0\.0090\.110\\pm 0\.0090\.083±0\.0040\.083\\pm 0\.0040\.917±0\.0040\.917\\pm 0\.004 Among the 341 participants in the analyzed cohort, 48 had at least one recorded fall and 293 had no recorded falls during follow\-up\. Among the no\-fall participants, 253, 227, and 158 had valid PAFIR feature\-group selections at both visits for T1–T2, T2–T3, and T3–T4, respectively\. Overall, 253 participants contributed to at least one adjacent\-visit comparison\. As shown in Table[4](https://arxiv.org/html/2608.18450#A2.T4), PAFIR exhibits distinct stability patterns at the group and feature levels\. The overall group\-level Jaccard similarity is0\.890±0\.0090\.890\\pm 0\.009, with a corresponding turnover of0\.110±0\.0090\.110\\pm 0\.009, indicating that the broader selected risk domains remain highly consistent across adjacent visits\. In contrast, the overall feature\-level Jaccard similarity is0\.083±0\.0040\.083\\pm 0\.004, corresponding to a turnover of0\.917±0\.0040\.917\\pm 0\.004\. These results indicate that PAFIR maintains stable group\-level selections while adaptively updating the specific features selected within those groups across visits\. ### B\.4Feature\- and Group\-Level Benchmark and Ablation Results Table 5:Feature\- and group\-level fall\-risk factor selection results on the PEER study\. Precision, recall, and F1 score are computed relative to the predefined feature\- and group\-level reference sets\.MethodFeature\-Level SelectionGroup\-Level SelectionPrecision↑\\uparrowRecall↑\\uparrowF1 Score↑\\uparrowPrecision↑\\uparrowRecall↑\\uparrowF1 Score↑\\uparrowPAFIR \(Ours\)0\.887±\\pm0\.0080\.874±\\pm0\.0070\.880±\\pm0\.0070\.781±\\pm0\.0110\.793±\\pm0\.0090\.787±\\pm0\.010GraphGPS\+HDSE \([44](https://arxiv.org/html/2608.18450#bib.bib14)\)0\.891±\\pm0\.0090\.874±\\pm0\.0080\.882±\\pm0\.0080\.736±\\pm0\.0090\.744±\\pm0\.0100\.740±\\pm0\.009TAR \([89](https://arxiv.org/html/2608.18450#bib.bib8)\)0\.788±\\pm0\.0100\.795±\\pm0\.0090\.791±\\pm0\.0100\.767±\\pm0\.0080\.791±\\pm0\.0080\.779±\\pm0\.008SAT \([6](https://arxiv.org/html/2608.18450#bib.bib7)\)0\.723±\\pm0\.0090\.731±\\pm0\.0100\.727±\\pm0\.0090\.677±\\pm0\.0130\.689±\\pm0\.0110\.683±\\pm0\.012PCA \([77](https://arxiv.org/html/2608.18450#bib.bib5)\)0\.643±\\pm0\.0130\.624±\\pm0\.0150\.638±\\pm0\.0140\.573±\\pm0\.0130\.566±\\pm0\.0120\.569±\\pm0\.011TTG \([30](https://arxiv.org/html/2608.18450#bib.bib6)\)0\.677±\\pm0\.0080\.664±\\pm0\.0090\.670±\\pm0\.0080\.643±\\pm0\.0120\.621±\\pm0\.0120\.611±\\pm0\.012 Table[5](https://arxiv.org/html/2608.18450#A2.T5)reports feature\- and group\-level selection performance on the PEER study\. PAFIR achieves the highest group\-level precision, recall, and F1 score, indicating that the hierarchical selection policy effectively identifies clinically relevant feature groups\. At the individual\-feature level, GraphGPS\+HDSE obtains slightly higher precision and F1 score, while PAFIR achieves comparable recall\. The ablation results show that removing hierarchical graph encoding, temporal encoding, or the three\-agent comparison\-driven selection mechanism reduces feature\-level performance\. Removing the three\-agent mechanism also decreases group\-level F1 from 0\.787 to 0\.716, supporting the contribution of the candidate, reference, and operation agents to group\-level feature selection\. ### B\.5Computational Efficiency Table 6:Computational cost of PAFIR on the PEER study\.MeasurementResultGPUNVIDIA Quadro RTX 6000GPU memory capacity24 GBCUDA version13\.0Runtime per random seed∼\\sim101 sRuntime for 20 seeds∼\\sim33\.7 minPeak allocated GPU memory∼\\sim240 MB ## Appendix CPreprocessing ### C\.1Hierarchical State Representation To effectively model both the structural health assessment and temporal dynamics in longitudinal health data, we propose a state representation𝐬i,t\\mathbf\{s\}\_\{i,t\}that integrates graph\-based and time\-series modalities via a cross\-attention mechanism\. This architecture captures health status evolution across visits through three key components that jointly encode structurally informed features, such as fear of falling, at each visit, and temporally aligned physical activity patterns, such as step counts and vector magnitudes\. #### Hierarchical Structural Encoding First, structurally informed features are derived by applying the hierarchical distance structure encoding \(HDSE\)\([44](https://arxiv.org/html/2608.18450#bib.bib14)\), which reveals and encodes latent multi\-level feature relationships among features and their association with the primary outcome,fall incidence\. At each visitτ\\tau, each node represents a feature \(e\.g\.,BMI,weight, andheight\), and each feature group \(e\.g\.,body composition\) forms a fully connected graph𝒢m=\(𝒱m,ℰm\)\\mathcal\{G\}\_\{m\}=\(\\mathcal\{V\}\_\{m\},\\mathcal\{E\}\_\{m\}\), where𝒱m\\mathcal\{V\}\_\{m\}is the set of features belonging to groupmm\. We construct a multi\-level hierarchy for each feature group by iteratively coarsening the graph to generateK\+1K\+1levels\. At each levelkk, we compute the graph hierarchy distance \(GHD\)\([44](https://arxiv.org/html/2608.18450#bib.bib14)\)between nodesv∈𝒱mv\\in\\mathcal\{V\}\_\{m\}andu∈𝒱mu\\in\\mathcal\{V\}\_\{m\}, denoted as, GHDk\(v,u\)=SPD\(ψk−1∘⋯∘ψ0\(v\),ψk−1∘⋯∘ψ0\(u\)\),\\displaystyle\\text\{GHD\}^\{k\}\(v,u\)=\\text\{SPD\}\\\!\\left\(\\psi\_\{k\-1\}\\circ\\cdots\\circ\\psi\_\{0\}\(v\),\\psi\_\{k\-1\}\\circ\\cdots\\circ\\psi\_\{0\}\(u\)\\right\),whereψk−1∘⋯∘ψ0\(v\)\\psi\_\{k\-1\}\\circ\\cdots\\circ\\psi\_\{0\}\(v\)is the mapping of nodevvfrom level00to levelkkin the graph hierarchy, andSPD\(⋅,⋅\)\\text\{SPD\}\(\\cdot,\\cdot\)denotes shortest path distance\. For each feature nodevv, we construct a structural distance tensor𝐃v∈ℝ\(K\+1\)×\|𝒱m\|\\mathbf\{D\}\_\{v\}\\in\\mathbb\{R\}^\{\(K\+1\)\\times\|\\mathcal\{V\}\_\{m\}\|\}by stacking its pairwise distances to all other nodesu∈𝒱mu\\in\\mathcal\{V\}\_\{m\}across all hierarchy levels\. Each entry𝐃v,u\\mathbf\{D\}\_\{v,u\}captures the multi\-level distances from nodevvto nodeuu, 𝐃v,u=\[GHD0\(v,u\),GHD1\(v,u\),⋯,GHDK\(v,u\)\]\.\\mathbf\{D\}\_\{v,u\}=\\left\[\\text\{GHD\}^\{0\}\(v,u\),\\text\{GHD\}^\{1\}\(v,u\),\\cdots,\\text\{GHD\}^\{K\}\(v,u\)\\right\]\.\(4\)This tensor captures the multi\-level structural relationships of nodevvwithin groupmm\. To obtain a compact representation, we apply positional encoding to𝐃v\\mathbf\{D\}\_\{v\}, yielding a dense embedding that encodes the hierarchical position of the feature\. ### C\.2Temporal Evolving for Sequence and Time Series To capture the physical activity dynamics observed after each clinical visit, we adopt a temporal state encoder based on the iTransformer architecture\([43](https://arxiv.org/html/2608.18450#bib.bib13)\)\. For each feature, such as step counts and VMs, we extract a 7\-day time series collected starting from the visit day, and treat it as an input token\. Let𝐗i,:,b\(τ\)∈ℝL×1\\mathbf\{X\}\_\{i,:,b\}^\{\(\\tau\)\}\\in\\mathbb\{R\}^\{L\\times 1\}denote the time series of featurebbduring the 7\-day period following visitτ\\taufor participantii, whereLLis the sequence length \(e\.g\., minute\-level for 7 days\)\. Each sequence is projected into a latent space via a shared embedding function, 𝐳i,t,b\(0\)=Embed\(𝐗i,:,b\(τ\)\)∈ℝd,\\mathbf\{z\}\_\{i,t,b\}^\{\(0\)\}=\\text\{Embed\}\\\!\\left\(\\mathbf\{X\}\_\{i,:,b\}^\{\(\\tau\)\}\\right\)\\in\\mathbb\{R\}^\{d\},\(5\)whereddmeans the embedding dimension, which controls the capacity of the latent representation\. We then apply a stack ofMtempM\_\{\\mathrm\{temp\}\}Transformer encoder blocks to capture temporal dependencies, 𝐳i,t,b\(ℓ\+1\)=TrmBlock\(ℓ\)\(𝐳i,t,b\(ℓ\)\),ℓ=0,…,Mtemp−1\.\\mathbf\{z\}\_\{i,t,b\}^\{\(\\ell\+1\)\}=\\text\{TrmBlock\}^\{\(\\ell\)\}\\left\(\\mathbf\{z\}\_\{i,t,b\}^\{\(\\ell\)\}\\right\),\\qquad\\ell=0,\\ldots,M\_\{\\mathrm\{temp\}\}\-1\.\(6\)The final encoder output𝐳i,t,b\\mathbf\{z\}\_\{i,t,b\}summarizes the temporal dynamics of featurebband serves as the query for the cross\-attention mechanism\. #### Cross\-attention Mechanism In the cross\-attention mechanism, we regard the temporal embeddings of physical activity patterns as queries𝐐i,t∈ℝ\|ℬ\|×d′\\mathbf\{Q\}\_\{i,t\}\\in\\mathbb\{R\}^\{\|\\mathcal\{B\}\|\\times d^\{\\prime\}\}, the structural embeddings of feature groups as keys𝐊i,t∈ℝM×d′\\mathbf\{K\}\_\{i,t\}\\in\\mathbb\{R\}^\{M\\times d^\{\\prime\}\}, and the corresponding within\-group structural embeddings as values𝐕i,t∈ℝM×d′\\mathbf\{V\}\_\{i,t\}\\in\\mathbb\{R\}^\{M\\times d^\{\\prime\}\}, whereℬ\\mathcal\{B\}is the set of temporal features,MMis the number of feature groups, andd′d^\{\\prime\}is the hidden dimension per attention head\. To incorporate structural priors, we introduce a bias matrix𝐁i,tstruct∈ℝ\|ℬ\|×M\\mathbf\{B\}\_\{i,t\}^\{\\mathrm\{struct\}\}\\in\\mathbb\{R\}^\{\|\\mathcal\{B\}\|\\times M\}to encourage each temporal query to attend more strongly to structurally relevant feature groups, where each entry encodes the relative GHD\-based proximity between temporal featurebband feature group𝒱m\\mathcal\{V\}\_\{m\}\. The attention is then computed as CrossAttn\(𝐐i,t,𝐊i,t,𝐕i,t\)=softmax\(𝐐i,t𝐊i,t⊤d′\+𝐁i,tstruct\)𝐕i,t\.\\text\{CrossAttn\}\\left\(\\mathbf\{Q\}\_\{i,t\},\\mathbf\{K\}\_\{i,t\},\\mathbf\{V\}\_\{i,t\}\\right\)=\\text\{softmax\}\\left\(\\frac\{\\mathbf\{Q\}\_\{i,t\}\\mathbf\{K\}\_\{i,t\}^\{\\top\}\}\{\\sqrt\{d^\{\\prime\}\}\}\+\\mathbf\{B\}\_\{i,t\}^\{\\mathrm\{struct\}\}\\right\)\\mathbf\{V\}\_\{i,t\}\.\(7\)This fusion mechanism enables the model to align temporal physical activity patterns with the longitudinal structure of clinical survey data, allowing for more informed and interpretable inter\-feature association\. ### C\.3Three\-Agent RL Feature Selection Framework We leverage a three\-agent decision framework\([89](https://arxiv.org/html/2608.18450#bib.bib8)\)that performs personalized and temporally adaptive feature selection under RL\. At each clinical visit, the state representation𝐬i,t\\mathbf\{s\}\_\{i,t\}incorporates both temporal behavior dynamics from physical activity and structural information from clinical feature topology, and is used to guide the decision\-making in a three\-agent RL structure composed of a candidate feature group agent, a reference feature group agent, and an operation agent\. In the RL framework, time stepttcan correspond to a clinical visitτ\\tauor to the occurrence of a fall incident for participantii\. Let𝒞=\{𝒱1,𝒱2,…,𝒱M\}\\mathcal\{C\}=\\\{\\mathcal\{V\}\_\{1\},\\mathcal\{V\}\_\{2\},\\dots,\\mathcal\{V\}\_\{M\}\\\}denote the collection of clinically defined feature groups\. #### Group\-Level Feature Selection At each iterationtt, the framework operates - •Candidate Feature Group Agent: Selects a feature groupCi,tcand∈𝒞C\_\{i,t\}^\{\\mathrm\{cand\}\}\\in\\mathcal\{C\}based on the current state𝐬i,t\\mathbf\{s\}\_\{i,t\}\. - •Reference Feature Group Agent: Selects a comparison feature groupCi,tref∈𝒞∖\{Ci,tcand\}C\_\{i,t\}^\{\\mathrm\{ref\}\}\\in\\mathcal\{C\}\\setminus\\\{C\_\{i,t\}^\{\\mathrm\{cand\}\}\\\}from the remaining feature groups\. - •Operation Agent: Selects an add, remove, or retain operation based on the relative incremental utility ofCi,tcandC\_\{i,t\}^\{\\mathrm\{cand\}\}andCi,trefC\_\{i,t\}^\{\\mathrm\{ref\}\}under the available outcome signalyi,ty\_\{i,t\}\. The resulting actionai,t=\(Ci,tcand,Ci,tref,oi,t\)a\_\{i,t\}=\(C\_\{i,t\}^\{\\mathrm\{cand\}\},C\_\{i,t\}^\{\\mathrm\{ref\}\},o\_\{i,t\}\)governs the inclusion or exclusion of existing feature groups\. The group with higher relevance to the outcome is retained in the selected feature setFi,tF\_\{i,t\}, and the other is used as the reference in the next iteration\. #### Within\-Group Feature Selection To further refine feature granularity, we extend the three\-agent structure to operate within each selected group𝒱m∈Fi,tgroup\\mathcal\{V\}\_\{m\}\\in F\_\{i,t\}^\{\\mathrm\{group\}\}\. For each group, a second round of selection is performed at the individual feature level\. A localized state representation𝐬i,t\(𝒱m\)\\mathbf\{s\}\_\{i,t\}^\{\(\\mathcal\{V\}\_\{m\}\)\}encodes the temporal and structural context of individual features in𝒱m\\mathcal\{V\}\_\{m\}\. Within\-group agents operate similarly\. TheCandidate Feature Agentselects a featurexi,tcand∈𝒱mx\_\{i,t\}^\{\\mathrm\{cand\}\}\\in\\mathcal\{V\}\_\{m\}\(e\.g\.,BMI\), while theReference Feature Agentchooses a comparison featurexi,tref∈𝒱mx\_\{i,t\}^\{\\mathrm\{ref\}\}\\in\\mathcal\{V\}\_\{m\}\(e\.g\.,fat\-free mass\) from the remaining features in the same group\. Finally, theOperation Agentdecides whether to retainxi,tcandx\_\{i,t\}^\{\\mathrm\{cand\}\}\. The temporal\-difference \(TD\) learning and reward mechanisms follow the same procedure as the group level\. The Q\-networks are randomly initialized, while the HDSE node\-level importance scores are used to initialize the feature\-level rewards\. ## Appendix DBenchmarking and Results In this section, we evaluate our proposed method on ten public datasets and demonstrate state\-of\-the\-art performance\. We investigate the research problems: RQ1: Can our method effectively reconstruct high\-quality feature spaces to improve performance on downstream tasks in both graph and time series domains? RQ2: Does our unified architecture outperform state\-of\-the\-art baselines in terms of both structural and temporal evaluation metrics across diverse datasets? ### D\.1Baselines We compare our method against state\-of\-the\-art baselines\. For graph\-based benchmarks, we include HDSE\([44](https://arxiv.org/html/2608.18450#bib.bib14)\), TAR\([89](https://arxiv.org/html/2608.18450#bib.bib8)\), SAT\([6](https://arxiv.org/html/2608.18450#bib.bib7)\), PCA\([77](https://arxiv.org/html/2608.18450#bib.bib5)\), and TTG\([30](https://arxiv.org/html/2608.18450#bib.bib6)\)\. For time series forecasting, we consider iTransformer\([43](https://arxiv.org/html/2608.18450#bib.bib13)\), RLinear\([38](https://arxiv.org/html/2608.18450#bib.bib12)\), PatchTST\([53](https://arxiv.org/html/2608.18450#bib.bib11)\), TimesNet\([82](https://arxiv.org/html/2608.18450#bib.bib10)\), and SCINet\([41](https://arxiv.org/html/2608.18450#bib.bib9)\)\. ### D\.2Benchmark Dataset Description We evaluate our method on six public benchmark datasets across different domains\. For graph\-based comparisons, we use two bioinformatics datasets: ENZYMES\([65](https://arxiv.org/html/2608.18450#bib.bib15)\)and PROTEINS\([13](https://arxiv.org/html/2608.18450#bib.bib16)\), and a small molecules dataset: AIDS\([59](https://arxiv.org/html/2608.18450#bib.bib48)\)\. For time series evaluation, we use three datasets from the energy and transportation domains: Solar\-Energy\([34](https://arxiv.org/html/2608.18450#bib.bib17)\), Traffic\([83](https://arxiv.org/html/2608.18450#bib.bib18)\), Weather\([83](https://arxiv.org/html/2608.18450#bib.bib18)\), and ETT\([93](https://arxiv.org/html/2608.18450#bib.bib19)\)\(details in Table[7](https://arxiv.org/html/2608.18450#A4.T7)\.\)\. Table 7:Detailed Dataset Description\. Dim denotes the variate number of each dataset\. Dataset Size denotes the total number of time points in \(Train, Validation, Test\), respectively\. Prediction Length denotes the future time points to be predicted in each dataset\. Frequency denotes the sampling interval of timepoints\.Top Side: Time Series Public Datasets\.Bottom Side: Graph Public Datasets\.Time\-Series DatasetDimPrediction LengthDataset SizeFrequencyInformationETTh1, ETTh27\{96, 192, 336, 720\}\(8545, 2881, 2881\)HourlyElectricityETTm1, ETTm27\{96, 192, 336, 720\}\(34465, 11521, 11521\)15minElectricityTraffic862\{96, 192, 336, 720\}\(12185, 1757, 3509\)HourlyTransportationSolar\-Energy137\{96, 192, 336, 720\}\(36601, 5161, 10417\)HourlyEnergyWeather21\{96, 192, 336, 720\}\(36792, 5271, 10540\)10minWeatherGraph DatasetGraphs CountsNodes CountsGraph ClassesNode ClassesLabelsENZYMES6001958063YesPROTEINS11134347123YesAIDS200031385238Yes For the graph prediction, we follow the data processing and train\-validate\-test set split protocol used in TAR[89](https://arxiv.org/html/2608.18450#bib.bib8)\. For the time\-series prediction, we follow the same data processing and train\-validate\-test set split protocol used in iTransformer[43](https://arxiv.org/html/2608.18450#bib.bib13), where the train, validation, and test datasets are strictly divided according to chronological order to make sure there are no data leakage issues\. ### D\.3Environment Setup and Metrics For graph benchmarks, we follow the data\-processing and split protocol of TAR\([89](https://arxiv.org/html/2608.18450#bib.bib8)\)\. For time\-series benchmarks, we use the predefined chronological train/validation/test splits adopted by iTransformer\([43](https://arxiv.org/html/2608.18450#bib.bib13)\)\. To ensure robustness, each experiment is repeated 10 times with different random seeds, and the mean performance is reported\. For graph\-based tasks, we evaluate the quality of the transformed feature space using Precision, Recall, and F1 score\. Higher values indicate better performance\. For reinforcement feature space reconstruction, training was limited to 10 epochs, each with 10 exploration steps\. All agents were implemented using a DQN with two ReLU\-activated linear layers, optimized with Adam \(learning rate 0\.01\), an experience replay memory size of 32, and a batch size of 8\. For time series forecasting tasks, we report prediction accuracy using Mean Squared Error \(MSE\) and Mean Absolute Error \(MAE\)\. Lower values indicate better performance\. For the forecasting settings, we fixed the lookback window length to 96 for the ETT, Solar\-Energy, Traffic, and Weather datasets, while the prediction horizons were set to vary among \{96, 192, 336, 720\}\. #### Model Regularization and Generalization Given the high\-dimensional feature space and sparse fall outcomes, we adopt multiple strategies to mitigate overfitting in the proposed PAFIR framework\. First, strong regularization is applied throughout the model, including dropout \(rate = 0\.5\) and residual connections to stabilize deep representations\. Second, hierarchical and group\-wise structural representations reduce the effective dimensionality by leveraging clinically meaningful feature groupings rather than treating all variables independently\. For temporal modeling, a shared encoder is used across longitudinal visits, preventing visit\-specific overfitting and encouraging the learning of consistent temporal patterns\. Finally, the adaptive policy is optimized using temporally aggregated reward signals rather than directly fitting individual fall events, reducing sensitivity to sparse and delayed outcome labels\. ### D\.4Benchmarks Results To address RQ1, we evaluate the performance of our proposed method, PAFIR, in comparison with several state\-of\-the\-art baselines across six benchmark datasets covering both graph classification and time series forecasting tasks\. On the PROTEINS, ENZYMES, and AIDS datasets, as shown in Table[8](https://arxiv.org/html/2608.18450#A4.T8), PAFIR achieves the highest scores across all evaluation metrics, including Precision, Recall, and F1 Score\. These results confirm the effectiveness of PAFIR in capturing hierarchical structure patterns within graph data, consistently outperforming baselines such as GraphGPS\+HDSE, TAR, SAT, PCA, and TTG\. Table 8:Overall Performance Comparison on Hierarchical Graph Datasets\. PAFIR outperforms existing methods on both node and graph classification tasks across PROTEINS, ENZYMES, and AIDS datasets, following the TAR\([89](https://arxiv.org/html/2608.18450#bib.bib8)\)setting\.DatasetMethodNode ClassificationGraph ClassificationPrecisionRecallF1 ScorePrecisionRecallF1 ScorePROTEINSPAFIR0\.921±\\pm0\.0070\.914±\\pm0\.0070\.918±\\pm0\.0090\.843±\\pm0\.0060\.822±\\pm0\.0070\.832±\\pm0\.006GraphGPS\+HDSE \([44](https://arxiv.org/html/2608.18450#bib.bib14)\)0\.867±\\pm0\.0550\.785±\\pm0\.0970\.780±\\pm0\.1800\.830±\\pm0\.0270\.812±\\pm0\.0550\.821±\\pm0\.012TAR[89](https://arxiv.org/html/2608.18450#bib.bib8)0\.916±\\pm0\.0170\.925±\\pm0\.0170\.916±\\pm0\.0170\.767±\\pm0\.0060\.766±\\pm0\.0070\.765±\\pm0\.006SAT \([6](https://arxiv.org/html/2608.18450#bib.bib7)\)0\.779±\\pm0\.1040\.785±\\pm0\.0970\.780±\\pm0\.0370\.769±\\pm0\.0370\.643±\\pm0\.0390\.701±\\pm0\.104PCA[77](https://arxiv.org/html/2608.18450#bib.bib5)0\.781±\\pm0\.0030\.719±\\pm0\.0030\.738±\\pm0\.0030\.643±\\pm0\.0010\.648±\\pm0\.0020\.644±\\pm0\.001TTG \([30](https://arxiv.org/html/2608.18450#bib.bib6)\)0\.847±\\pm0\.0050\.856±\\pm0\.0050\.847±\\pm0\.0050\.752±\\pm0\.0040\.751±\\pm0\.0040\.751±\\pm0\.004ENZYMESPAFIR0\.937±\\pm0\.0040\.940±\\pm0\.0070\.941±\\pm0\.0060\.866±\\pm0\.0060\.848±\\pm0\.0070\.857±\\pm0\.006GraphGPS\+HDSE \([44](https://arxiv.org/html/2608.18450#bib.bib14)\)0\.816±\\pm0\.0060\.862±\\pm0\.2770\.885±\\pm0\.1210\.842±\\pm0\.0030\.819±\\pm0\.0040\.831±\\pm0\.106TAR[89](https://arxiv.org/html/2608.18450#bib.bib8)0\.936±\\pm0\.0060\.934±\\pm0\.0050\.937±\\pm0\.0060\.324±\\pm0\.0530\.358±\\pm0\.0290\.325±\\pm0\.029SAT \([6](https://arxiv.org/html/2608.18450#bib.bib7)\)0\.744±\\pm0\.0070\.782±\\pm0\.0060\.753±\\pm0\.0040\.694±\\pm0\.0080\.833±\\pm0\.0370\.757±\\pm0\.104PCA \([77](https://arxiv.org/html/2608.18450#bib.bib5)\)0\.753±\\pm0\.0000\.768±\\pm0\.0000\.756±\\pm0\.0000\.239±\\pm0\.0000\.292±\\pm0\.0000\.260±\\pm0\.000TTG \([30](https://arxiv.org/html/2608.18450#bib.bib6)\)0\.919±\\pm0\.0030\.928±\\pm0\.0020\.920±\\pm0\.0030\.257±\\pm0\.0420\.308±\\pm0\.0150\.261±\\pm0\.026AIDSPAFIR0\.987±\\pm0\.0030\.991±\\pm0\.0070\.989±\\pm0\.0020\.986±\\pm0\.0010\.985±\\pm0\.0060\.986±\\pm0\.001GraphGPS\+HDSE \([44](https://arxiv.org/html/2608.18450#bib.bib14)\)0\.982±\\pm0\.0020\.989±\\pm0\.0970\.985±\\pm0\.0210\.984±\\pm0\.0060\.985±\\pm0\.0040\.984±\\pm0\.096TAR[89](https://arxiv.org/html/2608.18450#bib.bib8)0\.988±\\pm0\.0020\.991±\\pm0\.0010\.988±\\pm0\.0020\.984±\\pm0\.0020\.984±\\pm0\.0020\.984±\\pm0\.002SAT \([6](https://arxiv.org/html/2608.18450#bib.bib7)\)0\.944±\\pm0\.0070\.982±\\pm0\.0060\.953±\\pm0\.0040\.946±\\pm0\.0050\.933±\\pm0\.0170\.957±\\pm0\.094PCA \([77](https://arxiv.org/html/2608.18450#bib.bib5)\)0\.381±\\pm0\.0000\.615±\\pm0\.0000\.471±\\pm0\.0000\.899±\\pm0\.0000\.899±\\pm0\.0000\.893±\\pm0\.000TTG \([30](https://arxiv.org/html/2608.18450#bib.bib6)\)0\.925±\\pm0\.0230\.947±\\pm0\.0160\.933±\\pm0\.0200\.896±\\pm0\.0000\.899±\\pm0\.0000\.893±\\pm0\.000 For time series forecasting, we conduct experiments on Solar\-Energy, Traffic, ETT \(including ETTh1, ETTh2, ETTm1, and ETTm2\), and Weather datasets with a fixed input sequence length of 96 and prediction lengths varying in \{96, 192, 336, 720\}\. As summarized in Table[9](https://arxiv.org/html/2608.18450#A4.T9), PAFIR achieves the lowest MSE and MAE in most cases across all datasets, outperforming recent baselines including iTransformer, RLinear, PatchTST, TimesNet, and SCINet\. These results demonstrate the model’s strong ability to align temporal resolution with evolving feature dynamics\. Overall, the consistent superiority of PAFIR in both graph and time series domains highlights its robustness and adaptability for complex learning tasks involving heterogeneous and temporal data\. Table 9:Overall Performance for Time Series Forecasting\. We compare extensive competitive methods following the setting of iTransformer\([43](https://arxiv.org/html/2608.18450#bib.bib13)\)\. The input sequence length is set to 96 for all baselines\.DatasetSeq\.PAFIR \(Ours\)iTransformer\([43](https://arxiv.org/html/2608.18450#bib.bib13)\)RLinear\([38](https://arxiv.org/html/2608.18450#bib.bib12)\)PatchTST\([53](https://arxiv.org/html/2608.18450#bib.bib11)\)TimesNet\([82](https://arxiv.org/html/2608.18450#bib.bib10)\)SCINet\([41](https://arxiv.org/html/2608.18450#bib.bib9)\)MSEMAEMSEMAEMSEMAEMSEMAEMSEMAEMSEMAESolarEnergy960\.2130\.236\\mathbf\{0\.236\}0\.203\\mathbf\{0\.203\}0\.2370\.3220\.3390\.2710\.3070\.2500\.2920\.2370\.3441920\.226\\mathbf\{0\.226\}0\.244\\mathbf\{0\.244\}0\.2330\.2610\.3590\.3560\.2670\.3100\.2960\.4160\.2810\.3833360\.243\\mathbf\{0\.243\}0\.258\\mathbf\{0\.258\}0\.2480\.2730\.3970\.3690\.2900\.3150\.3190\.4320\.3020\.4007200\.246\\mathbf\{0\.246\}0\.268\\mathbf\{0\.268\}0\.2500\.2760\.3970\.3560\.2890\.3170\.3380\.4250\.3100\.400Traffic960\.412\\mathbf\{0\.412\}0\.272\\mathbf\{0\.272\}0\.4170\.2760\.6490\.3980\.4620\.3040\.5930\.3210\.8040\.5091920\.424\\mathbf\{0\.424\}0\.269\\mathbf\{0\.269\}0\.4280\.2820\.6010\.3660\.4660\.2960\.6170\.3360\.7890\.5053360\.428\\mathbf\{0\.428\}0\.277\\mathbf\{0\.277\}0\.4330\.2830\.6090\.3690\.4820\.3040\.6290\.3350\.8000\.5087200\.451\\mathbf\{0\.451\}0\.296\\mathbf\{0\.296\}0\.4670\.3010\.6470\.3870\.5140\.3220\.6450\.3510\.8410\.523ETTh1960\.381\\mathbf\{0\.381\}0\.398\\mathbf\{0\.398\}0\.3830\.4050\.3860\.4000\.4140\.4190\.3840\.4020\.6540\.5991920\.434\\mathbf\{0\.434\}0\.422\\mathbf\{0\.422\}0\.4410\.4360\.4370\.4240\.4600\.4450\.4360\.4290\.7190\.6313360\.433\\mathbf\{0\.433\}0\.4500\.4870\.4580\.4790\.446\\mathbf\{0\.446\}0\.5010\.4660\.4910\.4690\.7780\.6597200\.4980\.4870\.5030\.4910\.481\\mathbf\{0\.481\}0\.470\\mathbf\{0\.470\}0\.5000\.4880\.5210\.5000\.8360\.697ETTh2960\.288\\mathbf\{0\.288\}0\.3400\.2970\.3490\.2880\.3380\.3020\.3480\.3400\.3740\.7070\.6211920\.377\{0\.377\}0\.398\{0\.398\}0\.3800\.4000\.3740\.3900\.3880\.4000\.4020\.4140\.8600\.6893360\.4220\.4290\.4280\.4320\.4150\.426\\mathbf\{0\.426\}0\.4260\.4330\.4520\.4521\.0000\.7447200\.4190\.4380\.4270\.4450\.4200\.4400\.4310\.4460\.4620\.4681\.2490\.838ETTm1960\.3300\.366\\mathbf\{0\.366\}0\.3340\.3680\.3550\.3760\.3290\.3670\.3380\.3750\.4180\.4381920\.3700\.385\\mathbf\{0\.385\}0\.3770\.3910\.3910\.3920\.3670\.3850\.3740\.3870\.4390\.4503360\.4080\.4130\.4260\.4200\.4240\.4150\.3990\.4100\.4100\.4110\.4900\.4857200\.4840\.4480\.4910\.4590\.4870\.4500\.4540\.4390\.4780\.4500\.5950\.550ETTm2960\.173\\mathbf\{0\.173\}0\.259\\mathbf\{0\.259\}0\.1800\.2640\.1820\.2650\.1750\.2590\.1870\.2670\.2860\.3771920\.241\\mathbf\{0\.241\}0\.3050\.2500\.3090\.2460\.3040\.2410\.3020\.2490\.3090\.3990\.4453360\.304\\mathbf\{0\.304\}0\.3430\.3110\.3480\.3070\.3420\.3050\.3430\.3210\.3510\.3690\.3427200\.4070\.4030\.4120\.4070\.407\{0\.407\}0\.398\\mathbf\{0\.398\}0\.4020\.4000\.4080\.4030\.9600\.735Weather960\.170\\mathbf\{0\.170\}0\.211\\mathbf\{0\.211\}0\.1740\.2140\.1920\.2320\.1770\.2180\.1720\.2200\.2210\.3061920\.206\\mathbf\{0\.206\}0\.244\\mathbf\{0\.244\}0\.2210\.2540\.2400\.2710\.2250\.2590\.2190\.2610\.2610\.3403360\.277\\mathbf\{0\.277\}0\.2950\.2780\.2960\.2920\.3070\.2780\.2970\.2800\.3060\.3090\.3787200\.3510\.3400\.3580\.3490\.3640\.3530\.3540\.3480\.3650\.3590\.3770\.427 ## Appendix EAblation Study To evaluate the contribution of each key component in our unified architecture and address RQ2, we perform a comprehensive ablation study across both graph\-based and time series forecasting tasks\. Specifically, we examine the individual impact of \(1\) the hierarchical graph encoder, \(2\) the time series encoder, and \(3\) the multi\-agent policy mechanism\. Each component is systematically removed or replaced, and the resulting performance is compared against the full model\. Experiments are conducted on representative benchmarks, including node classification, graph classification, and multivariate time series forecasting, to assess the effectiveness and generalizability of each architectural element\. ### E\.1Impact of Hierarchical Graph Encoding To evaluate the effectiveness of the hierarchical graph encoder in capturing structural dependencies, we perform an ablation by removing this component from the PAFIR architecture\. As shown in Table[10](https://arxiv.org/html/2608.18450#A5.T10), the performance drops significantly across both node and graph classification tasks on the PROTEINS, ENZYMES, and AIDS datasets\. Table 10:Ablation Experiment for Hierarchical Graph Encoding\.DatasetModel VariantNode ClassificationGraph ClassificationPrecisionRecallF1 ScorePrecisionRecallF1 ScorePROTEINSPAFIR \(Full\)0\.921±0\.007\\mathbf\{0\.921\\pm 0\.007\}0\.914±0\.007\\mathbf\{0\.914\\pm 0\.007\}0\.918±0\.009\\mathbf\{0\.918\\pm 0\.009\}0\.843±0\.006\\mathbf\{0\.843\\pm 0\.006\}0\.822±0\.007\\mathbf\{0\.822\\pm 0\.007\}0\.832±0\.006\\mathbf\{0\.832\\pm 0\.006\}w/o HierarchicalGraph Encoding0\.678±\\pm0\.0050\.673±\\pm0\.0130\.675±\\pm0\.0200\.765±\\pm0\.0330\.757±\\pm0\.0050\.766±\\pm0\.003ENZYMESPAFIR \(Full\)0\.937±0\.004\\mathbf\{0\.937\\pm 0\.004\}0\.940±0\.007\\mathbf\{0\.940\\pm 0\.007\}0\.941±0\.006\\mathbf\{0\.941\\pm 0\.006\}0\.866±0\.006\\mathbf\{0\.866\\pm 0\.006\}0\.848±0\.007\\mathbf\{0\.848\\pm 0\.007\}0\.857±0\.006\\mathbf\{0\.857\\pm 0\.006\}w/o HierarchicalGraph Encoding0\.926±\\pm0\.0060\.923±\\pm0\.0050\.927±\\pm0\.0060\.340±\\pm0\.0090\.344±\\pm0\.0070\.348±\\pm0\.001AIDSPAFIR \(Full\)0\.987±0\.003\\mathbf\{0\.987\\pm 0\.003\}0\.991±0\.007\\mathbf\{0\.991\\pm 0\.007\}0\.989±0\.002\\mathbf\{0\.989\\pm 0\.002\}0\.986±0\.001\\mathbf\{0\.986\\pm 0\.001\}0\.985±0\.006\\mathbf\{0\.985\\pm 0\.006\}0\.986±0\.001\\mathbf\{0\.986\\pm 0\.001\}w/o HierarchicalGraph Encoding0\.896±\\pm0\.0040\.863±\\pm0\.0060\.884±\\pm0\.0010\.808±\\pm0\.0060\.820±\\pm0\.0030\.814±\\pm0\.009 On PROTEINS, the F1 score for node classification drops from 0\.918 to 0\.675, while graph classification performance also declines from 0\.832 to 0\.766\. The impact is even more pronounced on the ENZYMES dataset, where graph classification F1 decreases drastically from 0\.857 to 0\.348\. On the AIDS dataset, although the overall performance remains high, the consistent drop in F1 scores for both node and graph classification after removing hierarchical encoding further confirms its role in stabilizing and enhancing structural representation learning\. This suggests that hierarchical structural information plays a crucial role in enabling PAFIR to identify task\-relevant topological and semantic patterns\. The results confirm that incorporating multi\-level structural representations allows the model to generalize better across diverse graph distributions and enhances its capacity to model complex relationships between features and labels\. \\subfigure\[Node Classification\]\\subfigure\[Graph Classification\] Figure 7:Ablation experiment on the graph domain\. We report Precision under two tasks: node classification and graph classification\. ### E\.2Impact of Temporal Encoding To assess the contribution of the time series encoding module in capturing temporal dynamics, we remove it from the PAFIR architecture and compare performance across four forecasting benchmarks: Solar\-Energy, Traffic, ETT, and Weather\. As shown in Table[11](https://arxiv.org/html/2608.18450#A5.T11), removing the temporal encoder results in substantial performance degradation across all datasets\. Table 11:Ablation Experiment for Time Series Encoding\. The input sequence length is fixed to 96 for all methods\.Model VariantSolar\-EnergyTrafficETTh1ETTh2ETTm1ETTm2WeatherMSEMAEMSEMAEMSEMAEMSEMAEMSEMAEMSEMAEMSEMAEPAFIR \(Full\)0\.213\\mathbf\{0\.213\}0\.236\\mathbf\{0\.236\}0\.412\\mathbf\{0\.412\}0\.272\\mathbf\{0\.272\}0\.381\\mathbf\{0\.381\}0\.398\\mathbf\{0\.398\}0\.2880\.3400\.3300\.3660\.1730\.2590\.1700\.211w/o Time Series Encoding0\.7160\.7540\.7880\.7610\.7250\.7390\.7330\.7480\.6560\.6710\.4590\.4700\.4330\.451 For instance, on the Solar\-Energy dataset with a sequence length of 96, removing time series encoding causes the MSE to increase from 0\.213 to 0\.716 and the MAE from 0\.236 to 0\.754\. Consistent and substantial performance degradations are also observed across Traffic, ETTh1, ETTh2, ETTm1, ETTm2, and Weather datasets, where both MSE and MAE increase markedly, in several cases by nearly two to three times\. These results demonstrate that the temporal encoding module plays a crucial role in capturing sequential dependencies and aligning feature evolution with predictive accuracy across diverse time\-series forecasting tasks\. Figure 8:Ablation Experiment on Time Series Domain\. We report MSE under Three Datasets\. ### E\.3Impact of Multi\-Agent Policy To investigate the role of the multi\-agent policy in PAFIR, we conduct an ablation by removing this component and evaluating performance on both node and graph classification tasks\. As presented in Table[12](https://arxiv.org/html/2608.18450#A5.T12), removing the multi\-agent policy leads to a notable drop in graph classification performance, while the impact on node classification is relatively modest\. Table 12:Ablation Experiment for Multi\-Agent Module in RL Framework\.DatasetModel VariantNode ClassificationGraph ClassificationPrecisionRecallF1 ScorePrecisionRecallF1 ScorePROTEINSPAFIR \(Full\)0\.921±0\.007\\mathbf\{0\.921\\pm 0\.007\}0\.914±0\.007\\mathbf\{0\.914\\pm 0\.007\}0\.918±0\.009\\mathbf\{0\.918\\pm 0\.009\}0\.843±0\.006\\mathbf\{0\.843\\pm 0\.006\}0\.822±0\.007\\mathbf\{0\.822\\pm 0\.007\}0\.832±0\.006\\mathbf\{0\.832\\pm 0\.006\}w/o Multi\-Agent Policy0\.916±\\pm0\.0030\.913±0\.002\{0\.913\\pm 0\.002\}0\.917±\\pm0\.0020\.761±\\pm0\.0100\.760±\\pm0\.0070\.759±\\pm0\.007ENZYMESPAFIR \(Full\)0\.937±0\.004\\mathbf\{0\.937\\pm 0\.004\}0\.940±0\.007\\mathbf\{0\.940\\pm 0\.007\}0\.941±0\.006\\mathbf\{0\.941\\pm 0\.006\}0\.866±0\.006\\mathbf\{0\.866\\pm 0\.006\}0\.848±0\.007\\mathbf\{0\.848\\pm 0\.007\}0\.857±0\.006\\mathbf\{0\.857\\pm 0\.006\}w/o Multi\-Agent Policy0\.926±\\pm0\.0030\.932±\\pm0\.0020\.927±\\pm0\.0030\.339±\\pm0\.0400\.348±\\pm0\.0020\.306±\\pm0\.004AIDSPAFIR \(Full\)0\.987±0\.003\\mathbf\{0\.987\\pm 0\.003\}0\.991±0\.007\\mathbf\{0\.991\\pm 0\.007\}0\.989±0\.002\\mathbf\{0\.989\\pm 0\.002\}0\.986±0\.001\\mathbf\{0\.986\\pm 0\.001\}0\.985±0\.006\\mathbf\{0\.985\\pm 0\.006\}0\.986±0\.001\\mathbf\{0\.986\\pm 0\.001\}w/o Multi\-Agent Policy0\.934±\\pm0\.0060\.948±\\pm0\.0020\.940±\\pm0\.0010\.453±\\pm0\.0110\.488±\\pm0\.0020\.474±\\pm0\.005 \\subfigure\[Node Classification\]\\subfigure\[Graph Classification\] Figure 9:Ablation experiment on the graph domain\. We report Precision under two tasks: node classification and graph classification\.For example, on the PROTEINS dataset, the F1 score for graph classification drops from 0\.832 to 0\.759 when the multi\-agent policy is excluded\. A similar but more severe degradation is observed on the ENZYMES dataset, where the F1 score decreases from 0\.857 to 0\.306\. Similarly, on the AIDS dataset, removing the multi\-agent policy leads to a substantial decline in graph classification performance from 0\.986 to 0\.474\. This contrast indicates that the multi\-agent strategy plays a critical role in guiding feature transformation and aggregation at the graph level, where complex interactions between feature subsets are more pronounced\. These results confirm that the cooperative mechanism among agents enables PAFIR to better explore and select informative features, ultimately enhancing its capacity for structured representation learning\. #### Summary of Ablation Findings Across all ablation studies, we observe that each core component of PAFIR contributes uniquely and substantially to the model’s overall performance\. The hierarchical graph encoder is essential for capturing multi\-level structural patterns, with its removal leading to significant performance degradation in both node and graph classification\. The time series encoding module plays a key role in modeling temporal patterns; excluding it results in dramatically increased forecasting errors across all benchmarks\. Finally, the multi\-agent policy proves particularly effective in guiding feature selection at the graph level, as evidenced by the sharp drop in graph classification performance when this component is removed\. Collectively, these results demonstrate that the integrated design of structural encoding, temporal modeling, and the multi\-agent module in the reinforcement learning framework is key to PAFIR’s potential in learning expressive and generalizable representations across diverse data modalities\. \\subfigure\[ENZYMES: Precision\]\\subfigure\[TRAFFIC: MSE\] Figure 10:Experiment Result of Different Methods on Two Datasets\. ## Appendix FCross\-Domain Generality Results We apply the proposed PAFIR framework to two representative tasks, graph\-based classification and multivariate time series forecasting, to evaluate its generality across heterogeneous data modalities\. Specifically, we test PAFIR on the ENZYMES dataset, which involves structural learning over biological graph data, and the TRAFFIC dataset, which captures temporal dynamics in real\-world sensor measurements \(Fig\.[10](https://arxiv.org/html/2608.18450#A5.F10)\)\. These benchmarks represent two fundamentally different problem settings and data characteristics, allowing us to assess the robustness and flexibility of our unified framework\. These two tasks not only differ in terms of data modality, structured graphs versus sequential time series, but also in the nature of their feature dependencies and temporal or topological patterns\. On ENZYMES, PAFIR is compared with state\-of\-the\-art graph learning models, including GraphGPS\+HDSE, TAR, SAT, PCA, and TTG\. On TRAFFIC, we benchmark PAFIR against competitive forecasting models such as iTransformer, RLinear, PatchTST, TimesNet, and SCINet\. On the ENZYMES dataset \(Fig\.[10](https://arxiv.org/html/2608.18450#A5.F10)[10](https://arxiv.org/html/2608.18450#A5.F10)\), PAFIR achieves the highest precision among all competing graph\-based models\. This result highlights PAFIR’s ability to capture discriminative structural features and maintain high predictive accuracy in complex biological graph classification tasks\. On the TRAFFIC dataset \(Fig\.[10](https://arxiv.org/html/2608.18450#A5.F10)[10](https://arxiv.org/html/2608.18450#A5.F10)\), where the evaluation metric is MSE, PAFIR again outperforms temporal forecasting baselines\. Its consistently lower MSE reflects the model’s robustness in capturing temporal dependencies and adapting to traffic dynamics\. The consistent gains across both structural and temporal settings highlight the effectiveness of PAFIR’s unified design in handling heterogeneous data\. This cross\-domain generality underscores its practical value in real\-world scenarios where multimodal health, behavioral, or sensor data are commonly encountered\.
Similar Articles
AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems
This paper introduces AgentForesight, a framework for online auditing and early failure prediction in LLM-based multi-agent systems. It presents a new dataset, AFTraj-22K, and a specialized model, AgentForesight-7B, which outperforms leading proprietary models in detecting decisive errors during trajectory execution.
A Personalized Computational Framework for Assessing the Sufficiency of Partially Observed Data in Healthcare AI models
This paper introduces Feature Sufficiency Analysis (FSA), a framework to determine whether a subset of clinical features is sufficient for AI model predictions, with case studies in postoperative ventilation and mortality prediction.
Adaptive data selection improves wearable prediction under low baseline performance
This paper evaluates adaptive data selection strategies for wearable health prediction, finding they significantly improve AUROC for participants with low baseline performance but offer limited gains for strong baselines.
PATHFinder Agent for Tailored Prenatal Care
This paper presents PATHFinder Agent, an end-to-end conversational AI system that generates personalized prenatal care plans following ACOG's PATH guidelines, integrating patient intake, dynamic dialogue, plan synthesis, and clinician oversight. Evaluation of frontier LLMs, including GPT-5.2, shows promising but incomplete performance, highlighting gaps in antenatal testing recommendations.
Physical activities enable scalable foundation modelling for broad-spectrum health prediction
StepFM is a foundation model that uses only step counter data for broad-spectrum health prediction, offering a privacy-preserving and scalable alternative to high-frequency sensor models.