Benchmarking Multi-Modal Graph-based Social Media Popularity Prediction

arXiv cs.AI Papers

Summary

This paper introduces MMG-Pop, a unified benchmark for multi-modal graph-based social media popularity prediction, and proposes MMG-PopNet, a model that jointly models multimodal content and temporal social interactions. Experiments on Bluesky and Reddit datasets demonstrate superior performance and provide insights into cross-platform generalization and multi-task prediction.

arXiv:2606.27539v1 Announce Type: cross Abstract: Social media popularity prediction aims to forecast the future reach or influence of online content from early-stage observations. Accurate prediction enables key downstream applications, such as advertising optimization and strategic content planning by users, creators, and platforms. Despite substantial progress, existing popularity prediction works often fail to jointly consider multimodal content and temporal social interaction signals. Moreover, the literature remains highly fragmented across datasets, modalities, observation windows, prediction targets, and evaluation protocols. This fragmentation prevents fair comparison and obscures a systematic understanding of how textual, visual, temporal, and interaction-based signals jointly shape popularity dynamics. To address these challenges, we introduce MMG-Pop, a Multi-modal Graph-based Popularity Prediction benchmark, which unifies datasets, modalities, temporal interaction signals, and representative baselines under a standardized evaluation protocol. Furthermore, we propose MMG-PopNet, a unified multi-modal graph-based network that jointly models the aforementioned multi-modal signals and graph-structured social interactions. Extensive experiments on MMG-Pop, comprising four datasets across Bluesky and Reddit platforms, demonstrate the superior performance of MMG-PopNet and yield new insights into cross-platform training generalization, multi-task prediction benefits, multi-modality contributions, and LLM prediction limitation. These findings establish a unified foundation for future research on social dynamics modeling and intervention under heterogeneous modalities and socially-aware agentic ecosystem paradigms.
Original Article
View Cached Full Text

Cached at: 06/29/26, 05:29 AM

# Benchmarking Multi-Modal Graph-based Social Media Popularity Prediction
Source: [https://arxiv.org/html/2606.27539](https://arxiv.org/html/2606.27539)
Utkarsh Sahu1Zhisheng Qi2Li Zhu2Yizhao Yang1Jun Li1 Ryan Rossi3Yu Wang2 1University of Oregon2University of Georgia3Adobe Research \{utkarsh, yizhao, lijun\}@uoregon\.edu \{zq03788, ryanlizhu, Yu\.Wang6\}@uga\.edu ryrossi@adobe\.com

###### Abstract

Social media popularity prediction aims to forecast the future reach or influence of online content from early\-stage observations\. Accurate prediction enables key downstream applications, such as advertising optimization and strategic content planning by users, creators, and platforms\. Despite substantial progress, existing popularity prediction works often fail to jointly consider multimodal content and temporal social interaction signals\. Moreover, the literature remains highly fragmented across datasets, modalities, observation windows, prediction targets, and evaluation protocols\. This fragmentation prevents fair comparison and obscures a systematic understanding of how textual, visual, temporal, and interaction\-based signals jointly shape popularity dynamics\. To address these challenges, we introduceMMG\-Pop, aMulti\-modalGraph\-basedPopularity Prediction benchmark, which unifies datasets, modalities, temporal interaction signals, and representative baselines under a standardized evaluation protocol\. Furthermore, we proposeMMG\-PopNet, a unified multi\-modal graph\-based network that jointly models the aforementioned multi\-modal signals and graph\-structured social interactions\. Extensive experiments on MMG\-Pop, comprising four datasets across Bluesky and Reddit platforms, demonstrate the superior performance of MMG\-PopNet and yield new insights into cross\-platform training generalization, multi\-task prediction benefits, multi\-modality contributions, and LLM prediction limitation\. These findings establish a unified foundation for future research on social dynamics modeling and intervention under heterogeneous modalities and socially\-aware agentic ecosystem paradigms\. The MMG\-Pop benchmark and MMG\-PopNet code are available at this[Link](https://github.com/utkarshxsahu/MM-Pop)\.

### 1Introduction

Social dynamics refers to the patterns of interactions and relationships among individuals within a society[UNESCO](https://arxiv.org/html/2606.27539#bib.bib5); Farmeret al\.\([2018](https://arxiv.org/html/2606.27539#bib.bib38)\); Brock and Durlauf \([2001](https://arxiv.org/html/2606.27539#bib.bib42)\), emerging across diverse real\-world contexts such as public health behavior change, collective responses during crises, and political mobilizationCentola \([2010](https://arxiv.org/html/2606.27539#bib.bib39)\); Newman \([2001](https://arxiv.org/html/2606.27539#bib.bib40)\); Zhouet al\.\([2021](https://arxiv.org/html/2606.27539#bib.bib7)\); Bailet al\.\([2018](https://arxiv.org/html/2606.27539#bib.bib41)\)\. Accurately modeling social dynamics provides critical insights for analyzing, anticipating, and potentially intervening in collective social behaviors \(e\.g\., early detection of toxic information cascades and timely intervention strategies on online platforms\)Bak\-Colemanet al\.\([2022](https://arxiv.org/html/2606.27539#bib.bib43)\); Zhaoet al\.\([2015b](https://arxiv.org/html/2606.27539#bib.bib44)\); Shaoet al\.\([2018](https://arxiv.org/html/2606.27539#bib.bib45)\); Chenget al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib46)\); Linet al\.\([2021](https://arxiv.org/html/2606.27539#bib.bib47)\)In this work, we focus on one of the most important social dynamics modeling problems, social media popularity prediction, which aims to leverage early observations of social content to predict its future popularity/influence \(e\.g\., number of likes, reposts\) across diverse social contexts and modalitiesLerman and Hogg \([2010](https://arxiv.org/html/2606.27539#bib.bib60)\); Meghawatet al\.\([2018](https://arxiv.org/html/2606.27539#bib.bib16)\); Szabo and Huberman \([2010](https://arxiv.org/html/2606.27539#bib.bib12)\)\. Effectively predicting popularity on social media has important implications for both platforms and users\. For platforms, it supports content recommendation, trend forecasting, advertising, and efficient allocation of moderation resourcesPintoet al\.\([2013](https://arxiv.org/html/2606.27539#bib.bib61)\); Cobbe \([2021](https://arxiv.org/html/2606.27539#bib.bib62)\); Tanget al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib63)\); Javari and Jalili \([2014](https://arxiv.org/html/2606.27539#bib.bib64)\)by estimating which content is likely to attract substantial future engagement\. For users, creators, and organizations, it helps estimate future reach, optimize posting content and promotion strategiesMazloomet al\.\([2016](https://arxiv.org/html/2606.27539#bib.bib65)\); Zhanget al\.\([2018c](https://arxiv.org/html/2606.27539#bib.bib66)\), and plan social or marketing campaigns more effectivelyYuet al\.\([2011](https://arxiv.org/html/2606.27539#bib.bib67)\); Van Aelstet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib68)\)\.

Existing WorkSocial Modality SignalPopularity MetricPopularity Prediction MethodGraphTextImageTimeVideo@UsernameYang and Counts\([2010](https://arxiv.org/html/2606.27539#bib.bib26)\)✔✘✘✔✘Speed/Width/DepthCox PH \+ Log\-linear RegressionResubmissionsLakkarajuet al\.\([2013](https://arxiv.org/html/2606.27539#bib.bib27)\)✘✔✘✔✘Reddit KarmaSupervised LDA \+ Linear RegressionSzabo\-HubermanSzabo and Huberman\([2010](https://arxiv.org/html/2606.27539#bib.bib12)\)✘✘✘✔✘View Count/ Digg votesLog\-linear RegressionFlickr\-SVRKhoslaet al\.\([2014](https://arxiv.org/html/2606.27539#bib.bib15)\)✘✔✔✘✘View CountSupport Vector RegressionSEISMICZhaoet al\.\([2015a](https://arxiv.org/html/2606.27539#bib.bib19)\)✘✘✘✔✘Cascade sizeSelf\-exciting point processGalton–WatsonMedvedevet al\.\([2019](https://arxiv.org/html/2606.27539#bib.bib4)\)✔✘✘✔✘Cascade sizeBranching Hawkes processHIPLakkarajuet al\.\([2013](https://arxiv.org/html/2606.27539#bib.bib27)\)✘✔✘✔✘View CountHawkes Intensity ProcessesDeepHawkesCaoet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib1)\)✔✘✘✔✘Cascade sizeGRU \+ Time DecayDeepCasLiet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib2)\)✔✘✘✘✘Cascade sizeBi\-directional GRUCasSeqGCNWanget al\.\([2022](https://arxiv.org/html/2606.27539#bib.bib3)\)✔✘✘✔✘Cascade sizeGCN \+ LSTMTSGNNLiuet al\.\([2021](https://arxiv.org/html/2606.27539#bib.bib25)\)✔✘✘✘✘Cascade sizeGAT \+ GLUGraphLSTMZayats and Ostendorf\([2018](https://arxiv.org/html/2606.27539#bib.bib22)\)✔✔✘✔✘Reddit KarmaGraph\-structured LSTMUHANZhanget al\.\([2018b](https://arxiv.org/html/2606.27539#bib.bib30)\)✘✔✔✘✘View CountMulti\-modal AttentionHMMVEDXieet al\.\([2021](https://arxiv.org/html/2606.27539#bib.bib57)\)✘✔✔✘✔Comment/Repost/Likes/ViewHierarchical Multimodal VAEMMRAZhonget al\.\([2024](https://arxiv.org/html/2606.27539#bib.bib58)\)✘✔✔✘✔View CountMulti\-modal Attention \+ RetrievalMASSLZhanget al\.\([2022](https://arxiv.org/html/2606.27539#bib.bib59)\)✘✔✔✘✔View CountMultimodal VAE

Table 1:Prior social media popularity prediction methods, organized by modality signal, prediction metric, modeling approach, and method category: feature engineering, statistical, and deep learning\.Prior social media popularity prediction can be categorized into three lines\. The first predicts social media popularity based on the initial media content without considering its subsequent spreadSzabo and Huberman \([2010](https://arxiv.org/html/2606.27539#bib.bib12)\); Bandariet al\.\([2012](https://arxiv.org/html/2606.27539#bib.bib13)\); Tsur and Rappoport \([2012](https://arxiv.org/html/2606.27539#bib.bib14)\); Khoslaet al\.\([2014](https://arxiv.org/html/2606.27539#bib.bib15)\); Gelliet al\.\([2015](https://arxiv.org/html/2606.27539#bib.bib17)\); Dinget al\.\([2019](https://arxiv.org/html/2606.27539#bib.bib18)\)\. However, they are content\-centric and fail to account for how early social interaction dynamics, such as high\-impact comments and influencers’ resharesGarciaet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib6)\), impact eventual popularity\. The second leverages observed interaction patterns \(e\.g\., reshares, replies, and user engagements\) to model how a post propagates through the social interaction network\. These interaction histories are represented as cascades and modeled using point processes or geometric deep learning \(e\.g\., RNNs/GNNs\) to capture the temporal/structural dynamics of information diffusionZhaoet al\.\([2015a](https://arxiv.org/html/2606.27539#bib.bib19)\); Medvedevet al\.\([2019](https://arxiv.org/html/2606.27539#bib.bib4)\); Caoet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib1)\); Liet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib2)\); Wanget al\.\([2022](https://arxiv.org/html/2606.27539#bib.bib3)\)\. However, they under\-utilize the semantic content of posts and responses, such as textual meaning and visual signals that shape engagement\. A third line of work jointly leverages both post content and interaction structureAragónet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib20)\); Zayats and Ostendorf \([2018](https://arxiv.org/html/2606.27539#bib.bib22)\); Zhanget al\.\([2018a](https://arxiv.org/html/2606.27539#bib.bib23)\); Zubiagaet al\.\([2016](https://arxiv.org/html/2606.27539#bib.bib24)\)\. However, these studies primarily focus on tasks such as toxicity detection, sentiment analysis, rumor propagation, or conversation modeling, rather than directly addressing popularity prediction\. More recently, emerging generative models and agentic AI have been applied to social dynamics modeling, either through LLM\-based multi\-agent simulationsLiuet al\.\([2025](https://arxiv.org/html/2606.27539#bib.bib35)\); Yanget al\.\([2024](https://arxiv.org/html/2606.27539#bib.bib70)\)or by directly repurposing LLMs as autoregressive cascade predictors or reasoning\-augmented regressors for popularity forecastingZhenget al\.\([2025](https://arxiv.org/html/2606.27539#bib.bib34)\); Xuet al\.\([2025a](https://arxiv.org/html/2606.27539#bib.bib33)\)\. However, none of these approaches has been systematically benchmarked against prior non\-LLM baselines\.

In addition to the above limitations, existing social media popularity prediction studies are constructed under heterogeneous yet inconsistent experimental settings as in Table[1](https://arxiv.org/html/2606.27539#S1.T1)\. These differences span dataset versions \(e\.g\., XCaoet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib1)\); Liet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib2)\); Wanget al\.\([2022](https://arxiv.org/html/2606.27539#bib.bib3)\)versus \(vs\.\) RedditMedvedevet al\.\([2019](https://arxiv.org/html/2606.27539#bib.bib4)\); Zayats and Ostendorf \([2018](https://arxiv.org/html/2606.27539#bib.bib22)\)\), modality signals \(e\.g\., text, image and videoZhanget al\.\([2022](https://arxiv.org/html/2606.27539#bib.bib59)\); Zhonget al\.\([2024](https://arxiv.org/html/2606.27539#bib.bib58)\); Xieet al\.\([2021](https://arxiv.org/html/2606.27539#bib.bib57)\)vs\. graph topologyLiet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib2)\); Caoet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib1)\); Liuet al\.\([2021](https://arxiv.org/html/2606.27539#bib.bib25)\)\), prediction targets \(e\.g\., cascade sizeLiuet al\.\([2021](https://arxiv.org/html/2606.27539#bib.bib25)\); Zhaoet al\.\([2015a](https://arxiv.org/html/2606.27539#bib.bib19)\); Wanget al\.\([2022](https://arxiv.org/html/2606.27539#bib.bib3)\)vs\. content view countSzabo and Huberman \([2010](https://arxiv.org/html/2606.27539#bib.bib12)\); Zhanget al\.\([2018b](https://arxiv.org/html/2606.27539#bib.bib30)\)\)\. This fragmentation prevents meaningful comparison of existing methods, thereby hindering the derivation of key insights, such as which modalities contribute most and whether cross\-platform transferability exists\. Furthermore, existing popularity prediction benchmarksXuet al\.\([2025b](https://arxiv.org/html/2606.27539#bib.bib55)\); Wuet al\.\([2023](https://arxiv.org/html/2606.27539#bib.bib56)\)mainly focus on initial media content, overlooking the evolving cascade of user interactions over time\. This motivates us to develop both a unified benchmarkMMG\-Popand a popularity prediction modelMMG\-Pop\-Netthat jointly capture multi\-modal content and temporal social interactions\. The key contributions are summarized as follows:

- •Unified Popularity Prediction Benchmark\.We introduce theMMG\-Popbenchmark, standardizing datasets, social modality signals, observation windows, prediction horizons, and popularity measures to enable consistent evaluation across in\-domain forecasting, future\-horizon prediction, and cross\-platform transfer, with representative baselines\.
- •Unified Multi\-Modal Model\.We proposeMMG\-Pop\-Net, the first unified architecture to jointly model multimodal content, graph\-structured interaction dynamics, and temporal signals through bidirectional graph message passing, supporting multi\-objective popularity prediction\.
- •Comprehensive Experiments and Novel Insights\.We conduct extensive experiments to demonstrate the advantages of MMG\-Pop\-Net and the insights enabled by the MMG\-Pop benchmark, highlighting the importance of jointly modeling multimodal content and cascade structure, the generalization gains from cross\-community training, the benefits of multi\-task training for engagement prediction, and the limited ability of LLMs to predict social popularity\.

### 2Design Space of MMG\-Pop Popularity Prediction Benchmark

This section outlines the design space of our proposedMMG\-Popbenchmark for social media popularity prediction, encompassing the notation and problem formulation, popularity measurement, training and evaluation, dataset curation, and existing baselines\.

![Refer to caption](https://arxiv.org/html/2606.27539v1/figure/fig_1.png)Figure 1:Overview of MMG\-Pop Benchmark\.Social cascades from Bluesky and Reddit are represented as tree\-structured graphs, where each node carries multi\-modal attributes\. Given only an early observed prefixGtG^\{t\}, the task is to predict six complementary popularity dimensions characterizing the future cascade stateGt′G^\{t^\{\\prime\}\}\. The benchmark evaluates baselines alongside our proposedMMG\-PopNetacross multiple observation windows\.#### 2\.1Notation and Problem Formulation

Notation of Social Media Popularity Prediction\.We represent an online social cascade as a directed tree\-structured graphG=\(𝒱,ℰ\)G=\(\\mathcal\{V\},\\mathcal\{E\}\), where𝒱\\mathcal\{V\}denotes the set of nodes andℰ⊆𝒱×𝒱\\mathcal\{E\}\\subseteq\\mathcal\{V\}\\times\\mathcal\{V\}denotes the set of directed edges\. Each nodev∈𝒱v\\in\\mathcal\{V\}corresponds to a content item \(e\.g\., a post, comment, or reply\), and a directed edge\(u,v\)∈ℰ\(u,v\)\\in\\mathcal\{E\}indicates that nodevvis generated in response to nodeuu, capturing the information propagation \(e\.g\., reply\-to, repost/reshare, or quote relationships\)\. The graph is rooted at a unique nodevroot∈𝒱v\_\{\\text\{root\}\}\\in\\mathcal\{V\}representing the initial item of the cascade \(e\.g\., a starter post on Bluesky or a submission on Reddit\)\. Each nodev∈𝒱v\\in\\mathcal\{V\}is associated with multimodal attributesXv=\(XvText,XvVisual,XvGraph,XvSocial,XvTime\)X\_\{v\}=\(X\_\{v\}^\{\\text\{Text\}\},\\,X\_\{v\}^\{\\text\{Visual\}\},\\,X\_\{v\}^\{\\text\{Graph\}\},\\,X\_\{v\}^\{\\text\{Social\}\},\\,X\_\{v\}^\{\\text\{Time\}\}\)whereXvTextX\_\{v\}^\{\\text\{Text\}\}denotes textual content features \(e\.g\., post/comment text, hashtags, or semantic embeddingsZhanget al\.\([2018b](https://arxiv.org/html/2606.27539#bib.bib30)\); Lakkarajuet al\.\([2013](https://arxiv.org/html/2606.27539#bib.bib27)\)\),XvVisualX\_\{v\}^\{\\text\{Visual\}\}denotes visual features when media is present \(e\.g\., images, videos, or visual descriptorsKhoslaet al\.\([2014](https://arxiv.org/html/2606.27539#bib.bib15)\); Zhanget al\.\([2022](https://arxiv.org/html/2606.27539#bib.bib59)\)\),XvGraphX\_\{v\}^\{\\text\{Graph\}\}denotes local structural or network context \(e\.g\., neighborhood subgraph statistics or position within the cascadeZayats and Ostendorf \([2018](https://arxiv.org/html/2606.27539#bib.bib22)\)\),XvSocialX\_\{v\}^\{\\text\{Social\}\}denotes author\-level social context features \(e\.g\., user profile attributesKeneshlooet al\.\([2016](https://arxiv.org/html/2606.27539#bib.bib74)\)and interaction graph\-derived influence proxies, including centrality and PageRank scoresGuo and Shakarian \([2016](https://arxiv.org/html/2606.27539#bib.bib75)\); Brin and Page \([1998](https://arxiv.org/html/2606.27539#bib.bib71)\)\), andXvTimeX\_\{v\}^\{\\text\{Time\}\}denotes temporal features \(e\.g\., global timestamp or relative time to parent nodes within the cascadeZayats and Ostendorf \([2018](https://arxiv.org/html/2606.27539#bib.bib22)\)\)\. In addition, each cascade may be associated with thread\-level contextual metadataXThreadX^\{\\text\{Thread\}\}, capturing properties of the root postvrootv\_\{\\text\{root\}\}\(e\.g\., topic, presence of visual media, or root\-author follower count\)\.

Formulation of Social Media Popularity Prediction\.The core objective is to predict the future evolution of a social cascade given only its early\-stage observations\. Lett≥0t\\geq 0denote the elapsed time since the root post\. Given a cascade graphG=\(𝒱,ℰ\)G=\(\\mathcal\{V\},\\mathcal\{E\}\), we define the observed prefix at timettasGt=\(𝒱t,ℰt\)G^\{t\}=\(\\mathcal\{V\}^\{t\},\\mathcal\{E\}^\{t\}\), where𝒱t=\{v∈𝒱∣tv≤t\}\\mathcal\{V\}^\{t\}=\\\{v\\in\\mathcal\{V\}\\mid t\_\{v\}\\leq t\\\}andℰt\\mathcal\{E\}^\{t\}contains all edges among nodes in𝒱t\\mathcal\{V\}^\{t\}\. This prefix captures the historical context of the cascade, including time\-truncated multimodal node attributes\{Xvt\}v∈𝒱t\\\{X\_\{v\}^\{t\}\\\}\_\{v\\in\\mathcal\{V\}^\{t\}\}and thread\-level contextXThread,tX^\{\\text\{Thread\},t\}observable up to timett\. The prediction target is the future state of the cascade,𝐘G∈ℝK\\mathbf\{Y\}\_\{G\}\\in\\mathbb\{R\}^\{K\}, consisting ofKKpopularity measures defined in Section[2\.2](https://arxiv.org/html/2606.27539#S2.SS2)\. We aim to learn a parametric mappingℱ𝚯:\(Gt,\{Xvt\}v∈𝒱t,XThread,t\)↦𝐘G\\mathcal\{F\}\_\{\\bm\{\\Theta\}\}:\\left\(G^\{t\},\\\{X\_\{v\}^\{t\}\\\}\_\{v\\in\\mathcal\{V\}^\{t\}\},X^\{\\text\{Thread\},t\}\\right\)\\mapsto\\mathbf\{Y\}\_\{G\}\.

#### 2\.2Popularity Measurement

Social popularity can be quantified in multiple ways\. Our benchmark, MMG\-Pop, considers six distinct dimensions to comprehensively capture popularity dynamics, following prior literatureVosoughiet al\.\([2018](https://arxiv.org/html/2606.27539#bib.bib36)\); Goelet al\.\([2016](https://arxiv.org/html/2606.27539#bib.bib28)\); Zhanget al\.\([2021](https://arxiv.org/html/2606.27539#bib.bib37)\); Szabo and Huberman \([2010](https://arxiv.org/html/2606.27539#bib.bib12)\)\. We categorize these intostructural,participation, andengagementtasks:

- •Max Width:It measures the largest breadth of the cascadeVosoughiet al\.\([2018](https://arxiv.org/html/2606.27539#bib.bib36)\); Zhanget al\.\([2021](https://arxiv.org/html/2606.27539#bib.bib37)\)\. It is quantified as the maximum number of nodes appearing at any depth level in the cascade graphGG\.
- •Max Depth:It measures the length of the longest reply chain in the cascadeZhanget al\.\([2021](https://arxiv.org/html/2606.27539#bib.bib37)\); Szabo and Huberman \([2010](https://arxiv.org/html/2606.27539#bib.bib12)\)\. It is quantified as the maximum distance fromvrootv\_\{\\text\{root\}\}to any node in the cascade graphGG\.
- •Structural Virality:It quantifies whether cascade diffusion is dominated by shallow broadcast spread or deeper multi\-hop propagationGoelet al\.\([2016](https://arxiv.org/html/2606.27539#bib.bib28)\)\. It is measured as the average shortest\-path distance between all pairs of distinct nodes in the cascade graph\.
- •Size:It measures how many content items are generated in the cascadeZhanget al\.\([2021](https://arxiv.org/html/2606.27539#bib.bib37)\); Liet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib2)\)\. It is quantified as the number of nodes in the cascade graphGG, i\.e\.,\|𝒱\|\|\\mathcal\{V\}\|\.
- •Unique Users:It measures the number of unique users who participate in the cascadeZhanget al\.\([2021](https://arxiv.org/html/2606.27539#bib.bib37)\)\.
- •Like Score:It captures platform\-visible engagement received by the root post contentvrootv\_\{\\text\{root\}\}through platform\-specific metrics such as likes or up\-votes \(e\.g\., Reddit Karma score\)Szabo and Huberman \([2010](https://arxiv.org/html/2606.27539#bib.bib12)\)\.

#### 2\.3Training and Evaluation

Training\.Let𝒢\\mathcal\{G\}denote the set of cascades, partitioned into𝒢Train∪𝒢Val∪𝒢Test\\mathcal\{G\}^\{\\text\{Train\}\}\\cup\\mathcal\{G\}^\{\\text\{Val\}\}\\cup\\mathcal\{G\}^\{\\text\{Test\}\}\. For each cascadeG∈𝒢TrainG\\in\\mathcal\{G\}^\{\\text\{Train\}\}, we construct its observed prefixGtG^\{t\}with truncated node/thread\-level features\{Xvt\}\\\{X\_\{v\}^\{t\}\\\}andXThread,tX^\{\\text\{Thread\},t\}\. The objective is to predict future outcomes at a later timet′\>tt^\{\\prime\}\>t\. Given heavy\-tailed nature of social popularity signalsChaet al\.\([2009](https://arxiv.org/html/2606.27539#bib.bib72)\); Tataret al\.\([2014](https://arxiv.org/html/2606.27539#bib.bib73)\), we define popularity targets in the log\-transformed space:Y~Gt′=log⁡\(𝐈\+YGt′\)\\widetilde\{\\textbf\{Y\}\}\_\{G\}^\{t^\{\\prime\}\}=\\log\(\\mathbf\{I\}\+\\textbf\{Y\}\_\{G\}^\{t^\{\\prime\}\}\)\. The model is trained to predictY~Gt′\\widetilde\{\\textbf\{Y\}\}\_\{G\}^\{t^\{\\prime\}\}by optimizing:𝚯∗=arg⁡min𝚯​∑G∈𝒢Trainℒ​\(ℱ𝚯​\(Gt,\{Xvt\}v∈𝒱t,XThread,t\),Y~Gt′\)\\bm\{\\Theta\}^\{\*\}=\\arg\\min\_\{\\bm\{\\Theta\}\}\\sum\_\{G\\in\\mathcal\{G\}^\{\\text\{Train\}\}\}\\mathcal\{L\}\(\\mathcal\{F\}\_\{\\bm\{\\Theta\}\}\(G^\{t\},\\\{X\_\{v\}^\{t\}\\\}\_\{v\\in\\mathcal\{V\}^\{t\}\},X^\{\\text\{Thread\},t\}\),\\widetilde\{\\textbf\{Y\}\}\_\{G\}^\{t^\{\\prime\}\}\)whereℒ\\mathcal\{L\}is the mean squared error \(MSE\) loss:ℒ​\(Y^,Y~\)=1d​‖Y^−Y~‖22,\\mathcal\{L\}\(\\widehat\{\\textbf\{Y\}\},\\widetilde\{\\textbf\{Y\}\}\)=\\frac\{1\}\{d\}\|\|\\widehat\{\\textbf\{Y\}\}\-\\widetilde\{\\textbf\{Y\}\}\|\|\_\{2\}^\{2\},withdddenoting the number of prediction targets\.

Evaluation\.During evaluation, the learned modelℱ𝚯∗\\mathcal\{F\}\_\{\\bm\{\\Theta\}\}^\{\*\}is applied to unseen cascadesG∈𝒢TestG\\in\\mathcal\{G\}^\{\\text\{Test\}\}, using only their observed prefixesGtG^\{t\}and corresponding features\. We evaluate performance by comparing the predicted future cascade dynamics at timet′t^\{\\prime\}with the corresponding ground\-truth values\. We report MSE,R2R^\{2\}, and Spearman correlation, all computed in the log\-transformed space\.

#### 2\.4Dataset Curation

We curate social cascades from Bluesky and Reddit, two platforms with distinct platform dynamics\. Each discussion thread is represented as a tree\-structured social cascade following the formulation in Section[2\.1](https://arxiv.org/html/2606.27539#S2.SS1), where the root post isvrootv\_\{\\text\{root\}\}, posts or comments are nodes in𝒱\\mathcal\{V\}, and parent–reply relations define directed edges\(u,v\)∈ℰ\(u,v\)\\in\\mathcal\{E\}\. Node attributesXvX\_\{v\}and thread\-level contextXThreadX^\{\\text\{Thread\}\}are instantiated from the available platform metadata\. Bluesky\.We curate our Bluesky subset from the large\-scale collection ofFailla and Rossetti \([2024](https://arxiv.org/html/2606.27539#bib.bib29)\), which originally contains approximately 235 million posts from 4 million users between February 2023 and March 2024\. We construct cascades from reply interactions, which provide explicit conversational content for modeling discussion dynamics\. Reddit\.We use Pushshift data dumpsBaumgartneret al\.\([2020](https://arxiv.org/html/2606.27539#bib.bib32)\)from three communities: r/AMA, r/Gaming, and r/Futurology\. These subreddits capture complementary discussion styles: centralized Q&A, media\-rich entertainment discussion, and speculative scientific discourse\. We construct cascades from submissions and their comment reply trees\. The datasets cover July 2021–December 2024 for r/AMA, January 2023–August 2024 for r/Gaming, and August 2019–December 2024 for r/Futurology\. Detailed dataset information is provided in Appendix[A](https://arxiv.org/html/2606.27539#A1)\.

#### 2\.5Representative Baselines

We evaluate representative baselines spanning structure\-agnostic, temporal, sequence\-based, graph\-based, and content\-aware cascade modeling baselines\.MLPis a structure\-agnostic baseline that represents each cascade using root\-post features, aggregated reply features, and global thread metadata\.DeepHawkesCaoet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib1)\)captures temporal diffusion dynamics through user embeddings, diffusion\-path encoding, and time\-decay modeling\.DeepCasLiet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib2)\)models cascades as sampled diffusion paths and learns sequence representations with attention\.CasSeqGCNWanget al\.\([2022](https://arxiv.org/html/2606.27539#bib.bib3)\)represents temporal graph evolution by encoding graph snapshots with GCN to model time progression\.GraphLSTMZayats and Ostendorf \([2018](https://arxiv.org/html/2606.27539#bib.bib22)\)is a content\-aware graph sequence baseline that models reply\-tree structure with textual, user, temporal, and structural features\. Additional details are provided in the Appendix[B](https://arxiv.org/html/2606.27539#A2)\.

![Refer to caption](https://arxiv.org/html/2606.27539v1/x1.png)Figure 2:Overview of MMG\-PopNet Model:The model embeds node\-level text and temporal signals for bidirectional graph message passing over the cascade, and root visual content and thread metadata are encoded as separate contextual features\. The learned root, graph, visual, and metadata representations are fused at the prediction stage to support multi\-task popularity forecasting\.

### 3Foundational Multi\-modal Graph\-based Popularity Prediction Network

This section proposes a unified framework that integrates multimodal content and graph\-structured social interactions for social dynamics prediction, as shown in Figure[2](https://arxiv.org/html/2606.27539#S2.F2)\. Our architecture first encodes heterogeneous multimodal cascade content \(text semantics, visual media, and global context\) into a unified representation, and then applies graph message passing to incorporate cascade social interaction structure, yielding a shared representation for multi\-task popularity prediction\.

Multi\-Modal Feature Embedding\.To encode heterogeneous cascade attributes, we design modality\-specific encoders tailored to the semantics of each signal\. For each nodev∈𝒱tv\\in\\mathcal\{V\}^\{t\}, we encode the available node\-level textual and temporal attributes\. For textual contentXvTextX\_\{v\}^\{\\text\{Text\}\}, we use Transformer to obtain𝐳vText,t=F𝚯EncTextText​\(XvText,t\)\\mathbf\{z\}\_\{v\}^\{\\text\{Text\},t\}=F\_\{\\bm\{\\Theta\}\_\{\\text\{Enc\}\}^\{\\text\{Text\}\}\}^\{\\text\{Text\}\}\(X\_\{v\}^\{\\text\{Text\},t\}\)\. For temporal information, we encode relative timing signals viaΔvTime=\[Δ​tv←parent,Δ​tv←root\]\\Delta\_\{v\}^\{\\text\{Time\}\}=\[\\Delta t\_\{v\\leftarrow\\text\{parent\}\},\\Delta t\_\{v\\leftarrow\\text\{root\}\}\], whereΔ​tv←parent\\Delta t\_\{v\\leftarrow\\text\{parent\}\}andΔ​tv←root\\Delta t\_\{v\\leftarrow\\text\{root\}\}denote the elapsed time since the parent node and the root post, reflecting the immediacy of user engagement and overall temporal stage\. These timing signals are log\-transformed and z\-score normalized to obtain fixed temporal features𝐳vTime,t\\mathbf\{z\}\_\{v\}^\{\\text\{Time,t\}\}\. We also encode thread\-level context asXThread,t=\(XrootVisual,t,XmetaThread,t\)X^\{\\text\{Thread\},t\}=\(X\_\{\\text\{root\}\}^\{\\text\{Visual\},t\},X\_\{\\text\{meta\}\}^\{\\text\{Thread\},t\}\), denoting root post visual content and non\-visual thread context related to the root post\. The root image is encoded with a CLIPRadfordet al\.\([2021](https://arxiv.org/html/2606.27539#bib.bib85)\),𝐳rootVisual,t=F𝚯EncVisualVisual​\(XrootVisual,t\)\\mathbf\{z\}\_\{\\text\{root\}\}^\{\\text\{Visual,t\}\}=F\_\{\\bm\{\\Theta\}\_\{\\text\{Enc\}\}^\{\\text\{Visual\}\}\}^\{\\text\{Visual\}\}\(X\_\{\\text\{root\}\}^\{\\text\{Visual\},t\}\)\. Finally, non\-visual thread context is modeled with initial user’s influence and posting time, i\.e\.,XmetaThread,t=\[Followers​\(vroot\),ϕ​\(tpost\)\]X\_\{\\text\{meta\}\}^\{\\text\{Thread\},t\}=\[\\mathrm\{Followers\}\(v\_\{\\text\{root\}\}\),\\phi\(t\_\{\\text\{post\}\}\)\], where the follower count is log\-transformed and standardized, andϕ​\(⋅\)\\phi\(\\cdot\)is a cyclic encoding of time\-of\-day as content posting time during a day also influences interactions\. The resulting root visual𝐳rootVisual,t\\mathbf\{z\}\_\{\\text\{root\}\}^\{\\text\{Visual,t\}\}and non\-visualXmetaThread,tX\_\{\\text\{meta\}\}^\{\\text\{Thread\},t\}representations together concatenate into the thread\-level contextual representation𝐳Thread,t\\mathbf\{z\}^\{\\text\{Thread,t\}\}\.

Bidirectional Graph Message\-Passing\.While node\-level textual and temporal encodings capture rich local signals at each node, modeling the structural context of the cascade is essential to understand how information propagates and evolves over time\. In particular, early\-stage cascade dynamics, such as branching patterns and response depth, provide strong indicators of future popularity\. To encode such structural dependencies, we perform graph message\-passing over the cascade\. We first initialize each nodev∈𝒱tv\\in\\mathcal\{V\}^\{t\}representation by concatenating semantic and temporal features:𝐡v\(0\)=Concat​\(𝐳vText,𝐳vTime\)\\mathbf\{h\}\_\{v\}^\{\(0\)\}=\\mathrm\{Concat\}\\bigl\(\\mathbf\{z\}\_\{v\}^\{\\text\{Text\}\},\\mathbf\{z\}\_\{v\}^\{\\text\{Time\}\}\\bigr\), capturing both content and temporal context\. We then apply bidirectional message\-passing to simultaneously model root\-to\-leaf and leaf\-to\-root propagation\. The node embeddings are initialized in both directions𝐡v,down\(0\)=𝐡v,up\(0\)=𝐡v\(0\)\\mathbf\{h\}\_\{v,\\text\{down\}\}^\{\(0\)\}=\\mathbf\{h\}\_\{v,\\text\{up\}\}^\{\(0\)\}=\\mathbf\{h\}\_\{v\}^\{\(0\)\}are updated at layerℓ\\ell:

𝐡v,down\(ℓ\)=SAGEdown\(ℓ\)​\(𝐡v,down\(ℓ−1\),\{𝐡u,down\(ℓ−1\):u∈𝒩down​\(v\)\}\),𝐡v,up\(ℓ\)=SAGEup\(ℓ\)​\(𝐡v,up\(ℓ−1\),\{𝐡u,up\(ℓ−1\):u∈𝒩up​\(v\)\}\),\\scriptsize\\mathbf\{h\}\_\{v,\\text\{down\}\}^\{\(\\ell\)\}=\\mathrm\{SAGE\}\_\{\\text\{down\}\}^\{\(\\ell\)\}\(\\mathbf\{h\}\_\{v,\\text\{down\}\}^\{\(\\ell\-1\)\},\\\{\\mathbf\{h\}\_\{u,\\text\{down\}\}^\{\(\\ell\-1\)\}:u\\in\\mathcal\{N\}\_\{\\text\{down\}\}\(v\)\\\}\),\\mathbf\{h\}\_\{v,\\text\{up\}\}^\{\(\\ell\)\}=\\mathrm\{SAGE\}\_\{\\text\{up\}\}^\{\(\\ell\)\}\(\\mathbf\{h\}\_\{v,\\text\{up\}\}^\{\(\\ell\-1\)\},\\\{\\mathbf\{h\}\_\{u,\\text\{up\}\}^\{\(\\ell\-1\)\}:u\\in\\mathcal\{N\}\_\{\\text\{up\}\}\(v\)\\\}\),\(1\)where𝒩down​\(v\)\\mathcal\{N\}\_\{\\text\{down\}\}\(v\)and𝒩up​\(v\)\\mathcal\{N\}\_\{\\text\{up\}\}\(v\)denote the parent/child\-side neighbors of nodevv\. AfterLLlayers, we aggregate node representations via mean pooling to obtain direction\-specific summaries𝐡G,downt,𝐡G,upt\\mathbf\{h\}\_\{G,\\text\{down\}\}^\{t\},\\mathbf\{h\}\_\{G,\\text\{up\}\}^\{t\}, which are then concatenated to form the final graph representation:𝐡Gt=Concat​\(𝐡G,downt,𝐡G,upt\)\\mathbf\{h\}\_\{G\}^\{t\}=\\mathrm\{Concat\}\(\\mathbf\{h\}\_\{G,\\text\{down\}\}^\{t\},\\mathbf\{h\}\_\{G,\\text\{up\}\}^\{t\}\)\.

Multi\-Modal Feature Fusion and Multi\-Task Prediction\.After encoding node/graph/thread\-level signals, we fuse them into a unified cascade representation for prediction\. Specifically, we aggregate \(1\) the raw root node representation𝐡root0\\mathbf\{h\}^\{0\}\_\{\\text\{root\}\}before message passing, \(2\) the final bidirectional root representation,𝐡root\(L\)=Concat​\(𝐡root,down\(L\),𝐡root,up\(L\)\)\\mathbf\{h\}\_\{\\text\{root\}\}^\{\(L\)\}=\\mathrm\{Concat\}\(\\mathbf\{h\}\_\{\\text\{root\},\\text\{down\}\}^\{\(L\)\},\\mathbf\{h\}\_\{\\text\{root\},\\text\{up\}\}^\{\(L\)\}\), \(3\) the graph structural representationhGt\\textbf\{h\}\_\{G\}^\{t\}, and \(4\) the thread\-level contextual representation, into a single vector:𝐱Gt=𝐡root\(0\)​‖𝐡root\(L\)‖​𝐠Gt∥𝐳Thread,t\.\\mathbf\{x\}\_\{G\}^\{t\}=\\mathbf\{h\}\_\{\\text\{root\}\}^\{\(0\)\}\\;\\\|\\;\\mathbf\{h\}\_\{\\text\{root\}\}^\{\(L\)\}\\;\\\|\\;\\mathbf\{g\}\_\{G\}^\{t\}\\;\\\|\\;\\mathbf\{z\}^\{\\text\{Thread,t\}\}\.We then apply a shared multi\-layer perceptron \(MLP\) to map the fused representation into a multi\-task output space:𝐘^Gt′=MLP𝚯Pred​\(𝐱Gt\)\\widehat\{\\mathbf\{Y\}\}\_\{G\}^\{t^\{\\prime\}\}=\\mathrm\{MLP\}\_\{\\bm\{\\Theta\}\_\{\\text\{Pred\}\}\}\(\\mathbf\{x\}\_\{G\}^\{t\}\), where𝐘^G∈ℝK\\widehat\{\\mathbf\{Y\}\}\_\{G\}\\in\\mathbb\{R\}^\{K\}contains log\-space predictions forKKtargets \(e\.g\., final cascade size, unique users, or structural properties\)\. The model is trained using a multi\-task objectiveℒ=1K​∑k=1Kℒk​\(Y^G\(k\),YG\(k\)\)\\mathcal\{L\}=\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}\\mathcal\{L\}\_\{k\}\(\\widehat\{Y\}\_\{G\}^\{\(k\)\},Y\_\{G\}^\{\(k\)\}\), where all parameters \(including the text encoder𝚯EncText\{\\bm\{\\Theta\}\_\{\\text\{Enc\}\}^\{\\text\{Text\}\}\}, image encoder𝚯EncImage\{\\bm\{\\Theta\}\_\{\\text\{Enc\}\}^\{\\text\{Image\}\}\}, two GNNs, and prediction heads𝚯Pred\{\\bm\{\\Theta\}\_\{\\text\{Pred\}\}\}\) are jointly optimized end\-to\-end\.

### 4Related Work

Social Dynamics Modeling\.Social dynamics modeling studies how local interactions among individuals give rise to collective outcomes such as consensus, segregation, polarization, and information diffusionCastellanoet al\.\([2009](https://arxiv.org/html/2606.27539#bib.bib76)\); Flacheet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib77)\)\. Classical models explain these phenomena through simple but expressive mechanisms, whereSchelling \([1971](https://arxiv.org/html/2606.27539#bib.bib78)\)showed how individual preferences can produce macro\-level segregation, while threshold and cascade models describe how behaviors spread once social reinforcement exceeds adoption barriersGranovetter \([1978](https://arxiv.org/html/2606.27539#bib.bib79)\); Watts \([2002](https://arxiv.org/html/2606.27539#bib.bib80)\)\. Opinion dynamics and social\-influence models further categorize how network structure, homophily, and repeated exposure shape agreement, diversity, and polarizationDeGroot \([1974](https://arxiv.org/html/2606.27539#bib.bib81)\); Friedkin and Johnsen \([1990](https://arxiv.org/html/2606.27539#bib.bib82)\); Flacheet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib77)\)\. With online platforms, this perspective has expanded to large\-scale information diffusion, misinformation spread, and intervention analysis, where temporal interactions and network topology jointly determine collective trajectoriesGuilleet al\.\([2013](https://arxiv.org/html/2606.27539#bib.bib83)\); Bak\-Colemanet al\.\([2022](https://arxiv.org/html/2606.27539#bib.bib43)\); Bailet al\.\([2018](https://arxiv.org/html/2606.27539#bib.bib41)\)\. Recent agent\-based and LLM\-driven simulations enrich this line by modeling adaptive, language\-mediated agents to simulate broad social dynamicsYanget al\.\([2024](https://arxiv.org/html/2606.27539#bib.bib70)\); Guoet al\.\([2024](https://arxiv.org/html/2606.27539#bib.bib84)\)\.

Social Media Popularity Prediction\.Social media popularity prediction models and forecasts social dynamics by predicting the future influence, engagement, or diffusion of online contentSzabo and Huberman \([2010](https://arxiv.org/html/2606.27539#bib.bib12)\); Zhouet al\.\([2021](https://arxiv.org/html/2606.27539#bib.bib7)\)\. It has broad applications in trend forecasting, advertisement targeting, and public opinion analysis\. Existing studies can be categorized based on social signals and modeling paradigms\. From the signal perspective, prior work includes content\-based methods leveraging textual or visual informationSzabo and Huberman \([2010](https://arxiv.org/html/2606.27539#bib.bib12)\); Bandariet al\.\([2012](https://arxiv.org/html/2606.27539#bib.bib13)\); Tsur and Rappoport \([2012](https://arxiv.org/html/2606.27539#bib.bib14)\); Khoslaet al\.\([2014](https://arxiv.org/html/2606.27539#bib.bib15)\); Gelliet al\.\([2015](https://arxiv.org/html/2606.27539#bib.bib17)\); Dinget al\.\([2019](https://arxiv.org/html/2606.27539#bib.bib18)\), structure\-based methods modeling diffusion topology and user interactionsZhaoet al\.\([2015a](https://arxiv.org/html/2606.27539#bib.bib19)\); Medvedevet al\.\([2019](https://arxiv.org/html/2606.27539#bib.bib4)\); Caoet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib1)\); Liet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib2)\); Wanget al\.\([2022](https://arxiv.org/html/2606.27539#bib.bib3)\), and hybrid approaches combining bothAragónet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib20)\); Zayats and Ostendorf \([2018](https://arxiv.org/html/2606.27539#bib.bib22)\); Zhanget al\.\([2018a](https://arxiv.org/html/2606.27539#bib.bib23)\); Zubiagaet al\.\([2016](https://arxiv.org/html/2606.27539#bib.bib24)\)\. From the modeling perspective, earlier studies mainly relied on handcrafted features with classical machine learningZhaoet al\.\([2015a](https://arxiv.org/html/2606.27539#bib.bib19)\); Medvedevet al\.\([2019](https://arxiv.org/html/2606.27539#bib.bib4)\); Lakkarajuet al\.\([2013](https://arxiv.org/html/2606.27539#bib.bib27)\), followed by sequential and geometric deep learning methods, including LSTMs and GNNs, to capture temporal and structural dynamicsCaoet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib1)\); Liet al\.\([2017](https://arxiv.org/html/2606.27539#bib.bib2)\); Wanget al\.\([2022](https://arxiv.org/html/2606.27539#bib.bib3)\); Liuet al\.\([2021](https://arxiv.org/html/2606.27539#bib.bib25)\); Zayats and Ostendorf \([2018](https://arxiv.org/html/2606.27539#bib.bib22)\); Zhanget al\.\([2018b](https://arxiv.org/html/2606.27539#bib.bib30)\); Xieet al\.\([2021](https://arxiv.org/html/2606.27539#bib.bib57)\); Zhonget al\.\([2024](https://arxiv.org/html/2606.27539#bib.bib58)\); Zhanget al\.\([2022](https://arxiv.org/html/2606.27539#bib.bib59)\)\. More recently, agentic social simulation approachesLiuet al\.\([2025](https://arxiv.org/html/2606.27539#bib.bib35)\); Yanget al\.\([2024](https://arxiv.org/html/2606.27539#bib.bib70)\); Zhenget al\.\([2025](https://arxiv.org/html/2606.27539#bib.bib34)\); Xuet al\.\([2025a](https://arxiv.org/html/2606.27539#bib.bib33)\)have been explored to model contextual user behaviors and social interactions\. However, they remain fragmented across modalities, platforms, prediction targets, and evaluation protocols, motivating the need for a unified benchmark\.

### 5Experiment

We conduct extensive experiments with the MMG\-Pop benchmark, designed to systematically address the following research questions:

𝐐1\\mathbf\{Q\}\_\{1\}:How do different baselines and MMG\-Pop\-Net perform in predicting cascade popularity across varying future horizons and early observation windows under our MMG\-Pop benchmark?

𝐐2\\mathbf\{Q\}\_\{2\}:Does unified training across communities and platforms improve social popularity prediction?

𝐐3\\mathbf\{Q\}\_\{3\}:Can MM\-LLMs serve as competitive predictors for multimodal social popularity forecasting?

𝐐4\\mathbf\{Q\}\_\{4\}:Does multi\-objective prediction benefit from jointly predicting popularity targets?

𝐐5\\mathbf\{Q\}\_\{5\}:How do different modalities contribute to popularity prediction?

Detailed experiment settings are described in Appendix[C](https://arxiv.org/html/2606.27539#A3)\.

#### 5\.1𝐐1\\mathbf\{Q\}\_\{1\}: Popularity Prediction across Varying Future Horizons/Early Observation Windows\.

Popularity prediction of final cascade state under different early observation\.Given an early observation windowtt, we aim to predict the final popularity of a social cascade based on the cascade state observed up to the data collection time\. To assess the role of early information, we consider a root\-only setting and three early\-observation windows for each dataset\. The root\-only setting, denoted as window 0, includes only the root post and its thread\-level context\. The remaining windows capture increasingly mature cascade prefixes: 2, 10, and 20 minutes for Bluesky; 15, 30, and 60 minutes for r/AMA; 20, 50, and 90 minutes for r/Gaming; and 30, 90, and 180 minutes for r/Futurology\. Window lengths vary by dataset because cascades unfold at different speeds across platforms and communities\. Table[2](https://arxiv.org/html/2606.27539#S5.T2)reports MSE in the log\-transformed target space across all datasets, observation windows, and prediction targets\. MMG\-PopNet achieves the best overall average MSE for every target category, with 4\.6% to 17\.0% reductions over the strongest non\-MMG\-PopNet baselines\. Among the baselines, Graph\-LSTM is the strongest competitor on most structural targets, reflecting the value of reply\-tree structure, temporal ordering, and textual content\. CasSeqGCN also performs competitively, likely by modeling evolving cascade as snapshots, which capture propagation topology\. MLP is particularly strong forLike Scoreprediction because this target measures engagement received by the root post, whose features are directly combined with the mean representation of early observed nodes\. However, without message passing, MLP cannot fully model cascade\-level dependencies, whereas MMG\-PopNet integrates these early signals through bidirectional graph propagation and achieves the lowest average MSE across all targets\.

Table 2:MSE results for final cascade\-state prediction under different early observation windows, coveringstructural prediction tasks\(max width, max depth, structural virality, and cascade size\),unique\-user prediction, andlike\-score prediction\. Lower values indicate better performance\. The best results are highlighted inbold, while the second\-best results areunderlined\. Statistical significance analyses show that improvements are significant across settings in Table[11](https://arxiv.org/html/2606.27539#A4.T11)\.TaskModelBlueskyr/AMAr/Gamingr/FuturologyAvg021020Avg0153060Avg0205090Avg03090180AvgMax WidthMLP0\.4580\.4270\.3710\.3500\.4020\.6640\.5280\.4930\.4150\.5251\.7551\.1820\.8770\.7081\.1281\.8101\.3410\.6970\.4571\.0760\.783DeepHawkes0\.6840\.6170\.4030\.3380\.5100\.7130\.5540\.4910\.3650\.5311\.8881\.2681\.1470\.8391\.2881\.8761\.6431\.1840\.7831\.3710\.925DeepCas0\.6330\.6750\.6760\.6760\.6660\.8410\.7570\.7670\.7090\.7692\.3611\.8951\.8331\.6901\.9451\.7381\.6431\.1901\.1011\.4431\.199CasSeqGCN0\.6760\.5380\.3570\.2860\.4650\.7130\.5340\.4460\.3130\.5021\.8641\.1400\.8410\.5871\.1081\.8201\.2580\.6790\.3661\.0310\.776Graph\-LSTM0\.6570\.5260\.3450\.2810\.4520\.6860\.5100\.4300\.3050\.4831\.7711\.1550\.8110\.5181\.0641\.7491\.1490\.6310\.3450\.9690\.742MMG\-PopNet0\.4570\.3690\.2840\.2340\.3360\.6250\.4910\.4290\.2760\.4551\.5571\.0960\.7140\.5620\.9821\.5811\.1320\.6650\.3720\.9380\.678Max DepthMLP0\.3560\.3460\.2960\.2700\.3180\.3180\.2670\.2250\.1980\.2520\.3440\.2790\.2140\.1860\.2560\.5930\.4840\.2890\.2210\.3970\.305DeepHawkes0\.3670\.3580\.3440\.3000\.3420\.3290\.2800\.2530\.2400\.2760\.3460\.2900\.2690\.2290\.2840\.6020\.5670\.4290\.3250\.4820\.345DeepCas0\.3610\.3640\.3640\.3640\.3630\.4320\.3750\.3410\.3330\.3700\.5770\.4090\.3870\.3400\.4270\.5960\.5540\.4100\.3740\.4830\.411CasSeqGCN0\.3640\.3480\.2940\.2580\.3160\.3280\.2810\.2270\.1930\.2570\.3470\.2960\.2470\.2100\.2750\.5860\.4770\.3170\.2300\.4030\.313Graph\-LSTM0\.3640\.3400\.2850\.2430\.3080\.3230\.2640\.2150\.1760\.2450\.3530\.2880\.2090\.1610\.2530\.5700\.4530\.2980\.2320\.3880\.298MMG\-PopNet0\.3630\.3250\.2730\.2380\.3000\.3140\.2560\.2190\.1720\.2400\.3350\.2750\.2000\.1600\.2430\.5170\.4180\.2820\.2050\.3560\.284StructuralViralityMLP0\.1470\.1430\.1230\.1140\.1320\.1290\.1050\.0890\.0740\.0990\.0960\.0780\.0590\.0490\.0710\.2090\.1640\.1020\.0770\.1380\.110DeepHawkes0\.1580\.1540\.1440\.1190\.1440\.1340\.1110\.1010\.0940\.1100\.0950\.0850\.0830\.0690\.0820\.2000\.1950\.1570\.1270\.1700\.127DeepCas0\.1550\.1570\.1570\.1570\.1570\.2180\.1680\.1460\.1410\.1670\.2650\.1540\.1400\.1170\.1690\.2130\.2000\.1650\.1450\.1810\.169CasSeqGCN0\.1570\.1480\.1240\.1090\.1350\.1340\.1110\.0890\.0700\.1020\.0950\.0860\.0740\.0660\.0800\.1960\.1620\.1160\.0810\.1390\.114Graph\-LSTM0\.1570\.1450\.1220\.1040\.1320\.1310\.1060\.0850\.0650\.0970\.1030\.0830\.0600\.0440\.0730\.1950\.1570\.1060\.0840\.1360\.109MMG\-PopNet0\.1550\.1350\.1140\.1010\.1260\.1300\.1000\.0870\.0640\.0950\.0980\.0810\.0570\.0450\.0700\.1820\.1460\.1010\.0730\.1260\.104SizeMLP0\.6950\.6520\.5560\.5150\.6050\.9180\.7110\.6260\.5190\.6892\.0181\.3990\.9930\.8011\.3032\.7472\.0941\.0660\.6371\.6371\.059DeepHawkes0\.8990\.8170\.6180\.4990\.7090\.9820\.7380\.6350\.4780\.7092\.1371\.4721\.3140\.9391\.4662\.9032\.5561\.9041\.1592\.1301\.253DeepCas0\.8320\.8790\.8800\.8800\.8681\.2621\.0931\.0531\.0021\.1032\.9522\.2282\.1521\.9182\.3132\.6732\.5171\.7671\.6052\.0161\.606CasSeqGCN0\.8810\.7510\.5510\.4590\.6600\.9790\.7350\.5990\.4210\.6842\.1011\.4001\.0210\.7151\.3092\.7602\.0151\.1240\.5831\.6211\.068Graph\-LSTM0\.8590\.7330\.5310\.4460\.6420\.9450\.7000\.5720\.4050\.6562\.0311\.4180\.9880\.6411\.2702\.6401\.8521\.0520\.5631\.5271\.024MMG\-PopNet0\.7050\.5870\.4700\.3970\.5400\.8630\.6610\.5730\.3660\.6161\.8291\.3170\.8420\.6531\.1602\.3561\.7430\.9970\.5591\.4140\.932UniqueUsersMLP0\.4650\.4320\.3740\.3510\.4050\.6570\.5310\.5020\.4180\.5271\.8831\.3060\.9490\.7561\.2242\.1391\.6420\.8470\.5121\.2900\.860DeepHawkes0\.7340\.6600\.4390\.3710\.5510\.7070\.5590\.5050\.3700\.5332\.0091\.3921\.2570\.9171\.3942\.2281\.9961\.4390\.8961\.6391\.030DeepCas0\.6660\.7210\.7230\.7220\.7060\.8690\.7670\.7710\.7160\.7792\.6292\.0192\.0021\.7962\.1122\.0731\.9761\.3811\.2691\.6751\.319CasSeqGCN0\.7230\.5840\.3950\.3180\.5050\.7060\.5430\.4680\.3280\.5111\.9791\.2940\.9660\.6711\.2282\.1571\.5830\.8720\.4491\.2650\.877Graph\-LSTM0\.7000\.5600\.3680\.2960\.4810\.6790\.5140\.4480\.3200\.4901\.8981\.3070\.9280\.5991\.1832\.0741\.4410\.8160\.4561\.1970\.838MMG\-PopNet0\.4670\.3830\.3020\.2550\.3520\.6220\.4940\.4470\.2900\.4631\.6931\.2180\.7920\.6131\.0791\.8591\.3740\.7850\.4391\.1140\.752Like ScoreMLP1\.3111\.2981\.2541\.2321\.2741\.3511\.2841\.2921\.1931\.2805\.6335\.5934\.9914\.6605\.2196\.4626\.1044\.0613\.7465\.0933\.217DeepHawkes2\.5322\.4002\.0101\.8552\.1991\.4381\.4111\.4061\.3251\.3956\.2406\.1225\.9736\.1306\.1166\.8856\.8806\.0465\.3046\.0313\.997DeepCas2\.3562\.5022\.5072\.5052\.4931\.4441\.4281\.4831\.4351\.4506\.3596\.0996\.0486\.2686\.1945\.7255\.7854\.5764\.9055\.2733\.839CasSeqGCN2\.5082\.2661\.9011\.7492\.1041\.4361\.4061\.4001\.3071\.3876\.2066\.1276\.0436\.0046\.0956\.7786\.6045\.4285\.0005\.4523\.885Graph\-LSTM2\.4672\.1861\.7911\.6072\.0131\.3601\.3111\.3371\.2371\.3115\.7075\.5765\.3024\.8665\.3636\.2395\.3954\.3434\.3935\.0923\.445MMG\-PopNet1\.2601\.0871\.0260\.9501\.0811\.3111\.2121\.1661\.0281\.1795\.0674\.6544\.3074\.0734\.5254\.9994\.5963\.2142\.7503\.8902\.669

![Refer to caption](https://arxiv.org/html/2606.27539v1/x2.png)Figure 3:MSE\-Loss trajectories across datasets, comparing different models for targetSize\.Lower is better\. MMG\-PopNet achieves the lowest MSE, with the strongest gains at later horizons\.Popularity Prediction of Cascade States at Future Horizons\.Beyond prediction at the terminal state, we evaluate social media popularity prediction at intermediate future horizons of cascades\. Given an early observed prefixGtG^\{t\}\(as described in previous section\), each method predicts future cascade outcomes at later timest′\>tt^\{\\prime\}\>t, using targets computed from the cascade state at horizons\{4​h,8​h,16​h,24​h\}\\\{4\\text\{h\},8\\text\{h\},16\\text\{h\},24\\text\{h\}\\\}\. This setting tests whether a model can forecast not only terminal popularity, but also the trajectory of cascade growth over time\. Figure[3](https://arxiv.org/html/2606.27539#S5.F3)reports MSE trajectories for theSizetarget across datasets\. Prediction error generally increases with the forecasting horizon, reflecting the greater uncertainty of longer\-range cascade growth\. Despite this increased difficulty, MMG\-PopNet consistently achieves the lowest error across datasets and horizons\. Graph\-LSTM and CasSeqGCN are the closest baselines, with comparable performance at 4h and 8h, but they fall behind at 16h and 24h as forecasting uncertainty increases\. In contrast, DeepCas has substantially higher error in several cases, with clipped values indicating that its trajectory predictions fall outside the plotted range\. Similar trends are observed for other popularity targets in Appendix[D](https://arxiv.org/html/2606.27539#A4)\.

#### 5\.2𝐐2\\mathbf\{Q\}\_\{2\}: Unified Training Across Communities and Platforms

We investigate whether popularity prediction benefits from unified training across datasets from multiple platforms\. Instead of training isolated MMG\-PopNet models per dataset and observation window, we train a single model on the combined cascades from all datasets\. This evaluates whether joint supervision over heterogeneous cascades improves generalization compared to dataset\-specific training\. Figure[4](https://arxiv.org/html/2606.27539#S5.F4)demonstrates that unified training yields substantial performance gains on Reddit while maintaining comparable accuracy on Bluesky\. On Reddit, the unified model reduces average MSE across all popularity targets, achieving dramatic error reductions forLike Score,Size, andUnique Users\. Conversely, dataset\-specific training retains a marginal edge on Bluesky\. This pattern suggests that the unified model benefits most when cross\-community training shares a platform\-level interaction structure, while Bluesky introduces distinct dynamics less represented in the combined training distribution\. Detailed results and analysis can be found in Appendix[E](https://arxiv.org/html/2606.27539#A5)\.

![Refer to caption](https://arxiv.org/html/2606.27539v1/x3.png)Figure 4:Dataset\-Specific vs\. Unified Training\.Avg MSE of MMG\-PopNet under dataset\-specific and unified training\. Lower is better\. Unified training greatly improves performance on Reddit communities and remains competitive on Bluesky\.![Refer to caption](https://arxiv.org/html/2606.27539v1/x4.png)Figure 5:Normalized LLM Performance on Bluesky\.Scores are normalized with MMG\-PopNet as the reference baseline, fixed at1\.01\.0on all axes, where smaller areas indicate worse performance\. MMG\-PopNet outperforms LLM baselines across all settings\. Among LLMs, retrieval\-augmented few\-shot prompting performs better in sparse early windows, while fine\-tuning becomes stronger as longer cascade prefixes provide richer temporal and structural training signals\.
#### 5\.3𝐐3\\mathbf\{Q\}\_\{3\}: Comparison with LLM\-Based Approaches

We compare MMG\-PopNet with multimodal LLM models under three settings: zero\-shot prompting, retrieval\-augmented few\-shot prompting, and supervised fine\-tuning\. Zero\-shot setting serialized the observed cascade prefix as a structured JSON input prompt for LLM\-based prediction\. The retrieval\-augmented few\-shot setting includes four training examples with similar root posts\. Fine\-tuning setting trains LLM with early\-observation cascade inputs\. Figure[5](https://arxiv.org/html/2606.27539#S5.F5)shows the normalized comparison on Bluesky\. Scores are computed asMSEmodel/MSEMMG−PopNet\\mathrm\{MSE\}\_\{\\mathrm\{model\}\}/\\mathrm\{MSE\}\_\{\\mathrm\{MMG\-PopNet\}\}, so MMG\-PopNet forms the reference score of1\.01\.0on every axis, and smaller polygons indicate worse performance\. MMG\-PopNet uniformly outperforms all LLM baselines across all targets and observation windows\. Zero\-shot prompting yields the highest error, proving that direct prompting lacks the numerical calibration of LLMs for popularity prediction despite structured inputs\. Few\-shot prompting rivals or exceeds fine\-tuning given root\-only or early windows, where historical examples provide crucial context for sparse cascades\. Fine\-tuning overtakes prompting as the window expands, likely because it better exploits the richer reply structure, temporal progression, and participation signals in longer observation windows\. Overall, MMG\-PopNet performs better likely because it models topology, timing features, and multimodal context for popularity prediction more directly, rather than relying on prompt\-driven inference over serialized cascade inputs\. Detailed setting in Appendix[F](https://arxiv.org/html/2606.27539#A6)\.

#### 5\.4𝐐5\\mathbf\{Q\}\_\{5\}: Single\-Task vs\. Multi\-Task Training

We examine whether multi\-objective prediction benefits from jointly modeling complementary popularity targets\. Here, we compare the multi\-task MMG\-PopNet with task\-specific variants trained independently for each target\. Table[3](https://arxiv.org/html/2606.27539#S5.T3)shows that the benefit of joint training depends on the target\. Single\-task training is stronger for topology\-driven objectives\. It achieves lower MSE forMax WidthandMax Depthacross all datasets, indicating that these structural properties benefit from dedicated training\. Similar trends appear forStructural ViralityandUnique Users, although the gaps are smaller\. Multi\-task training remains competitive forSize, matching or slightly improving over single\-task models on three datasets\. Joint modeling is most beneficial for engagement\. Multi\-task MMG\-PopNet lowersLike ScoreMSE on every dataset, with large gains on Bluesky and r/Futurology\. This suggests that engagement prediction can benefit from shared signals of cascade structure and user participation\. Overall, joint modeling does not improve every target\. However, it offers a useful deployment trade\-off by remaining competitive on most outcomes while consistently improvingLike Scoreprediction\.

Table 3:Comparison of multi\-task versus single\-task MSE results across datasets using the following windows: Bluesky @ 10min, r/AMA @ 30min, r/Gaming @ 50min, and r/Futurology @ 90min\.Best in bold\.TaskBlueskyr/AMAr/Gamingr/FuturologySingleMultiSingleMultiSingleMultiSingleMultiMax Width0\.2800\.2840\.4110\.4290\.6630\.7140\.5830\.665Max Depth0\.2720\.2730\.2070\.2190\.1820\.2000\.2720\.282StructuralVirality0\.1150\.1140\.0820\.0870\.0500\.0570\.0940\.101Size0\.4940\.4700\.5780\.5730\.8430\.8420\.9870\.997Unique Users0\.2950\.3020\.4310\.4470\.7660\.7920\.7740\.785Like Score1\.5751\.0261\.2841\.1664\.4824\.3073\.7083\.214Figure 6:Modality Ablation: Avg\. MSEincreaseper excluded modality relative to full MMG\-PopNet\.
![Refer to caption](https://arxiv.org/html/2606.27539v1/x5.png)

#### 5\.5𝐐6\\mathbf\{Q\}\_\{6\}: Modality Ablation Analysis

We evaluate how each modality contributes by removing one input source from MMG\-PopNet at a time and measuring the relative MSE increase over the full model\. The ablations remove textual semanticsXTextX^\{\\text\{Text\}\}, root visual contentXVisualX^\{\\text\{Visual\}\}, node temporal featuresXTimeX^\{\\text\{Time\}\}, or reply\-tree topologyXGraphX^\{\\text\{Graph\}\}\.Figure[6](https://arxiv.org/html/2606.27539#S5.F6)reports results on r/Gaming and r/Futurology, where all four modalities are available\. All removals increase error, showing that each modality adds useful information\. Temporal features have the largest effect on cascade growth and participation\. They encode each node’s time since its parent reply and since the root post\. Removing these features sharply hurtsMax Width\(21\.2%\),Unique Users\(21\.0%\), andSize\(18\.7%\)\. These features capture the pace of early discussion\. Fast replies signal bursty growth relevant to width and size, while slower temporal medians can indicate longer\-lived discussions with more distinct users\. Text is most important for engagement\. Removing textual semantics increasesLike Scoreerror by 16\.5%, while only mildly affecting structural targets\. This suggests that audience approval depends strongly on what is said, not only how the cascade grows\. Reply\-tree topology mainly supports structural prediction, especiallyStructural ViralityandMax Depth\. Root visual content has the smallest effect, but its consistent gains indicate a modest complementary role\. Detailed results in Appendix[G](https://arxiv.org/html/2606.27539#A7)\.

### 6Conclusion and Future Work

In this paper, we introduced MMG\-Pop, a unified benchmark for multi\-modal social media popularity prediction, and MMG\-PopNet, a unified model that captures content, temporal dynamics, and reply structure to forecast multiple forms of popularity\. Our experiments show that MMG\-Pop enables systematic evaluation across datasets, communities, observation windows, prediction horizons, and engagement targets\. Results show that MMG\-PopNet improves prediction by jointly modeling multimodal content and cascade structure, with different modalities offering complementary signals\. In addition, cross\-community training improves generalization, while multi\-task training captures shared engagement patterns\. In contrast, LLMs remain limited in predicting social popularity, suggesting that language understanding alone is insufficient for modeling social dynamics\. Together, these findings establish MMG\-Pop as a useful benchmark and MMG\-PopNet as an effective model for integrating the signals that shape social media popularity\. Furthermore, we have conducted real\-world case studies with MMG\-PopNet in Appendix[H](https://arxiv.org/html/2606.27539#A8)\. Future work will explore agentic social simulation to predict popularity\. Limitations of this work and additional discussion are in Appendix[I](https://arxiv.org/html/2606.27539#A9)\.

### References

- \[1\]\(2017\)Generative models of online discussion threads: state of the art and research challenges\.Journal of Internet Services and Applications\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[2\]C\. A\. Bail, L\. P\. Argyle, T\. W\. Brown, J\. P\. Bumpus, H\. Chen, M\. F\. Hunzaker, J\. Lee, M\. Mann, F\. Merhout, and A\. Volfovsky\(2018\)Exposure to opposing views on social media can increase political polarization\.Proceedings of the National Academy of Sciences\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1),[§4](https://arxiv.org/html/2606.27539#S4.p1.1)\.
- \[3\]J\. B\. Bak\-Coleman, I\. Kennedy, M\. Wack, A\. Beers, J\. S\. Schafer, E\. S\. Spiro, K\. Starbird, and J\. D\. West\(2022\)Combining interventions to reduce the spread of viral misinformation\.Nature Human Behaviour\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1),[§4](https://arxiv.org/html/2606.27539#S4.p1.1)\.
- \[4\]R\. Bandari, S\. Asur, and B\. Huberman\(2012\)The pulse of news in social media: forecasting popularity\.InProceedings of the International AAAI Conference on Web and Social Media,Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[5\]O\. Barkan, E\. Hauon, A\. Caciularu, O\. Katz, I\. Malkiel, O\. Armstrong, and N\. Koenigstein\(2021\)Grad\-sam: explaining transformers via gradient self\-attention maps\.InProceedings of the 30th ACM International Conference on Information & Knowledge Management,pp\. 2882–2887\.Cited by:[Appendix H](https://arxiv.org/html/2606.27539#A8.p2.1)\.
- \[6\]J\. Baumgartner, S\. Zannettou, B\. Keegan, M\. Squire, and J\. Blackburn\(2020\)The pushshift reddit dataset\.InProceedings of the international AAAI conference on web and social media,Vol\.14,pp\. 830–839\.Cited by:[§2\.4](https://arxiv.org/html/2606.27539#S2.SS4.p1.5)\.
- \[7\]S\. Brin and L\. Page\(1998\)The anatomy of a large\-scale hypertextual web search engine\.Computer networks and ISDN systems\.Cited by:[§2\.1](https://arxiv.org/html/2606.27539#S2.SS1.p1.17)\.
- \[8\]W\. A\. Brock and S\. N\. Durlauf\(2001\)Discrete choice with social interactions\.The Review of Economic Studies\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[9\]Q\. Cao, H\. Shen, K\. Cen, W\. Ouyang, and X\. Cheng\(2017\)Deephawkes: bridging the gap between prediction and understanding of information cascades\.InProceedings of the 2017 ACM on Conference on Information and Knowledge Management,Cited by:[§B\.2](https://arxiv.org/html/2606.27539#A2.SS2.p1.1),[Table 1](https://arxiv.org/html/2606.27539#S1.T1.2.1.10.1),[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§1](https://arxiv.org/html/2606.27539#S1.p3.1),[§2\.5](https://arxiv.org/html/2606.27539#S2.SS5.p1.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[10\]C\. Castellano, S\. Fortunato, and V\. Loreto\(2009\)Statistical physics of social dynamics\.Reviews of modern physics\.Cited by:[§4](https://arxiv.org/html/2606.27539#S4.p1.1)\.
- \[11\]D\. Centola\(2010\)The spread of behavior in an online social network experiment\.science\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[12\]M\. Cha, A\. Mislove, and K\. P\. Gummadi\(2009\)A measurement\-driven analysis of information propagation in the flickr social network\.InProceedings of the 18th international conference on World wide web,Cited by:[§2\.3](https://arxiv.org/html/2606.27539#S2.SS3.p1.13)\.
- \[13\]J\. Cheng, M\. Bernstein, C\. Danescu\-Niculescu\-Mizil, and J\. Leskovec\(2017\)Anyone can become a troll: causes of trolling behavior in online discussions\.InProceedings of the 2017 ACM conference on computer supported cooperative work and social computing,Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[14\]J\. Cobbe\(2021\)Algorithmic censorship by social platforms: power and resistance\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[15\]M\. H\. DeGroot\(1974\)Reaching a consensus\.Journal of the American Statistical association69\(345\),pp\. 118–121\.Cited by:[§4](https://arxiv.org/html/2606.27539#S4.p1.1)\.
- \[16\]K\. Ding, R\. Wang, and S\. Wang\(2019\)Social media popularity prediction: a multiple feature fusion approach with deep neural networks\.InProceedings of the 27th ACM International Conference on Multimedia,Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[17\]B\. Efron and R\. J\. Tibshirani\(1994\)An introduction to the bootstrap\.Chapman and Hall/CRC\.Cited by:[§D\.3](https://arxiv.org/html/2606.27539#A4.SS3.p1.10)\.
- \[18\]A\. Failla and G\. Rossetti\(2024\)“I’m in the bluesky tonight”: insights from a year worth of social data\.PloS one\.Cited by:[§2\.4](https://arxiv.org/html/2606.27539#S2.SS4.p1.5)\.
- \[19\]T\. W\. Farmer, B\. Talbott, M\. Dawes, H\. B\. Huber, D\. S\. Brooks, and E\. E\. Powers\(2018\)Social dynamics management: what is it and why is it important for intervention?\.Journal of Emotional and Behavioral Disorders\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[20\]A\. Flache, M\. Mäs, T\. Feliciani, E\. Chattoe\-Brown, G\. Deffuant, S\. Huet, and J\. Lorenz\(2017\)Models of social influence: towards the next frontiers\.Journal of Artificial Societies and Social Simulation\.Cited by:[§4](https://arxiv.org/html/2606.27539#S4.p1.1)\.
- \[21\]N\. E\. Friedkin and E\. C\. Johnsen\(1990\)Social influence and opinions\.Journal of mathematical sociology15\(3\-4\),pp\. 193–206\.Cited by:[§4](https://arxiv.org/html/2606.27539#S4.p1.1)\.
- \[22\]D\. Garcia, P\. Mavrodiev, D\. Casati, and F\. Schweitzer\(2017\)Understanding popularity, reputation, and social influence in the twitter society\.Policy & Internet9\(3\),pp\. 343–364\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p2.1)\.
- \[23\]F\. Gelli, T\. Uricchio, M\. Bertini, A\. Del Bimbo, and S\. Chang\(2015\)Image popularity prediction in social media using sentiment and context features\.InProceedings of the 23rd ACM international conference on Multimedia,Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[24\]J\. Gildenblat and contributors\(2021\)PyTorch library for cam methods\.GitHub\.Note:[https://github\.com/jacobgil/pytorch\-grad\-cam](https://github.com/jacobgil/pytorch-grad-cam)Cited by:[Appendix H](https://arxiv.org/html/2606.27539#A8.p2.1)\.
- \[25\]S\. Goel, A\. Anderson, J\. Hofman, and D\. J\. Watts\(2016\)The structural virality of online diffusion\.Management science\.Cited by:[3rd item](https://arxiv.org/html/2606.27539#S2.I1.i3.p1.1),[§2\.2](https://arxiv.org/html/2606.27539#S2.SS2.p1.1)\.
- \[26\]M\. Granovetter\(1978\)Threshold models of collective behavior\.American journal of sociology\.Cited by:[§4](https://arxiv.org/html/2606.27539#S4.p1.1)\.
- \[27\]A\. Guille, H\. Hacid, C\. Favre, and D\. A\. Zighed\(2013\)Information diffusion in online social networks: a survey\.ACM Sigmod Record42\(2\),pp\. 17–28\.Cited by:[§4](https://arxiv.org/html/2606.27539#S4.p1.1)\.
- \[28\]R\. Guo and P\. Shakarian\(2016\)A comparison of methods for cascade prediction\.In2016 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining \(ASONAM\),Cited by:[§2\.1](https://arxiv.org/html/2606.27539#S2.SS1.p1.17)\.
- \[29\]T\. Guo, X\. Chen, Y\. Wang, R\. Chang, S\. Pei, N\. V\. Chawla, O\. Wiest, and X\. Zhang\(2024\)Large language model based multi\-agents: a survey of progress and challenges\.arXiv preprint arXiv:2402\.01680\.Cited by:[§4](https://arxiv.org/html/2606.27539#S4.p1.1)\.
- \[30\]W\. Hamilton, Z\. Ying, and J\. Leskovec\(2017\)Inductive representation learning on large graphs\.Advances in neural information processing systems\.Cited by:[§C\.1](https://arxiv.org/html/2606.27539#A3.SS1.p1.1)\.
- \[31\]A\. Javari and M\. Jalili\(2014\)Accurate and novel recommendations: an algorithm based on popularity forecasting\.ACM Transactions on Intelligent Systems and Technology \(TIST\)\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[32\]Y\. Keneshloo, S\. Wang, E\. Han, and N\. Ramakrishnan\(2016\)Predicting the popularity of news articles\.InProceedings of the 2016 SIAM international conference on data mining,Cited by:[§2\.1](https://arxiv.org/html/2606.27539#S2.SS1.p1.17)\.
- \[33\]A\. Khosla, A\. Das Sarma, and R\. Hamid\(2014\)What makes an image popular?\.InProceedings of the 23rd international conference on World wide web,pp\. 867–876\.Cited by:[Table 1](https://arxiv.org/html/2606.27539#S1.T1.2.1.6.1),[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.27539#S2.SS1.p1.17),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[34\]H\. Lakkaraju, J\. McAuley, and J\. Leskovec\(2013\)What’s in a name? understanding the interplay between titles, content, and communities in social media\.InProceedings of the international AAAI conference on web and social media,Cited by:[Table 1](https://arxiv.org/html/2606.27539#S1.T1.2.1.4.1),[Table 1](https://arxiv.org/html/2606.27539#S1.T1.2.1.9.1),[§2\.1](https://arxiv.org/html/2606.27539#S2.SS1.p1.17),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[35\]K\. Lerman and T\. Hogg\(2010\)Using a model of social dynamics to predict popularity of news\.InProceedings of the 19th international conference on World wide web,Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[36\]C\. Li, J\. Ma, X\. Guo, and Q\. Mei\(2017\)Deepcas: an end\-to\-end predictor of information cascades\.InProceedings of the 26th international conference on World Wide Web,pp\. 577–586\.Cited by:[§B\.3](https://arxiv.org/html/2606.27539#A2.SS3.p1.1),[Table 1](https://arxiv.org/html/2606.27539#S1.T1.2.1.11.1),[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§1](https://arxiv.org/html/2606.27539#S1.p3.1),[4th item](https://arxiv.org/html/2606.27539#S2.I1.i4.p1.2),[§2\.5](https://arxiv.org/html/2606.27539#S2.SS5.p1.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[37\]K\. Lin, R\. K\. Lee, W\. Gao, and W\. Peng\(2021\)Early prediction of hate speech propagation\.In2021 International Conference on Data Mining Workshops \(ICDMW\),Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[38\]Y\. Liu, W\. Liu, X\. Gu, A\. He, W\. Wang, and Y\. Zhang\(2025\)PopSim: social network simulation for social media popularity prediction\.arXiv preprint arXiv:2512\.02533\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[39\]Y\. Liu, K\. Zeng, H\. Wang, X\. Song, and B\. Zhou\(2021\)Content matters: a gnn\-based model combined with text semantics for social network cascade prediction\.InPacific\-Asia Conference on Knowledge Discovery and Data Mining,Cited by:[Table 1](https://arxiv.org/html/2606.27539#S1.T1.2.1.13.1),[§1](https://arxiv.org/html/2606.27539#S1.p3.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[40\]D\. Mahajan, R\. Girshick, V\. Ramanathan, K\. He, M\. Paluri, Y\. Li, A\. Bharambe, and L\. Van Der Maaten\(2018\)Exploring the limits of weakly supervised pretraining\.InProceedings of the European conference on computer vision \(ECCV\),Cited by:[§A\.1\.5](https://arxiv.org/html/2606.27539#A1.SS1.SSS5.p1.1)\.
- \[41\]M\. Mazloom, R\. Rietveld, S\. Rudinac, M\. Worring, and W\. Van Dolen\(2016\)Multimodal popularity prediction of brand\-related social media posts\.InProceedings of the 24th ACM international conference on Multimedia,Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[42\]A\. N\. Medvedev, J\. Delvenne, and R\. Lambiotte\(2019\)Modelling structure and predicting dynamics of discussion threads in online boards\.Journal of Complex Networks\.Cited by:[Table 1](https://arxiv.org/html/2606.27539#S1.T1.2.1.8.1),[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§1](https://arxiv.org/html/2606.27539#S1.p3.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[43\]M\. Meghawat, S\. Yadav, D\. Mahata, Y\. Yin, R\. R\. Shah, and R\. Zimmermann\(2018\)A multimodal approach to predict social media popularity\.In2018 IEEE conference on multimedia information processing and retrieval \(MIPR\),Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[44\]M\. E\. Newman\(2001\)The structure of scientific collaboration networks\.Proceedings of the national academy of sciences\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[45\]H\. Pinto, J\. M\. Almeida, and M\. A\. Gonçalves\(2013\)Using early view patterns to predict the popularity of youtube videos\.InProceedings of the sixth ACM international conference on Web search and data mining,Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[46\]A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark,et al\.\(2021\)Learning transferable visual models from natural language supervision\.InInternational conference on machine learning,pp\. 8748–8763\.Cited by:[§3](https://arxiv.org/html/2606.27539#S3.p2.14)\.
- \[47\]T\. C\. Schelling\(1971\)Dynamic models of segregation\.Journal of mathematical sociology\.Cited by:[§4](https://arxiv.org/html/2606.27539#S4.p1.1)\.
- \[48\]R\. R\. Selvaraju, M\. Cogswell, A\. Das, R\. Vedantam, D\. Parikh, and D\. Batra\(2020\)Grad\-cam: visual explanations from deep networks via gradient\-based localization\.International journal of computer vision128\.Cited by:[Appendix H](https://arxiv.org/html/2606.27539#A8.p2.1)\.
- \[49\]C\. Shao, G\. L\. Ciampaglia, O\. Varol, K\. Yang, A\. Flammini, and F\. Menczer\(2018\)The spread of low\-credibility content by social bots\.Nature communications\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[50\]G\. Szabo and B\. A\. Huberman\(2010\)Predicting the popularity of online content\.Communications of the ACM\.Cited by:[Table 1](https://arxiv.org/html/2606.27539#S1.T1.2.1.5.1),[§1](https://arxiv.org/html/2606.27539#S1.p1.1),[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§1](https://arxiv.org/html/2606.27539#S1.p3.1),[2nd item](https://arxiv.org/html/2606.27539#S2.I1.i2.p1.2),[6th item](https://arxiv.org/html/2606.27539#S2.I1.i6.p1.1),[§2\.2](https://arxiv.org/html/2606.27539#S2.SS2.p1.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[51\]L\. Tang, Q\. Huang, A\. Puntambekar, Y\. Vigfusson, W\. Lloyd, and K\. Li\(2017\)Popularity prediction of facebook videos for higher quality streaming\.In2017 USENIX Annual Technical Conference \(USENIX ATC 17\),Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[52\]A\. Tatar, M\. D\. De Amorim, S\. Fdida, and P\. Antoniadis\(2014\)A survey on predicting the popularity of web content\.Journal of Internet Services and Applications\.Cited by:[§2\.3](https://arxiv.org/html/2606.27539#S2.SS3.p1.13)\.
- \[53\]O\. Tsur and A\. Rappoport\(2012\)What’s in a hashtag? content based prediction of the spread of ideas in microblogging communities\.InProceedings of the fifth ACM international conference on Web search and data mining,Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[54\]UNESCOSocial dynamics\.Note:https://www\.unesco\.org/en/tags/social\-dynamics\-0Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[55\]P\. Van Aelst, P\. Van Erkel, E\. D’heer, and R\. A\. Harder\(2017\)Who is leading the campaign charts? comparing individual popularity on old and new media\.Information, communication & society\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[56\]S\. Vosoughi, D\. Roy, and S\. Aral\(2018\)The spread of true and false news online\.science\.Cited by:[1st item](https://arxiv.org/html/2606.27539#S2.I1.i1.p1.1),[§2\.2](https://arxiv.org/html/2606.27539#S2.SS2.p1.1)\.
- \[57\]Y\. Wang, X\. Wang, Y\. Ran, R\. Michalski, and T\. Jia\(2022\)CasSeqGCN: combining network structure and temporal sequence to predict information cascades\.Expert Systems with Applications\.Cited by:[§B\.4](https://arxiv.org/html/2606.27539#A2.SS4.p1.2),[Table 1](https://arxiv.org/html/2606.27539#S1.T1.2.1.12.1),[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§1](https://arxiv.org/html/2606.27539#S1.p3.1),[§2\.5](https://arxiv.org/html/2606.27539#S2.SS5.p1.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[58\]D\. J\. Watts\(2002\)A simple model of global cascades on random networks\.Proceedings of the National Academy of Sciences99\.Cited by:[§4](https://arxiv.org/html/2606.27539#S4.p1.1)\.
- \[59\]B\. Wu, P\. Liu, W\. Cheng, B\. Liu, Z\. Zeng, J\. Wang, Q\. Huang, and J\. Luo\(2023\)Smp challenge: an overview and analysis of social media prediction challenge\.InProceedings of the 31st ACM International Conference on Multimedia,Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p3.1)\.
- \[60\]J\. Xie, Y\. Zhu, and Z\. Chen\(2021\)Micro\-video popularity prediction via multimodal variational information bottleneck\.IEEE Transactions on Multimedia\.Cited by:[Table 1](https://arxiv.org/html/2606.27539#S1.T1.2.1.16.1),[§1](https://arxiv.org/html/2606.27539#S1.p3.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[61\]Y\. Xu, J\. Wu, H\. Wan, Y\. Li, Z\. Hou, and M\. Kan\(2025\)Forecasting the buzz: enriching hashtag popularity prediction with llm reasoning\.InProceedings of the 34th ACM International Conference on Information and Knowledge Management,Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[62\]Y\. Xu, B\. Zheng, W\. Zhu, H\. Pan, Y\. Yao, N\. Xu, A\. Liu, Q\. Zhang, and C\. Yan\(2025\)SMTPD: a new benchmark for temporal prediction of social media popularity\.InProceedings of the Computer Vision and Pattern Recognition Conference,Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p3.1)\.
- \[63\]J\. Yang and S\. Counts\(2010\)Predicting the speed, scale, and range of information diffusion in twitter\.InProceedings of the International AAAI Conference on Web and Social Media,Cited by:[Table 1](https://arxiv.org/html/2606.27539#S1.T1.2.1.3.1)\.
- \[64\]Z\. Yang, Z\. Zhang, Z\. Zheng, Y\. Jiang, Z\. Gan, Z\. Wang, Z\. Ling, J\. Chen, M\. Ma, B\. Dong,et al\.\(2024\)Oasis: open agent social interaction simulations with one million agents\.arXiv preprint arXiv:2411\.11581\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§4](https://arxiv.org/html/2606.27539#S4.p1.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[65\]B\. Yu, M\. Chen, and L\. Kwok\(2011\)Toward predicting popularity of social marketing messages\.InInternational conference on social computing, behavioral\-cultural modeling, and prediction,Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[66\]V\. Zayats and M\. Ostendorf\(2018\)Conversation modeling on reddit using a graph\-structured lstm\.Transactions of the Association for Computational Linguistics6,pp\. 121–132\.Cited by:[§B\.5](https://arxiv.org/html/2606.27539#A2.SS5.p1.1),[Table 1](https://arxiv.org/html/2606.27539#S1.T1.2.1.14.1),[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§1](https://arxiv.org/html/2606.27539#S1.p3.1),[§2\.1](https://arxiv.org/html/2606.27539#S2.SS1.p1.17),[§2\.5](https://arxiv.org/html/2606.27539#S2.SS5.p1.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[67\]J\. Zhang, J\. Chang, C\. Danescu\-Niculescu\-Mizil, L\. Dixon, Y\. Hua, D\. Taraborelli, and N\. Thain\(2018\)Conversations gone awry: detecting early signs of conversational failure\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[68\]W\. Zhang, W\. Wang, J\. Wang, and H\. Zha\(2018\)User\-guided hierarchical attention network for multi\-modal social image popularity prediction\.InProceedings of the 2018 world wide web conference,Cited by:[Table 1](https://arxiv.org/html/2606.27539#S1.T1.2.1.15.1),[§1](https://arxiv.org/html/2606.27539#S1.p3.1),[§2\.1](https://arxiv.org/html/2606.27539#S2.SS1.p1.17),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[69\]Y\. Zhang, L\. Wang, J\. J\. Zhu, and X\. Wang\(2021\)Conspiracy vs science: a large\-scale analysis of online discussion cascades\.World wide web\.Cited by:[1st item](https://arxiv.org/html/2606.27539#S2.I1.i1.p1.1),[2nd item](https://arxiv.org/html/2606.27539#S2.I1.i2.p1.2),[4th item](https://arxiv.org/html/2606.27539#S2.I1.i4.p1.2),[5th item](https://arxiv.org/html/2606.27539#S2.I1.i5.p1.1),[§2\.2](https://arxiv.org/html/2606.27539#S2.SS2.p1.1)\.
- \[70\]Z\. Zhang, T\. Chen, Z\. Zhou, J\. Li, and J\. Luo\(2018\)How to become instagram famous: post popularity prediction with dual\-attention\.In2018 IEEE international conference on big data \(big data\),Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[71\]Z\. Zhang, S\. Xu, L\. Guo, and W\. Lian\(2022\)Multi\-modal variational auto\-encoder model for micro\-video popularity prediction\.InProceedings of the 8th International Conference on Communication and Information Processing,Cited by:[Table 1](https://arxiv.org/html/2606.27539#S1.T1.2.1.18.1),[§1](https://arxiv.org/html/2606.27539#S1.p3.1),[§2\.1](https://arxiv.org/html/2606.27539#S2.SS1.p1.17),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[72\]Q\. Zhao, M\. A\. Erdogdu, H\. Y\. He, A\. Rajaraman, and J\. Leskovec\(2015\)Seismic: a self\-exciting point process model for predicting tweet popularity\.InProceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining,Cited by:[Table 1](https://arxiv.org/html/2606.27539#S1.T1.2.1.7.1),[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§1](https://arxiv.org/html/2606.27539#S1.p3.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[73\]Z\. Zhao, P\. Resnick, and Q\. Mei\(2015\)Enquiring minds: early detection of rumors in social media from enquiry posts\.InProceedings of the 24th international conference on world wide web,Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1)\.
- \[74\]Y\. Zheng, C\. Gong, R\. Sun, J\. Zhang, L\. Pan, and L\. Lv\(2025\)Autocas: autoregressive cascade predictor in social networks via large language models\.arXiv preprint arXiv:2502\.18040\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[75\]T\. Zhong, J\. Lang, Y\. Zhang, Z\. Cheng, K\. Zhang, and F\. Zhou\(2024\)Predicting micro\-video popularity via multi\-modal retrieval augmentation\.InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval,Cited by:[Table 1](https://arxiv.org/html/2606.27539#S1.T1.2.1.17.1),[§1](https://arxiv.org/html/2606.27539#S1.p3.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[76\]F\. Zhou, X\. Xu, G\. Trajcevski, and K\. Zhang\(2021\)A survey of information cascade analysis: models, predictions, and recent advances\.ACM Computing Surveys \(CSUR\)\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p1.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.
- \[77\]A\. Zubiaga, M\. Liakata, R\. Procter, G\. Wong Sak Hoi, and P\. Tolmie\(2016\)Analysing how people orient to and spread rumours in social media by looking at conversational threads\.PloS one\.Cited by:[§1](https://arxiv.org/html/2606.27539#S1.p2.1),[§4](https://arxiv.org/html/2606.27539#S4.p2.1)\.

## Appendix

### Appendix ADataset Details

Here, we present the detailed information about the curation of the datasets and their statistics\.

#### A\.1Curation Details

##### A\.1\.1Cascade construction\.

For Bluesky, the metadata supports multiple interaction networks, including reply, repost, and quote networks\. We use the reply network because replies capture explicit conversational interactions and provide richer content signals than diffusion\-oriented actions such as reposts\. Posts are first grouped by their discussion thread identifier, and parent references are then used to add reply edges within each thread\. For Reddit, each submission defines a thread, and each comment provides a parent identifier indicating whether it replies to the root submission or to another comment\.

##### A\.1\.2Node attributes and thread\-level context\.

For both Bluesky and Reddit, each post is associated with textual content and timestamp information, which instantiateXvTextX\_\{v\}^\{\\text\{Text\}\}andXvTimeX\_\{v\}^\{\\text\{Time\}\}, respectively\. Platform\-specific engagement signals are also available at the post level\. Bluesky provides a like count for each post, whereas Reddit provides a karma score, computed as upvotes minus downvotes\. These engagement signals are available for individual posts, but our Like Score popularity target is defined only on the root postvrootv\_\{\\text\{root\}\}and is evaluated using its final engagement count\. For Bluesky, this translates to the number of likes that the social cascade initiating post receives, while on Reddit, it means the Karma score, which is defined as upvotes minus downvotes for the initial submission post\. Thread\-level contextXThreadX^\{\\text\{Thread\}\}is derived from metadata associated with the root post, including its timestamp\. For Bluesky, this additionally includes the follower count of the root post’s author, which provides a proxy for the initiating user’s social influence\.

##### A\.1\.3Modalities\.

The r/Gaming and r/Futurology data subsets include image content for root posts, although not every root post contains an image\. Specifically, 37\.9% of r/Gaming cascades and 68\.5% of r/Futurology cascades contain root\-post images\. In contrast, Bluesky and r/AMA are text\-based in our benchmark, soXvImageX\_\{v\}^\{\\text\{Image\}\}is empty for these datasets\. Since images are only available at the root\-post level in the Reddit multimodal subsets, visual features are included as part of the corresponding thread\-level context when present\.

##### A\.1\.4Missing and removed content\.

We discard posts whose parent post is missing\. If the root post of a cascade is missing but associated replies are present, we discard the entire cascade, sincevrootv\_\{\\text\{root\}\}is required to define the discussion tree\. For Reddit, some posts or comments may no longer contain their original textual content because they were deleted by users or removed by moderators at the time of collection and are instead represented by markers such as\[deleted\]or\[removed\]\. We retain such cascades when the remaining thread structure and metadata are intact, since these posts still correspond to observed participation in the cascade\. Although the original text is no longer accessible, the presence of deletion or removal markers may still provide information to the model about moderation or deletion patterns in the dataset communities\. Dataset sizes are computed after these filtering steps\.

##### A\.1\.5Dataset Sampling and Imbalance Mitigation\.

In the raw Bluesky dataset, cascade sizes are heavily skewed, with small conversation trees containing 3 to 10 posts comprising approximately 87% of the filtered data\. Training directly on this distribution would bias the model toward shallow dynamics and obscure structural patterns present in more complex cascades\. To address this imbalance, we employ square\-root sampling\[[40](https://arxiv.org/html/2606.27539#bib.bib31)\]\. Specifically, each size bin is sampled proportionally to the square root of its empirical frequency, yielding a more balanced subset of roughly 64,000 Bluesky discussion trees for robust model training and evaluation\.

#### A\.2Dataset Statistics\.

Table[4](https://arxiv.org/html/2606.27539#A1.T4)summarizes the dataset statistics for each benchmark subset after preprocessing and observation\-window filtering\. Each row corresponds to one dataset under one early observation window\. The observation window specifies how much of each cascade is visible to the model before prediction\. For example, a 2min window for Bluesky means that only the first 2 minutes of each cascade are observed, while a 15min window for r/AMA means that only the first 15 minutes are observed\.

TheDatasetcolumn indicates the benchmark subset, including Bluesky and the Reddit communities r/AMA, r/Gaming, and r/Futurology\. TheWindowcolumn gives the length of the early observation period\. TheCascadescolumn reports the number of cascades retained for that dataset and window\. Since cascades that have already completed before a given observation window are excluded, larger windows may retain fewer cascades\. Therefore, statistics are reported separately for each observation window\.

TheFinalcolumns describe the complete cascades after they have fully unfolded, up to the date of data collection\. These values represent the final cascade states used as prediction targets\. UnderFinal,Nodesreports the total number of nodes across all retained complete cascades,Avg\.reports the average number of final nodes per cascade, andMed\.reports the median number of final nodes per cascade\.

TheEarlycolumns describe the observed cascade prefixes within the specified observation window\. These values represent the information available to the model at prediction time\. UnderEarly,Nodesreports the total number of nodes observed during the early window across all retained cascades,Avg\.reports the average number of observed nodes per cascade, andMed\.reports the median number of observed nodes per cascade\.

Finally, theEarly % of Finalcolumn reports the fraction of the complete cascade that is visible within the observation window\. It is computed by comparing the total number of early observed nodes with the total number of final nodes for the same retained cascade set\. Larger values indicate that a greater portion of the final cascade is available to the model before prediction\. Overall, the table shows that longer observation windows provide more early cascade information, while often reducing the number of retained cascades because completed cascades are filtered out\.

Table 4:Dataset statistics across observation windows\. Final columns describe complete cascades retained for each window, while early columns describe the corresponding observed prefixesGtG^\{t\}\.DatasetWindowCascadesFinalEarlyEarly% of FinalNodesAvg\.Med\.NodesAvg\.Med\.Bluesky2min63,9041,198,63418\.767\.0107,9201\.621\.09\.0%10min60,3211,183,95619\.637\.0218,6192\.932\.018\.4%20min52,9721,144,61721\.618\.0305,9555\.763\.026\.7%r/AMA15min38,9821,500,18138\.515\.0156,1394\.03\.010\.4%30min37,6821,491,36539\.616\.0259,1856\.95\.017\.4%60min35,6771,471,68441\.317\.0392,44411\.07\.026\.7%r/Gaming20min13,3981,739,720129\.822\.070,9395\.34\.04\.1%50min12,8821,735,771134\.724\.0143,47511\.17\.08\.3%90min12,4131,730,833139\.425\.0237,26319\.19\.013\.7%r/Futurology30min9,7801,220,687124\.814\.035,0773\.62\.02\.9%90min8,7651,215,134138\.618\.095,95310\.94\.07\.9%180min8,3491,211,576145\.119\.0197,39823\.67\.016\.3%

### Appendix BBaseline Implementation Details

#### B\.1MLP\.

MLP is a structure\-agnostic baseline that represents each social cascade as a fixed vector rather than a reply tree\. For each observed cascade prefix, the model builds a cascade\-level representation by concatenating the initial root post’s feature vector, the mean feature vector over all observed posts, and the thread\-level context vector\. Each post feature vector contains a projected text embedding and two temporal features: time since the root post and time since the parent post\. The thread\-level context contains thread\-level metadata, including posting\-time and follower count\. The model passes the resulting cascade representation through a shared multilayer perceptron where it predicts all 6 popularity targets\. All targets are predicted in log\-transformed space\.

Hyperparameter Settings:The precomputed text embedding dimension is 384, and the projected text embedding dimension is 32\. The temporal features are log\-transformed and standardized using training\-set statistics\. The MLP has 2 layers of hidden dimensions 128, dropout is 0\.3, and the output dimension is 6, corresponding to the six popularity targets\. Training uses Adam with learning rate10−310^\{\-3\}, batch size 256, a maximum of 200 epochs and with early stopping patience 10\.

#### B\.2DeepHawkes\.

We implemented a DeepHawkes\-style baseline\[[9](https://arxiv.org/html/2606.27539#bib.bib1)\]for cascade prediction\. DeepHawkes represents an information cascade as a set of diffusion paths and learns path\-level representations that capture user influence, self\-excitation, and temporal decay in an end\-to\-end neural architecture\. In our setting, each MMG\-Pop training instance corresponds to one social cascade, represented by the observed root\-to\-node diffusion paths within the observation window\.

Following the DeepHawkes formulation, each path is encoded as a sequence of user embeddings and summarized with a recurrent encoder\. To incorporate temporal effects, we assign each path to one of 10 recency bins according to the time of its final post relative to the end of the observation window\. The model learns a positive scalar weight for each recency bin, and the cascade representation is obtained by a weighted aggregation of path representations\. This preserves the main DeepHawkes design while adapting it to the MMG\-Pop cascade format\.

Hyperparameter Settings:We use 50\-dimensional user embeddings and a GRU hidden size of 32\. The prediction MLP has hidden dimensions 32 and 16 with ReLU activations and dropout rate 0\.5\. The output dimension is 6, corresponding to the six prediction targets\. Models are trained with Adam using batch size 32, up to 200 epochs, early stopping patience 10, and weight decay10−410^\{\-4\}\. User embeddings use learning rate5×10−45\\times 10^\{\-4\}, while the remaining parameters use learning rate5×10−35\\times 10^\{\-3\}\.

#### B\.3DeepCas\.

We implement DeepCas\[[36](https://arxiv.org/html/2606.27539#bib.bib2)\]as a path\-based neural cascade encoder\. Each cascade is represented by random walks sampled from the observed early\-window conversation tree, using only nodes and edges available within the observation window\. Each walk starts at the initial post of the social cascade\. During sampling, the walker either moves to a child node or jumps to another observed node and where, both child transitions and jump targets are sampled using degree\-based weights\. If the current node is a leaf, the walker performs a jump using the same weighting rule\. The sampled user sequences are mapped to a pretrained user\-embedding vocabulary\. The embeddings are trained separately on a train\-split global interaction graph constructed from reply links across all cascades in the training set\. In this graph, each directed weighted edge connects a replying user to the author being replied to, and repeated reply interactions are aggregated as edge weights\. This graph is used to capture user\-user interactions across the dataset, rather than within a single cascade\. The pretrained embeddings are then loaded into DeepCas and kept fixed during supervised training\. Here, the model is adjusted to predict all 6 popularity targets as a 6 dimension vector\.

Hyperparameter Settings:We useKK= 200 sampled walks,TT= 10 steps per walk, pretrained interaction\-graph embeddings with dimension 128, GRU hidden dimension 128 per direction, random\-walk group size 5, two MLP hidden layers of size 128 with ReLU and dropout 0\.4, and a 6\-dimensional output layer\. Training uses Adam over trainable parameters only, learning rate10−310^\{\-3\}, weight decay5×10−45\\times 10^\{\-4\}batch size 256, maximum 200 epochs, early\-stopping patience 10, and gradient clipping with maximum norm 1\.0\.

#### B\.4CasSeqGCN\.

CasSeqGCN\[[57](https://arxiv.org/html/2606.27539#bib.bib3)\]represents each cascade as a temporal sequence of graph snapshots constructed from the observed early\-window conversation tree\. The original model combines structural and temporal cascade information by first encoding each snapshot with graph convolution, then aggregating node embeddings into a snapshot representation with dynamic routing, and finally processing the snapshot sequence with an LSTM before prediction\. Given this ordered sequence, snapshots are formed by adding posts in fixed increments ofQ=5Q=5\. We keep at mostKmax=15K\_\{\\max\}=15snapshots and force the final snapshot to include the last observed post if it is not already included by the fixed stride\. Thus, each cascade is represented as a sequence of up to 15 partial graph snapshots\. Here, the model is adjusted to predict all 6 popularity targets as a 6 dimension vector\.

Hyperparameter Settings:Node embedding dimension is 32, the snapshot embedding dimension is 32, the GCN hidden dimension is 32, the number of GCN layers is 2, the LSTM hidden dimension is 32, the number of LSTM layers is 2, the number of dynamic routing iterations is 3, and dropout is 0\.5\. For model selection, we search over learning rates\{0\.005,0\.01,0\.03,0\.05\}\\\{0\.005,0\.01,0\.03,0\.05\\\}separately for each dataset/window\. Each candidate is trained for at most 20 epochs with patience 5, and the candidate with the lowest validation loss is selected for full training\. Full training uses Adam with the selected learning rate, weight decay5×10−55\\times 10^\{\-5\}batch size 256, at most 200 epochs, early\-stopping patience 10, gradient clipping with maximum norm 1\.0\.

#### B\.5GraphLSTM\.

GraphLSTM\[[66](https://arxiv.org/html/2606.27539#bib.bib22)\]represents each cascade as a conversation tree\. Each post is a node, and each reply forms an edge from the parent post to the reply post\. For every node, the model builds an input vector by concatenating structural features with a mean\-pooled text embedding\. The structural features describe the node’s timing and position in the observed tree\. The text embedding summarizes the post text using learned token embeddings\. The model applies two graph\-LSTM passes over the tree\. The forward pass uses information from the parent node and the previous sibling node\. This gives each node a representation of the conversation context that came before it\. The backward pass uses information from the first child node and the next sibling node\. This gives each node a representation of the response context that follows it\. The forward and backward states are then concatenated to obtain a context\-aware representation for each node\. To obtain a cascade\-level representation, we mean\-pool the concatenated node states over all nodes in the tree\. Here, the model is adjusted to predict all 6 popularity targets as a 6 dimension vector\.

Hyperparameter Settings:The token embedding dimension is 100, the graph\-LSTM hidden dimension is 128, the final MLP hidden dimension is 128, dropout is 0\.3, the maximum post length is 100 tokens, the vocabulary minimum frequency is 10\. Training uses AdamW with learning rate10−310^\{\-3\}, embedding learning rate10−410^\{\-4\}, weight decay10−510^\{\-5\}, batch size 256, a maximum of 200 epochs, early stopping patience 10, and gradient clipping with maximum norm 1\.0\.

### Appendix CExperimental Setup

#### C\.1MMG\-PopNet Settings\.

MMG\-PopNet consists of three modality\-specific components and a prediction head\. We useall\-MiniLM\-L6\-v2as the text encoder,CLIP ViT\-B/32as the vision encoder, and GraphSAGE\[[30](https://arxiv.org/html/2606.27539#bib.bib69)\]as the graph message\-passing backbone\.

Hyperparameter Settings:The text encoderall\-MiniLM\-L6\-v2produces a 384\-dimensional representation\. This representation is mapped to a 32\-dimensional text embedding using a two\-layer projection network with hidden dimension 64 and dropout rate 0\.15\. All layers ofall\-MiniLM\-L6\-v2are updated during training\. The vision encoderCLIP ViT\-B/32produces a 512\-dimensional representation\. This representation is mapped to a 64\-dimensional image embedding using a linear projection layer with dropout rate 0\.3\. During training, only the final layer of the CLIP vision encoder is updated\. The graph component uses a 3\-layer GraphSAGE network with mean aggregation, hidden dimension 128, and dropout rate 0\.3\. The fused representation is passed to a two\-layer MLP prediction head with hidden dimension 128\. The prediction head outputs a 6\-dimensional vector, with one output corresponding to each popularity target\. We train MMG\-PopNet using Adam with separate learning rates for different parameter groups\. The learning rate is5×10−65\\times 10^\{\-6\}for the text encoder,10−610^\{\-6\}for the trainable CLIP vision parameters,10−310^\{\-3\}for the image projection parameters, and10−310^\{\-3\}for the GNN parameters\. The image projection parameters use weight decay10−210^\{\-2\}\. Training uses variable batch sizes with gradient accumulation to obtain an effective batch size of 256\. The maximum number of training epochs is 100, and early stopping is applied with patience 10\. We use mixed\-precision training and clip gradients to a maximum norm of 0\.5\. The learning\-rate schedule consists of linear warmup followed by cosine decay, with warmup ratio 0\.05 and minimum learning\-rate factor 0\.0\.

#### C\.2Completed Cascade Exclusion\.

Cascades that complete before the observation window are excluded from that setting to avoid trivial prediction cases\. For example, a cascade that ends after 8 minutes is excluded from the 10\-minute window because its observed prefix would already equal the final cascade\.

#### C\.3Training Split and Compute Resources\.

Dataset was divided into 80/10/10 splits of train, validation and test respectively\. For computation, 4×\\timesNvidia L40s GPUs were used to train the MMG\-PopNet to allow for handling the Out\-of\-Memory issue due to large text content associated with long social cascades\. To ensure fair comparison, all models had an effective batch size of 256\. Other representative baselines were trained on 1 L40s GPU\.

### Appendix DPopularity Prediction Across Future Horizons

#### D\.1Popularity prediction of final cascade state under different early observation\.

In addition to the MSE results reported in Table[2](https://arxiv.org/html/2606.27539#S5.T2), we reportR2R^\{2\}and Spearman rank correlation results for the same final cascade state prediction setting\. These two metrics offer complementary perspectives on model performance, whereR2R^\{2\}evaluates the exact predictive fit of the model’s estimates against true values, while Spearman correlation assesses the rank\-order agreement between predicted and actual outcomes\.

TheR2R^\{2\}metric \(coefficient of determination\) measures the proportion of variance in the final cascade states that can be explained by the early observation signals\. A higherR2R^\{2\}score indicates that a model’s numerical predictions tightly fit the actual popularity distributions\. As shown in Table[5](https://arxiv.org/html/2606.27539#A4.T5), MMG\-PopNet achieves the strongest overall performance, obtaining the highest average score for every target metric\. The gains are consistent across the Bluesky and Reddit datasets, indicating that the model successfully explains social cascade variance across platforms with diverse dynamics\. Furthermore, the advantage is maintained across all early observation windows\. In contrast, models relying on narrower temporal or structural signals, such as DeepHawkes and DeepCas, show unstable results withR2R^\{2\}values frequently approaching zero or dropping negative, meaning they fail to capture the variance better than simply predicting the mean\.

Conversely, the Spearman rank correlation evaluates how well a model preserves the relative ordering of cascades, independent of absolute numerical errors\. This is particularly important for downstream applications where the goal is to identify which threads will become the most viral or structurally complex\. Table[6](https://arxiv.org/html/2606.27539#A4.T6)demonstrates that MMG\-PopNet consistently achieves the highest overall average for every prediction task\. While baselines like Graph\-LSTM and MLP perform reasonably well at ranking tasks compared to theirR2R^\{2\}fit, they still fall short of MMG\-PopNet\. The strong Spearman results confirm that MMG\-PopNet’s multimodal design not only minimizes numerical error but reliably orders final cascade outcomes, making it highly effective for trend identification across heterogeneous platforms\.

Table 5:R2R^\{2\}results for final cascade state prediction under different early observation windows, includingstructural tasks\(max width, max depth, structural virality, size\),unique users, andlike score\. Higher is better\. Best values arebolded, and second\-best values areunderlined\.TaskModelBlueskyr/AMAr/Gamingr/FuturologyAvg021020Avg0153060Avg0205090Avg03090180AvgMax WidthMLP0\.3220\.3680\.4520\.4820\.4060\.0690\.2590\.3410\.3980\.2670\.0580\.3660\.5300\.6170\.393\-0\.0050\.2550\.5700\.6940\.3780\.361DeepHawkes\-0\.0120\.0870\.4040\.5000\.245\-0\.0010\.2230\.3440\.4700\.259\-0\.0130\.3200\.3850\.5470\.310\-0\.0420\.0870\.2690\.4750\.1970\.253DeepCas0\.0630\.0020\.0000\.0000\.016\-0\.180\-0\.062\-0\.025\-0\.030\-0\.074\-0\.267\-0\.0170\.0180\.087\-0\.0450\.0350\.0870\.2650\.2620\.1620\.015CasSeqGCN0\.0000\.2040\.4730\.5770\.3140\.0000\.2500\.4040\.5450\.3000\.0000\.3880\.5490\.6830\.405\-0\.0110\.3010\.5810\.7550\.4060\.356Graph\-LSTM0\.0290\.2220\.4900\.5840\.3310\.0380\.2850\.4250\.5570\.3260\.0500\.3800\.5650\.7200\.4290\.0280\.3610\.6110\.7690\.4420\.382MMG\-PopNet \(Ours\)0\.3250\.4540\.5800\.6540\.5030\.1240\.3120\.4270\.5980\.3650\.1640\.4120\.6180\.6960\.4720\.1210\.3710\.5900\.7510\.4580\.450Max DepthMLP0\.0220\.0490\.1880\.2580\.1290\.0310\.1860\.2990\.3860\.2260\.0080\.1940\.3870\.4440\.258\-0\.0170\.1700\.4340\.5370\.2810\.224DeepHawkes\-0\.0080\.0150\.0550\.1770\.060\-0\.0030\.1460\.2140\.2560\.1530\.0020\.1630\.2280\.3150\.177\-0\.0340\.0270\.1590\.3190\.1180\.127DeepCas0\.0090\.000\-0\.0000\.0010\.002\-0\.316\-0\.143\-0\.061\-0\.032\-0\.138\-0\.664\-0\.180\-0\.109\-0\.017\-0\.243\-0\.0230\.0500\.1970\.2150\.110\-0\.067CasSeqGCN0\.0000\.0450\.1930\.2930\.1330\.0000\.1430\.2950\.4020\.210\-0\.0010\.1460\.2940\.3730\.203\-0\.0060\.1820\.3800\.5170\.2680\.203Graph\-LSTM0\.0010\.0670\.2170\.3330\.1550\.0150\.1940\.3300\.4550\.248\-0\.0200\.1680\.4010\.5180\.2670\.0210\.2230\.4170\.5140\.2940\.241MMG\-PopNet \(Ours\)0\.0030\.1060\.2510\.3450\.1760\.0420\.2190\.3180\.4660\.2610\.0330\.2050\.4270\.5230\.2970\.1120\.2820\.4470\.5710\.3530\.272StructuralViralityMLP0\.0630\.0920\.2150\.2740\.1610\.0370\.2110\.3280\.4300\.252\-0\.0070\.1840\.3700\.4680\.254\-0\.0700\.1600\.4300\.5390\.2650\.233DeepHawkes\-0\.0070\.0210\.0870\.2410\.085\-0\.0070\.1700\.2360\.2830\.170\-0\.0010\.1110\.1120\.2520\.118\-0\.0270\.0020\.1260\.2440\.0860\.115DeepCas0\.016\-0\.000\-0\.0000\.0010\.004\-0\.632\-0\.257\-0\.099\-0\.078\-0\.267\-1\.780\-0\.613\-0\.490\-0\.260\-0\.786\-0\.091\-0\.0250\.0790\.1390\.026\-0\.256CasSeqGCN0\.0000\.0560\.2140\.3090\.145\-0\.0010\.1700\.3290\.4610\.2400\.0000\.0980\.2110\.2880\.149\-0\.0040\.1710\.3500\.5150\.2580\.198Graph\-LSTM0\.0010\.0750\.2240\.3400\.1600\.0160\.2090\.3590\.5050\.272\-0\.0790\.1270\.3610\.5210\.2320\.0010\.1950\.4070\.4980\.2750\.235MMG\-PopNet \(Ours\)0\.0150\.1420\.2730\.3600\.1980\.0230\.2510\.3430\.5080\.281\-0\.0250\.1510\.3890\.5130\.2570\.0670\.2510\.4370\.5640\.3300\.266SizeMLP0\.2110\.2600\.3690\.4150\.3140\.0630\.2740\.3800\.4550\.2930\.0390\.3340\.5290\.6160\.380\-0\.0090\.2310\.5600\.7170\.3750\.340DeepHawkes\-0\.0210\.0720\.2980\.4330\.196\-0\.0030\.2470\.3710\.4980\.278\-0\.0180\.2990\.3770\.5500\.302\-0\.0660\.0610\.2140\.4860\.1740\.237DeepCas0\.0550\.0020\.0000\.0010\.015\-0\.288\-0\.116\-0\.044\-0\.053\-0\.125\-0\.406\-0\.061\-0\.0210\.081\-0\.1020\.0180\.0760\.2700\.2880\.163\-0\.012CasSeqGCN0\.0000\.1470\.3740\.4780\.2500\.0000\.2500\.4070\.5580\.304\-0\.0010\.3330\.5160\.6570\.376\-0\.0130\.2600\.5360\.7420\.3810\.328Graph\-LSTM0\.0240\.1680\.3970\.4940\.2710\.0350\.2850\.4330\.5750\.3320\.0330\.3250\.5310\.6930\.3950\.0310\.3200\.5660\.7500\.4170\.354MMG\-PopNet \(Ours\)0\.1990\.3340\.4660\.5490\.3870\.1180\.3250\.4330\.6150\.3730\.1290\.3730\.6010\.6870\.4480\.1350\.3600\.5880\.7520\.4590\.416UniqueUsersMLP0\.3570\.4020\.4820\.5150\.4390\.0700\.2490\.3270\.3890\.2590\.0480\.3400\.5230\.6170\.382\-0\.0030\.2300\.5510\.7100\.3720\.363DeepHawkes\-0\.0160\.0870\.3920\.4870\.237\-0\.0010\.2090\.3220\.4590\.247\-0\.0150\.2960\.3680\.5360\.296\-0\.0440\.0640\.2380\.4920\.1880\.242DeepCas0\.0780\.0020\.0000\.0010\.020\-0\.231\-0\.086\-0\.034\-0\.048\-0\.100\-0\.329\-0\.021\-0\.0060\.090\-0\.0670\.0280\.0740\.2680\.2800\.1630\.004CasSeqGCN0\.0000\.1910\.4540\.5590\.3010\.0000\.2310\.3730\.5210\.2810\.0000\.3460\.5150\.6600\.380\-0\.0110\.2580\.5380\.7450\.3830\.336Graph\-LSTM0\.0310\.2250\.4910\.5900\.3340\.0390\.2720\.3990\.5320\.3100\.0400\.3390\.5340\.6970\.4020\.0280\.3250\.5680\.7410\.4150\.366MMG\-PopNet \(Ours\)0\.3540\.4700\.5820\.6480\.5130\.1190\.3000\.4000\.5760\.3490\.1440\.3840\.6020\.6900\.4550\.1290\.3560\.5840\.7510\.4550\.443Like ScoreMLP0\.4770\.4820\.5000\.5080\.4920\.0590\.1050\.1340\.1650\.1160\.0920\.0990\.2040\.3060\.1750\.0410\.0940\.3320\.4060\.2180\.250DeepHawkes\-0\.0100\.0430\.1980\.2600\.123\-0\.0020\.0170\.0570\.0730\.036\-0\.0060\.0130\.0480\.0870\.035\-0\.022\-0\.0210\.0060\.1580\.0300\.056DeepCas0\.0600\.0020\.0000\.0010\.016\-0\.0060\.0050\.006\-0\.0050\.000\-0\.0250\.0170\.0360\.0660\.0240\.1500\.1410\.2470\.2220\.1900\.057CasSeqGCN0\.0000\.0960\.2420\.3030\.160\-0\.0010\.0200\.0610\.0850\.0410\.0000\.0130\.0360\.1060\.039\-0\.0060\.0200\.1070\.2060\.0820\.081Graph\-LSTM0\.0160\.1280\.2860\.3590\.1970\.0520\.0870\.1040\.1340\.0940\.0800\.1010\.1550\.2750\.1530\.0740\.1990\.2860\.3030\.2150\.165MMG\-PopNet \(Ours\)0\.4970\.5660\.5910\.6210\.5690\.0870\.1560\.2180\.2810\.1850\.1830\.2500\.3130\.3930\.2850\.2580\.3180\.4710\.5640\.4030\.360

Description:TheR2R^\{2\}metric indicates the proportion of variance in the final popularity outcomes that the models successfully explain\. MMG\-PopNet consistently outperforms all baselines, achieving the highest average scores across structural, user, and engagement prediction targets\. Notably, traditional baselines like DeepCas and DeepHawkes struggle significantly, often yielding near\-zero or negative values, while MMG\-PopNet maintains a robust predictive fit regardless of the observation horizon\.

Table 6:Spearman results for final cascade state prediction under different early observation windows, includingstructural tasks\(max width, max depth, structural virality, size\),unique users, andlike score\. Higher is better\. Best values arebolded, and second\-best values areunderlined\.TaskModelBlueskyr/AMAr/Gamingr/FuturologyAvg021020Avg0153060Avg0205090Avg03090180AvgMax WidthMLP0\.5230\.5380\.5850\.6100\.5640\.2740\.5170\.6110\.6510\.5130\.2440\.5970\.7100\.7700\.5800\.1530\.5190\.7290\.7980\.5500\.552DeepHawkes0\.0300\.2380\.5240\.5530\.336\-0\.0060\.5070\.6360\.7280\.4660\.0040\.5520\.6900\.7730\.505\-0\.0350\.4090\.6150\.7310\.4300\.434DeepCas0\.0830\.1050\.0870\.0840\.0900\.0230\.1260\.1400\.1400\.1070\.1580\.2440\.2460\.3060\.2380\.2200\.3590\.5230\.5690\.4180\.213CasSeqGCN0\.0000\.3190\.5460\.6530\.3800\.0000\.5150\.6700\.7600\.4860\.0000\.6090\.7490\.8190\.5440\.0000\.5730\.7550\.8650\.5480\.490Graph\-LSTM0\.1630\.3350\.5530\.6460\.4240\.1970\.5360\.6770\.7710\.5450\.2210\.6050\.7810\.8460\.6130\.1920\.5980\.7740\.8540\.6050\.547MMG\-PopNet \(Ours\)0\.5340\.5890\.6670\.7140\.6260\.3660\.5780\.6810\.7790\.6010\.4300\.6520\.7830\.8380\.6760\.4100\.6340\.7710\.8550\.6670\.643Max DepthMLP0\.1300\.1800\.3670\.4670\.2860\.1720\.4280\.5530\.6220\.4440\.1210\.4490\.6230\.6620\.4640\.1220\.4130\.6480\.7180\.4750\.417DeepHawkes0\.0070\.1060\.2170\.3600\.172\-0\.0250\.3980\.5030\.5560\.3580\.0120\.4180\.5820\.6370\.412\-0\.0570\.3300\.5300\.6880\.3730\.329DeepCas0\.0380\.0370\.0440\.0590\.0440\.0570\.1260\.1350\.1390\.1140\.0810\.1610\.1640\.2300\.1590\.1960\.3040\.4640\.4910\.3640\.170CasSeqGCN0\.0000\.1760\.3530\.4920\.2550\.0000\.3880\.5420\.6260\.3890\.0000\.3860\.5560\.6110\.3880\.0000\.4020\.6130\.7050\.4300\.366Graph\-LSTM0\.1140\.2070\.4230\.5310\.3190\.1340\.4250\.5710\.6700\.4500\.1150\.4170\.6370\.7190\.4720\.1640\.4690\.6560\.7060\.4990\.435MMG\-PopNet \(Ours\)0\.1940\.2790\.4520\.5460\.3680\.2550\.4690\.5830\.6810\.4970\.2570\.4730\.6670\.7240\.5300\.3950\.5430\.6630\.7450\.5870\.495StructuralViralityMLP0\.2780\.3140\.4490\.5070\.3870\.1900\.4650\.5920\.6660\.4780\.1310\.4400\.6160\.6750\.4650\.1280\.4220\.6640\.7370\.4880\.455DeepHawkes0\.0130\.1630\.3460\.4780\.250\-0\.0260\.4390\.5560\.6210\.3980\.0010\.3950\.5500\.5900\.384\-0\.0510\.3460\.5350\.6830\.3780\.352DeepCas0\.0490\.0600\.0620\.0760\.0620\.0520\.1280\.1470\.1370\.1160\.0810\.1540\.1510\.2150\.1500\.2070\.3080\.4590\.4860\.3650\.173CasSeqGCN0\.0000\.2250\.4130\.5330\.2930\.0000\.4310\.5900\.6820\.4260\.0000\.3590\.5200\.5690\.3620\.0000\.4240\.6290\.7290\.4460\.382Graph\-LSTM0\.1280\.2360\.4450\.5390\.3370\.1490\.4560\.6090\.7120\.4810\.0880\.3890\.6260\.7260\.4570\.1780\.4820\.6740\.7150\.5120\.447MMG\-PopNet \(Ours\)0\.3030\.3720\.4960\.5740\.4360\.2760\.5080\.6180\.7210\.5310\.2390\.4760\.6680\.7330\.5290\.4050\.5380\.6840\.7610\.5970\.523SizeMLP0\.3930\.4210\.5180\.5660\.4750\.2530\.5270\.6410\.6980\.5300\.2070\.5840\.7210\.7750\.5720\.1370\.4950\.7320\.8140\.5440\.530DeepHawkes0\.0240\.2130\.4540\.5500\.310\-0\.0130\.5220\.6580\.7580\.4810\.0090\.5500\.7120\.7880\.515\-0\.0420\.3940\.6190\.7600\.4330\.435DeepCas0\.0710\.0840\.0790\.0880\.0800\.0410\.1410\.1570\.1510\.1220\.1510\.2370\.2320\.3070\.2320\.2220\.3440\.5240\.5730\.4160\.213CasSeqGCN0\.0000\.2800\.4770\.5880\.3360\.0000\.5120\.6630\.7630\.4850\.0000\.5750\.7410\.8050\.5300\.0000\.5250\.7310\.8470\.5260\.469Graph\-LSTM0\.1460\.2970\.5000\.5900\.3830\.1970\.5340\.6760\.7780\.5460\.1970\.5670\.7680\.8330\.5910\.1900\.5660\.7520\.8320\.5850\.526MMG\-PopNet \(Ours\)0\.4150\.4730\.5750\.6380\.5250\.3590\.5810\.6810\.7880\.6020\.4010\.6240\.7790\.8330\.6590\.4180\.6180\.7620\.8450\.6610\.612UniqueUsersMLP0\.5620\.5740\.6150\.6380\.5970\.2780\.5150\.6110\.6520\.5140\.2290\.5860\.7120\.7740\.5750\.1560\.4950\.7180\.8050\.5430\.557DeepHawkes0\.0420\.2220\.4990\.5170\.320\-0\.0090\.5100\.6350\.7280\.4660\.0060\.5460\.6920\.7730\.504\-0\.0330\.3900\.6020\.7510\.4270\.429DeepCas0\.0750\.1040\.0850\.0840\.0870\.0280\.1320\.1450\.1410\.1110\.1690\.2580\.2530\.3260\.2520\.2180\.3420\.5240\.5760\.4150\.216CasSeqGCN0\.0000\.2890\.5140\.6120\.3540\.0000\.5120\.6590\.7500\.4800\.0000\.5890\.7410\.8090\.5350\.0000\.5240\.7220\.8500\.5240\.473Graph\-LSTM0\.1720\.3440\.5600\.6460\.4300\.1990\.5300\.6680\.7620\.5400\.2090\.5840\.7740\.8400\.6020\.1890\.5660\.7510\.8240\.5820\.539MMG\-PopNet \(Ours\)0\.5690\.6160\.6750\.7070\.6420\.3700\.5740\.6690\.7670\.5950\.4220\.6410\.7800\.8360\.6700\.4220\.6190\.7580\.8430\.6600\.642Like ScoreMLP0\.6800\.6820\.6900\.6940\.6860\.2320\.2660\.2620\.2700\.2580\.2490\.2580\.3660\.4440\.3290\.2190\.2980\.5240\.5480\.3970\.418DeepHawkes0\.0570\.1650\.3760\.3720\.242\-0\.0140\.0500\.0950\.1100\.060\-0\.0200\.0100\.1120\.1840\.072\-0\.0020\.1510\.2130\.3170\.1700\.136DeepCas0\.0910\.1120\.0930\.0890\.0960\.0390\.0680\.1330\.1010\.0850\.0860\.1560\.1860\.2550\.1710\.4110\.4180\.4590\.4560\.4360\.197CasSeqGCN0\.0000\.2050\.3820\.4510\.2600\.0000\.0380\.0830\.1030\.0560\.0000\.0150\.0880\.1770\.0700\.0000\.1400\.2580\.3690\.1920\.144Graph\-LSTM0\.1270\.2750\.4410\.5200\.3410\.1890\.2070\.2080\.2460\.2120\.2980\.2920\.3430\.4330\.3420\.2800\.4620\.5430\.4450\.4330\.332MMG\-PopNet \(Ours\)0\.7110\.7500\.7650\.7760\.7500\.2510\.3080\.3470\.3560\.3160\.3970\.4320\.4890\.5250\.4610\.5320\.5750\.6890\.7260\.6300\.539

Description:The Spearman rank correlation evaluates the models’ ability to accurately preserve the relative ordinal ranking of cascades based on their final states\. MMG\-PopNet achieves the highest average correlation across all tasks and platforms, demonstrating superior capability in ranking future popularity trends\. While sequence\-based models like Graph\-LSTM provide competitive baseline rankings, MMG\-PopNet’s multimodal approach captures complementary signals that result in more reliable and consistent cascade orderings\.

#### D\.2Popularity Prediction of Cascade States at Future Horizon

We provide the full future\-horizon prediction results in Tables[7](https://arxiv.org/html/2606.27539#A4.T7),[8](https://arxiv.org/html/2606.27539#A4.T8),[9](https://arxiv.org/html/2606.27539#A4.T9), and[10](https://arxiv.org/html/2606.27539#A4.T10)\. These tables extend the main results by reporting all target variables across all datasets, observation windows, and prediction horizons\. In addition to the intermediate horizons\{4​h,8​h,16​h,24​h\}\\\{4\\text\{h\},8\\text\{h\},16\\text\{h\},24\\text\{h\}\\\}, we also report prediction error for the final cascade state\. Thus, in this setting, each observed prefixGtG^\{t\}can produce multiple supervised instances, where each instance pairs the same prefix with a different target stateYGt′\\textbf\{Y\}\_\{G\}^\{t^\{\\prime\}\}\. To specify which future state is being predicted, we append a one\-hot encoding of the target horizon to the input representation\. This training setup differs from terminal\-state\-only prediction because each model receives supervision from both intermediate cascade states and the final state\. The learned representation is therefore shaped by multiple stages of cascade evolution rather than only by the terminal outcome\.

If a cascade has already terminated before a given horizon, the corresponding intermediate target is unavailable and is excluded from that horizon\-specific training and evaluation set\. The final\-state target differs from the fixed intermediate horizons because it is defined for every completed cascade\. For this reason, the final\-state MSE can be lower than the error at longer horizons such as 16h or 24h\. Some cascades terminate before those later horizons, so their final state occurs earlier than the fixed horizon and is easier to infer from the observed prefix\. This effect should be considered when comparing the final\-state column with the 16h and 24h columns\.

For engagement targetLike Score, ground truth is available only at the final cascade state\. Consequently, those targets are evaluated only in the Final column and are marked as unavailable at intermediate horizons\.

Across the four datasets, MMG\-PopNet achieves the lowest MSE in most settings across the popularity targets\. It consistently outperforms the structure\-only baselines such as DeepCas and DeepHawkes, and this underscores the necessity of multimodal representations for capturing both the semantic and structural dynamics that influence a social cascade\. A notable pattern is that MLP is often competitive for intermediate horizons and sometimes obtains the best result\. This is particularly visible under root\-only inputs and shorter forecasting horizons, where a fixed cascade\-level representation can capture much of the available predictive signal\.

The MLP baseline uses the root\-post features, the mean representation of observed posts, and thread\-level metadata, so its strong performance indicates that early content, aggregate response features, posting context, and author\-level context are highly informative for near\-term cascade growth\. This result also shows that complex structural encoders are not always necessary when the target horizon is close to the observation window or when little cascade topology is available\.

Graph\-LSTM and CasSeqGCN form the strongest structure\-aware baselines in many settings\. They are especially competitive when longer observation windows expose more of the reply tree, and their strongest results occur forMax DepthandStructural Virality\.

Overall, the appendix results support the trends reported in the main section while adding a more detailed view across targets\.

Table 7:Bluesky future\-horizon prediction MSE across observation windows and popularity targets\.Lower is better\. Best values areboldedand second\-best values areunderlined\. Column groups correspond to Root Only, 2 min, 10 min, and 20 min observation windows, and columns within each group report prediction error at\{4​h,8​h,16​h,24​h\}\\\{4\\text\{h\},8\\text\{h\},16\\text\{h\},24\\text\{h\}\\\}and at the final cascade state\.Like Scoreis reported only for the final state because its intermediate\-horizon labels are unavailable\.TaskModelRoot Only2 min10 min20 min4h8h16h24hFinal4h8h16h24hFinal4h8h16h24hFinal4h8h16h24hFinalMax WidthMLP0\.4720\.5700\.6940\.8000\.4580\.4100\.4990\.6200\.7210\.4290\.3180\.3980\.4970\.5840\.3730\.2910\.3640\.4540\.5320\.348DeepHawkes0\.8090\.9991\.2621\.4440\.7280\.6830\.8561\.0871\.2700\.6340\.3690\.4940\.6610\.7860\.4360\.2610\.3470\.4620\.5490\.330DeepCas0\.7070\.8641\.0511\.1980\.6280\.7870\.9681\.1751\.3200\.6750\.7900\.9671\.1741\.3230\.6760\.7900\.9681\.1731\.3200\.677CasSeqGCN0\.7890\.9691\.1731\.3170\.6760\.5380\.6660\.8310\.9740\.5420\.2890\.3900\.5080\.6190\.3600\.2110\.2960\.3980\.4930\.292Graph\-LSTM0\.7510\.9051\.0831\.2170\.6600\.5170\.6320\.7780\.8860\.5230\.2840\.3820\.4980\.5990\.3540\.2120\.2930\.3860\.4740\.288MMG\-PopNet0\.4720\.5590\.6670\.7620\.4480\.3830\.4470\.5320\.6120\.3860\.2570\.3200\.4040\.4780\.2950\.1990\.2560\.3270\.3940\.243Max DepthMLP0\.4980\.5040\.5580\.5650\.3560\.4500\.4590\.5100\.5270\.3470\.3480\.3580\.4100\.4280\.2950\.3010\.3230\.3800\.4000\.268DeepHawkes0\.5960\.6120\.6790\.6890\.4030\.5600\.5560\.6150\.6260\.3840\.4280\.4720\.5010\.5620\.3400\.4040\.4140\.4670\.4830\.334DeepCas0\.5580\.5750\.6380\.6460\.3620\.5750\.5950\.6620\.6710\.3710\.5780\.5940\.6630\.6730\.3640\.5760\.5940\.6630\.6730\.364CasSeqGCN0\.5760\.5940\.6630\.6720\.3650\.4910\.5040\.5660\.5830\.3510\.3550\.3680\.4240\.4440\.3000\.3000\.3170\.3730\.3910\.261Graph\-LSTM0\.5540\.5650\.6220\.6270\.3640\.4700\.4780\.5320\.5450\.3410\.3370\.3500\.4000\.4180\.2790\.2870\.3060\.3590\.3750\.246MMG\-PopNet0\.4850\.4930\.5410\.5420\.3790\.4210\.4210\.4600\.4680\.3290\.3280\.3340\.3740\.3860\.2790\.2780\.2910\.3360\.3460\.240StructuralViralityMLP0\.2590\.2570\.2820\.2840\.1470\.2320\.2320\.2560\.2640\.1430\.1810\.1840\.2100\.2190\.1230\.1570\.1660\.1930\.2060\.113DeepHawkes0\.3240\.3240\.3550\.3570\.1740\.3020\.2960\.3260\.3300\.1700\.2240\.2440\.2570\.2910\.1410\.2100\.2110\.2380\.2470\.135DeepCas0\.3070\.3080\.3370\.3380\.1550\.3170\.3200\.3490\.3490\.1590\.3180\.3190\.3490\.3500\.1570\.3180\.3190\.3490\.3500\.157CasSeqGCN0\.3170\.3190\.3490\.3490\.1570\.2650\.2660\.2950\.3030\.1500\.1890\.1950\.2230\.2360\.1260\.1600\.1680\.1950\.2080\.111Graph\-LSTM0\.3060\.3040\.3280\.3270\.1580\.2560\.2560\.2800\.2870\.1460\.1820\.1870\.2120\.2220\.1180\.1520\.1630\.1890\.1990\.105MMG\-PopNet0\.2570\.2550\.2760\.2740\.1610\.2200\.2160\.2340\.2400\.1370\.1720\.1740\.1930\.2010\.1160\.1470\.1530\.1750\.1830\.102SizeMLP0\.8420\.9571\.1231\.2380\.6940\.7260\.8330\.9921\.1130\.6550\.5420\.6390\.7780\.8850\.5570\.4750\.5770\.7110\.8170\.512DeepHawkes1\.2621\.4901\.8162\.0070\.9671\.0841\.2611\.5451\.7400\.8520\.6300\.8020\.9911\.1690\.6260\.5200\.6190\.7820\.8940\.518DeepCas1\.1291\.3151\.5611\.7180\.8301\.2321\.4471\.7171\.8700\.8781\.2361\.4481\.7171\.8760\.8811\.2361\.4481\.7171\.8740\.881CasSeqGCN1\.2371\.4481\.7171\.8740\.8810\.8901\.0411\.2531\.4200\.7520\.5290\.6520\.8190\.9580\.5520\.4080\.5190\.6680\.8000\.463Graph\-LSTM1\.1731\.3481\.5681\.7050\.8650\.8570\.9901\.1751\.3020\.7310\.5090\.6290\.7840\.9100\.5320\.4040\.5120\.6480\.7610\.454MMG\-PopNet0\.8290\.9361\.0751\.1690\.7040\.6900\.7620\.8650\.9570\.6050\.4820\.5580\.6690\.7600\.4830\.3870\.4640\.5670\.6570\.408UniqueUsersMLP0\.5110\.6060\.7350\.8430\.4640\.4450\.5310\.6600\.7620\.4350\.3450\.4250\.5280\.6210\.3760\.3100\.3840\.4770\.5620\.348DeepHawkes0\.9071\.1191\.4171\.6200\.7820\.7610\.9441\.2031\.4000\.6780\.4240\.5600\.7460\.8870\.4670\.3120\.4050\.5370\.6370\.361DeepCas0\.7800\.9481\.1561\.3200\.6600\.8861\.0821\.3191\.4840\.7210\.8901\.0821\.3191\.4890\.7230\.8921\.0841\.3181\.4860\.724CasSeqGCN0\.8901\.0851\.3191\.4860\.7230\.6150\.7540\.9411\.0970\.5910\.3390\.4460\.5790\.6990\.3970\.2520\.3410\.4550\.5640\.324Graph\-LSTM0\.8451\.0081\.2081\.3550\.7040\.5860\.7090\.8740\.9910\.5580\.3260\.4300\.5570\.6680\.3750\.2450\.3300\.4340\.5320\.305MMG\-PopNet0\.5100\.5940\.7030\.7980\.4540\.4320\.4950\.5830\.6650\.3970\.3010\.3660\.4550\.5370\.3130\.2350\.2930\.3670\.4440\.258LikeScoreMLP––––1\.312––––1\.299––––1\.260––––1\.225DeepHawkes––––2\.533––––2\.398––––2\.114––––1\.898DeepCas––––2\.347––––2\.502––––2\.510––––2\.514CasSeqGCN––––2\.509––––2\.280––––1\.904––––1\.756Graph\-LSTM––––2\.469––––2\.205––––1\.839––––1\.649MMG\-PopNet––––1\.294––––1\.140––––1\.069––––1\.006

Description:The results demonstrate that MMG\-PopNet consistently achieves the lowest prediction error across all targets\. The data reveals a steep drop in MSE as the observation window expands from 0 minutes \("Root Only"\) to 20 minutes, highlighting the value of early engagement data\. Notably, while baseline models like DeepHawkes and DeepCas struggle significantly where they frequently exceed an MSE of 1\.0 on taregtsSizeandUnique Users\. Here, MMG\-PopNet maintains a substantial advantage over these targets\.

Table 8:r/AMA future\-horizon prediction MSE across observation windows and popularity targets\.Lower is better\. Best values areboldedand second\-best values areunderlined\. Column groups correspond to Root Only, 15 min, 30 min, and 60 min observation windows, and columns within each group report prediction error at\{4​h,8​h,16​h,24​h\}\\\{4\\text\{h\},8\\text\{h\},16\\text\{h\},24\\text\{h\}\\\}and at the final cascade state\.Like Scoreis reported only for the final state because its intermediate\-horizon labels are unavailable\.TaskModelRoot Only15 min30 min60 min4h8h16h24hFinal4h8h16h24hFinal4h8h16h24hFinal4h8h16h24hFinalMax WidthMLP0\.5450\.6470\.8070\.9600\.6700\.3830\.4780\.6220\.7430\.5330\.3240\.4250\.5650\.6870\.4930\.2640\.3430\.4630\.5520\.407DeepHawkes0\.6090\.7190\.9081\.0790\.7170\.4190\.5110\.6660\.8030\.5460\.3660\.4640\.6310\.7770\.5140\.2070\.2950\.4280\.5530\.361DeepCas0\.7900\.9001\.0661\.2280\.8980\.6270\.7240\.8871\.0680\.7440\.6600\.7841\.0071\.1390\.7960\.5810\.6920\.8621\.0010\.707CasSeqGCN0\.6060\.7140\.8911\.0570\.7140\.3930\.4920\.6550\.7980\.5280\.2730\.3860\.5330\.6520\.4470\.1550\.2360\.3460\.4360\.307Graph\-LSTM0\.5740\.6710\.8360\.9840\.6810\.3590\.4670\.6260\.7560\.5070\.2600\.3750\.5240\.6570\.4380\.1730\.2590\.3730\.4700\.316MMG\-PopNet0\.5360\.6300\.7870\.9470\.6630\.3480\.4400\.5790\.6940\.4910\.2590\.3590\.4890\.5860\.4260\.1530\.2230\.3180\.3920\.285Max DepthMLP0\.3710\.3640\.3640\.3610\.3210\.2800\.2860\.2890\.2940\.2680\.2230\.2240\.2250\.2340\.2280\.1610\.1790\.1870\.1980\.199DeepHawkes0\.3940\.3880\.3880\.3900\.3340\.3260\.3210\.3160\.3210\.2870\.3040\.2820\.2750\.2930\.2740\.2120\.2230\.2270\.2420\.231DeepCas0\.5560\.5350\.5230\.5110\.5140\.4200\.4010\.3890\.3920\.3650\.4360\.4100\.4090\.4000\.3720\.3780\.3750\.3610\.3660\.339CasSeqGCN0\.3940\.3880\.3840\.3840\.3280\.2990\.2980\.2990\.3090\.2730\.2240\.2260\.2300\.2430\.2240\.1510\.1680\.1750\.1880\.187Graph\-LSTM0\.3810\.3760\.3680\.3660\.3230\.2780\.2850\.2890\.2950\.2630\.2210\.2250\.2280\.2500\.2230\.1480\.1660\.1750\.1920\.186MMG\-PopNet0\.3780\.3690\.3550\.3510\.3300\.2680\.2750\.2750\.2840\.2620\.2070\.2070\.2100\.2180\.2150\.1330\.1490\.1560\.1700\.177StructuralViralityMLP0\.1710\.1580\.1510\.1410\.1300\.1250\.1210\.1180\.1140\.1060\.1010\.0950\.0900\.0930\.0900\.0710\.0750\.0730\.0740\.075DeepHawkes0\.1840\.1700\.1620\.1540\.1370\.1500\.1390\.1300\.1250\.1140\.1420\.1210\.1140\.1200\.1110\.0910\.0920\.0890\.0940\.089DeepCas0\.2910\.2690\.2480\.2290\.2470\.2150\.1900\.1700\.1620\.1640\.2370\.2080\.1980\.1800\.1810\.1910\.1740\.1560\.1500\.146CasSeqGCN0\.1840\.1700\.1600\.1500\.1340\.1370\.1280\.1220\.1190\.1080\.1010\.0940\.0920\.0980\.0880\.0650\.0690\.0670\.0720\.069Graph\-LSTM0\.1770\.1640\.1530\.1440\.1320\.1260\.1220\.1200\.1150\.1050\.1040\.0980\.0940\.1040\.0910\.0630\.0680\.0660\.0720\.068MMG\-PopNet0\.1790\.1650\.1470\.1380\.1390\.1210\.1170\.1120\.1100\.1050\.0940\.0870\.0850\.0880\.0870\.0600\.0630\.0610\.0650\.066SizeMLP0\.8380\.9341\.1101\.2650\.9230\.5530\.6610\.8270\.9600\.7140\.4430\.5510\.6870\.8280\.6310\.3310\.4330\.5590\.6610\.511DeepHawkes0\.9401\.0511\.2541\.4340\.9870\.6390\.7320\.8961\.0490\.7390\.5530\.6270\.7850\.9630\.6700\.2730\.3840\.5340\.6960\.456DeepCas1\.3981\.4721\.6341\.7921\.4600\.9961\.0731\.2361\.4351\.0571\.0891\.1891\.4111\.5381\.1390\.9341\.0511\.2181\.3741\.011CasSeqGCN0\.9371\.0451\.2351\.4070\.9800\.6070\.7110\.8881\.0490\.7260\.4180\.5400\.6910\.8350\.5960\.2260\.3310\.4470\.5600\.413Graph\-LSTM0\.8870\.9851\.1491\.3000\.9420\.5440\.6660\.8420\.9810\.6930\.4090\.5390\.6860\.8600\.5950\.2610\.3670\.4870\.6050\.436MMG\-PopNet0\.8320\.9191\.0751\.2420\.9250\.5170\.6170\.7680\.8970\.6680\.3830\.4870\.6180\.7320\.5650\.2090\.2980\.3980\.4900\.378UniqueUsersMLP0\.5290\.6610\.8571\.0370\.6630\.3720\.4950\.6690\.8170\.5320\.3170\.4400\.6120\.7670\.5020\.2510\.3510\.4970\.6000\.410DeepHawkes0\.5900\.7370\.9661\.1650\.7120\.4110\.5310\.7180\.8870\.5460\.3580\.4840\.6870\.8700\.5290\.1980\.3030\.4620\.6030\.370DeepCas0\.8000\.9371\.1291\.3060\.9370\.6120\.7390\.9381\.1400\.7470\.6580\.8111\.0721\.2380\.8180\.5540\.6940\.8961\.0570\.714CasSeqGCN0\.5860\.7310\.9511\.1450\.7060\.3960\.5270\.7230\.8970\.5370\.2900\.4250\.6060\.7650\.4700\.1610\.2570\.3890\.4990\.321Graph\-LSTM0\.5560\.6870\.8861\.0590\.6770\.3520\.4880\.6770\.8340\.5100\.2730\.4100\.5890\.7590\.4600\.1840\.2860\.4260\.5390\.338MMG\-PopNet0\.5230\.6500\.8391\.0270\.6650\.3440\.4590\.6240\.7650\.4940\.2670\.3840\.5410\.6710\.4410\.1600\.2410\.3560\.4450\.298LikeScoreMLP––––1\.367––––1\.316––––1\.283––––1\.204DeepHawkes––––1\.441––––1\.406––––1\.430––––1\.338DeepCas––––1\.468––––1\.422––––1\.497––––1\.419CasSeqGCN––––1\.436––––1\.405––––1\.403––––1\.305Graph\-LSTM––––1\.359––––1\.350––––1\.389––––1\.281MMG\-PopNet––––1\.403––––1\.281––––1\.239––––1\.086

Description:While MMG\-PopNet establishes dominance in long horizon predictions and larger observation windows, the data reveals a unique trend in the "Root Only" setting where the simpler MLP model frequently matches or outperforms complex graph models in predictingMax DepthandStructural Virality\.

Table 9:r/Gaming future\-horizon prediction MSE across observation windows and popularity targets\.Lower is better\. Best values areboldedand second\-best values areunderlined\. Column groups correspond to Root Only, 20 min, 50 min, and 90 min observation windows, and columns within each group report prediction error at\{4​h,8​h,16​h,24​h\}\\\{4\\text\{h\},8\\text\{h\},16\\text\{h\},24\\text\{h\}\\\}and at the final cascade state\.Like Scoreis reported only for the final state because its intermediate\-horizon labels are unavailable\.TaskModelRoot Only20 min50 min90 min4h8h16h24hFinal4h8h16h24hFinal4h8h16h24hFinal4h8h16h24hFinalMax WidthMLP1\.3011\.6231\.8822\.0671\.7780\.7010\.9521\.1761\.3001\.1740\.5060\.7150\.8680\.9700\.8790\.4060\.5670\.6930\.7830\.715DeepHawkes1\.4011\.7572\.1172\.3931\.9190\.9331\.1111\.3681\.5521\.2950\.5940\.9741\.2791\.5801\.2560\.3540\.5730\.8411\.0730\.866DeepCas1\.9962\.3172\.7222\.8982\.3531\.3671\.6691\.9752\.1471\.8151\.3361\.6521\.9001\.9401\.7861\.2501\.5381\.7571\.9131\.700CasSeqGCN1\.3931\.7131\.9922\.1641\.8680\.6640\.9311\.1861\.3451\.1350\.3880\.6490\.8700\.9710\.8550\.2250\.3930\.5720\.6650\.590Graph\-LSTM1\.3451\.6511\.9412\.1011\.8150\.6220\.8931\.1371\.2851\.1040\.3340\.5890\.7850\.8970\.7600\.2110\.4110\.6190\.7380\.587MMG\-PopNet1\.1601\.4421\.6871\.8181\.5540\.7180\.9401\.1161\.2131\.1030\.3430\.5110\.6590\.7330\.6620\.2140\.3480\.4940\.5750\.490Max DepthMLP0\.2950\.3200\.3420\.3660\.3420\.2150\.2380\.2730\.2960\.2800\.1570\.1750\.1880\.1920\.2080\.1260\.1500\.1680\.1730\.176DeepHawkes0\.3170\.3320\.3660\.4090\.3540\.2750\.2820\.3150\.3460\.3000\.1980\.2320\.2640\.3050\.2790\.1990\.2100\.2280\.2430\.234DeepCas0\.5770\.5640\.6050\.6000\.5890\.3370\.3490\.3800\.3970\.3710\.3380\.3460\.3610\.3530\.3750\.3130\.3230\.3340\.3260\.323CasSeqGCN0\.3090\.3290\.3550\.3810\.3470\.2370\.2620\.2970\.3240\.2950\.1870\.2120\.2340\.2480\.2390\.1720\.1920\.2090\.2100\.208Graph\-LSTM0\.3070\.3280\.3570\.3840\.3510\.2170\.2430\.2820\.3100\.2810\.1520\.1810\.2010\.2130\.2160\.1010\.1350\.1580\.1720\.166MMG\-PopNet0\.2960\.3100\.3290\.3460\.3360\.2250\.2420\.2610\.2780\.2730\.1420\.1630\.1790\.1890\.1980\.1120\.1300\.1460\.1510\.155StructuralViralityMLP0\.1180\.1130\.1060\.1090\.0940\.0840\.0850\.0880\.0930\.0780\.0540\.0540\.0590\.0560\.0570\.0480\.0490\.0500\.0510\.047DeepHawkes0\.1270\.1180\.1130\.1190\.0980\.1130\.1060\.1060\.1140\.0900\.0780\.0780\.0840\.0870\.0800\.0920\.0810\.0790\.0840\.071DeepCas0\.3510\.3070\.2800\.2480\.2820\.1570\.1370\.1280\.1290\.1180\.1460\.1240\.1200\.1060\.1200\.1450\.1210\.1090\.1050\.099CasSeqGCN0\.1240\.1170\.1100\.1140\.0950\.0950\.0960\.0990\.1060\.0860\.0660\.0690\.0750\.0750\.0710\.0720\.0710\.0740\.0780\.066Graph\-LSTM0\.1250\.1190\.1160\.1200\.1000\.0920\.0910\.0930\.1000\.0820\.0630\.0600\.0620\.0610\.0620\.0430\.0480\.0500\.0530\.047MMG\-PopNet0\.1260\.1190\.1090\.1100\.1000\.0980\.0930\.0880\.0900\.0790\.0610\.0580\.0580\.0550\.0590\.0580\.0500\.0460\.0450\.044SizeMLP1\.4941\.8942\.1742\.3842\.0340\.8251\.1431\.4171\.5701\.3950\.5450\.7920\.9901\.1090\.9900\.4220\.6310\.7910\.8900\.799DeepHawkes1\.6082\.1012\.4882\.7932\.2461\.0631\.3381\.6681\.9011\.5110\.6751\.1371\.5191\.8901\.4570\.4170\.6961\.0231\.3031\.015DeepCas2\.5792\.9493\.3763\.5752\.9581\.5851\.9492\.2842\.4772\.0701\.5921\.9442\.2262\.2462\.0641\.4361\.7742\.0192\.1591\.906CasSeqGCN1\.5891\.9752\.2742\.4762\.1030\.8401\.1861\.4911\.6841\.3930\.4720\.7931\.0691\.1951\.0100\.2750\.5000\.7270\.8420\.717Graph\-LSTM1\.5461\.9222\.2382\.4212\.0670\.7841\.1361\.4301\.6181\.3580\.4720\.7921\.0151\.1620\.9720\.2650\.5240\.7790\.9240\.725MMG\-PopNet1\.3801\.7151\.9702\.1261\.8280\.8731\.1481\.3361\.4501\.3220\.4050\.6090\.7930\.8870\.7910\.2500\.4150\.5900\.6780\.580UniqueUsersMLP1\.3971\.7962\.1032\.3191\.9050\.7761\.0921\.3641\.5131\.2960\.5310\.7800\.9771\.1000\.9490\.4080\.6130\.7740\.8740\.760DeepHawkes1\.4941\.9272\.3342\.6542\.0481\.0391\.2821\.5931\.8091\.4280\.6651\.1261\.5011\.8561\.4130\.3970\.6781\.0031\.2690\.974DeepCas2\.2462\.6333\.0743\.2652\.6231\.4551\.8262\.1822\.3811\.9261\.4511\.8252\.1222\.1791\.9281\.3231\.6661\.9322\.0881\.793CasSeqGCN1\.4801\.8762\.2012\.4071\.9830\.7801\.1201\.4281\.6201\.2890\.4600\.7801\.0591\.1910\.9720\.2570\.4750\.6980\.8100\.672Graph\-LSTM1\.4421\.8232\.1592\.3451\.9390\.7251\.0691\.3671\.5491\.2530\.4310\.7400\.9681\.1200\.8910\.2640\.5130\.7610\.8980\.684MMG\-PopNet1\.2711\.6081\.8932\.0491\.6900\.8001\.0741\.2791\.3911\.2110\.3810\.5820\.7660\.8570\.7370\.2320\.3940\.5670\.6460\.539LikeScoreMLP––––5\.856––––5\.410––––4\.883––––4\.614DeepHawkes––––6\.268––––6\.068––––6\.497––––6\.629DeepCas––––6\.466––––6\.291––––6\.253––––6\.673CasSeqGCN––––6\.208––––6\.124––––5\.992––––6\.048Graph\-LSTM––––5\.796––––6\.032––––5\.647––––5\.505MMG\-PopNet––––5\.180––––4\.819––––4\.325––––4\.073

Description:The results show that while MMG\-PopNet remains the most robust model for predicting “Final” outcomes, Graph\-LSTM and MLP are highly competitive and occasionally superior at forecasting short\-term 4\-hour horizons\. The table also underscores the extreme difficulty of predicting theLike Scorein gaming communities, with all models exhibiting massive error rates \(MSE \> 4\.0\)\. However, feeding the models 90 minutes of initial social cascade data reduces the error by over 60% compared to Root Only predictions\.

Table 10:r/Futurology future\-horizon prediction MSE across observation windows and popularity targets\.Lower is better\. Best values areboldedand second\-best values areunderlined\. Column groups correspond to Root Only, 30 min, 90 min, and 180 min observation windows, and columns within each group report prediction error at\{4​h,8​h,16​h,24​h\}\\\{4\\text\{h\},8\\text\{h\},16\\text\{h\},24\\text\{h\}\\\}and at the final cascade state\.Like Scoreis reported only for the final state because its intermediate\-horizon labels are unavailable\.MetricModelRoot Only30 min90 min180 min4h8h16h24hFinal4h8h16h24hFinal4h8h16h24hFinal4h8h16h24hFinalMax WidthMLP1\.1001\.5221\.8791\.9681\.7610\.6401\.0291\.4041\.4451\.3120\.3010\.5060\.7490\.8120\.7060\.2460\.3290\.4610\.5260\.454DeepHawkes1\.1411\.6362\.1182\.2491\.8940\.8131\.2601\.8191\.9691\.5560\.6320\.9571\.3941\.4681\.2170\.5570\.7301\.0161\.1860\.972DeepCas1\.1581\.5561\.9552\.0141\.7320\.6861\.4601\.8411\.9611\.6360\.8021\.0951\.3631\.4541\.2280\.7681\.0041\.2291\.3411\.122CasSeqGCN1\.1291\.5761\.9542\.0021\.8140\.5850\.9671\.3781\.4541\.2430\.2260\.4580\.7280\.8070\.6580\.0840\.2000\.3590\.4270\.347Graph\-LSTM1\.0901\.5081\.8891\.9511\.7210\.5410\.9311\.3261\.4191\.1820\.2060\.4480\.7130\.7790\.6470\.0910\.2190\.3820\.4650\.365MMG\-PopNet1\.0191\.4191\.7241\.7891\.5590\.6130\.9631\.2861\.3601\.1700\.3090\.4830\.7150\.7800\.6330\.1230\.2100\.3330\.3910\.320Max DepthMLP0\.4220\.4680\.4570\.4540\.5720\.2910\.3620\.3780\.3870\.4790\.1740\.2310\.2720\.2760\.2940\.1020\.1430\.1820\.1910\.212DeepHawkes0\.4430\.5080\.5290\.5400\.6210\.3430\.4250\.4910\.5180\.5550\.3070\.3640\.4460\.4550\.4440\.2970\.3270\.3780\.4230\.425DeepCas0\.4500\.5000\.5180\.5060\.5890\.4190\.4950\.5120\.5270\.5720\.3200\.3800\.4130\.4220\.4170\.2740\.3330\.3620\.3540\.361CasSeqGCN0\.4320\.4830\.4810\.4720\.5860\.3050\.3710\.4000\.4120\.4830\.2080\.2570\.2960\.3100\.3140\.1140\.1650\.1980\.2110\.227Graph\-LSTM0\.4210\.4690\.4690\.4600\.5690\.2790\.3590\.3830\.3970\.4720\.1590\.2250\.2680\.2730\.2920\.0680\.1310\.1750\.2010\.212MMG\-PopNet0\.3970\.4530\.4400\.4360\.5220\.2790\.3550\.3600\.3710\.4470\.1690\.2160\.2510\.2620\.2760\.0710\.1180\.1560\.1730\.198StructuralViralityMLP0\.2020\.1980\.1690\.1610\.1930\.1430\.1560\.1480\.1470\.1650\.0750\.0940\.1110\.1130\.1040\.0380\.0530\.0660\.0710\.072DeepHawkes0\.2100\.2110\.1980\.2000\.2090\.1660\.1790\.1870\.1920\.1900\.1550\.1670\.2000\.2000\.1730\.1450\.1360\.1470\.1620\.154DeepCas0\.2220\.2190\.2030\.1930\.2100\.2100\.2220\.2080\.2080\.2090\.1510\.1600\.1710\.1700\.1550\.1230\.1340\.1390\.1330\.132CasSeqGCN0\.2070\.2050\.1780\.1760\.1980\.1520\.1580\.1540\.1550\.1650\.0940\.1080\.1250\.1290\.1140\.0450\.0620\.0740\.0820\.081Graph\-LSTM0\.2010\.1990\.1760\.1650\.1920\.1380\.1530\.1480\.1480\.1610\.0680\.0900\.1070\.1070\.0990\.0280\.0500\.0650\.0730\.070MMG\-PopNet0\.1970\.1960\.1710\.1610\.1820\.1450\.1560\.1400\.1380\.1570\.0840\.0910\.1020\.1050\.0990\.0330\.0460\.0570\.0640\.068SizeMLP1\.6742\.2762\.6782\.7562\.6580\.9791\.5982\.0942\.1522\.0550\.3990\.7651\.1461\.2551\.0800\.2640\.4250\.6460\.7190\.642DeepHawkes1\.7432\.5283\.1363\.2532\.9431\.2471\.9482\.7332\.9542\.4371\.0621\.6332\.3482\.5032\.0040\.8451\.0931\.5131\.7391\.455DeepCas1\.7712\.3472\.8262\.8702\.6321\.6292\.2872\.7682\.9232\.5441\.1381\.6162\.0112\.1461\.8031\.0661\.4561\.7781\.8721\.602CasSeqGCN1\.7362\.3902\.8122\.8292\.7550\.9631\.5732\.1282\.2352\.0130\.3720\.7861\.2111\.3581\.1030\.0850\.3150\.5750\.6650\.566Graph\-LSTM1\.6502\.2582\.7042\.7492\.6040\.8901\.5092\.0372\.1561\.9190\.3500\.7621\.1581\.2501\.0700\.1310\.3620\.6290\.7500\.612MMG\-PopNet1\.5592\.1292\.4802\.5162\.3530\.9611\.5001\.9071\.9921\.8230\.4270\.7311\.0741\.1690\.9620\.1430\.3020\.4980\.5680\.506UniqueUsersMLP1\.3441\.8882\.2822\.3452\.0890\.8271\.3591\.8081\.8351\.6190\.3420\.6620\.9791\.0670\.8620\.2410\.3850\.5710\.6240\.515DeepHawkes1\.4082\.0952\.6502\.7362\.3261\.0241\.6362\.3242\.4811\.9280\.8051\.2821\.8541\.9661\.5160\.6760\.9251\.2921\.4731\.153DeepCas1\.4251\.9482\.3852\.4112\.0721\.3121\.8872\.3182\.4271\.9990\.9171\.3371\.6621\.7661\.4020\.8821\.2251\.5061\.5991\.276CasSeqGCN1\.3851\.9712\.3952\.4152\.1630\.8051\.3321\.8261\.8971\.5820\.3080\.6581\.0071\.1220\.8580\.0920\.2830\.5010\.5610\.437Graph\-LSTM1\.3321\.8772\.3002\.3442\.0430\.7241\.2571\.7241\.8081\.4890\.3040\.6641\.0011\.0750\.8340\.1370\.3160\.5360\.6180\.451MMG\-PopNet1\.2501\.7722\.1222\.1511\.8490\.7731\.2471\.6241\.6911\.4250\.3700\.6330\.9170\.9970\.7640\.1430\.2750\.4380\.4800\.389Like ScoreMLP––––6\.469––––6\.309––––4\.524––––3\.709DeepHawkes––––7\.002––––6\.877––––6\.688––––5\.646DeepCas––––5\.749––––6\.013––––4\.828––––4\.925CasSeqGCN––––6\.800––––6\.639––––5\.402––––5\.093Graph\-LSTM––––6\.447––––5\.822––––4\.674––––4\.012MMG\-PopNet––––5\.308––––4\.886––––3\.428––––2\.897

Description:The results that for given 180 min of observation data, graph\-based models like CasSeqGCN can achieve near\-perfect accuracy \(MSE < 0\.1\) for immediate 4\-hour horizon predictions regarding cascadeSizeandMax Width\. However, prediction accuracy decays steeply as the horizon extends to 24 hours\. MMG\-PopNet distinguishes itself most prominently in theLike Scorecategory, leveraging multimodal data to achieve an MSE of 2\.897 at the 180\-minute mark, vastly outperforming the nearest baseline \(MLP at 3\.709\)\.

#### D\.3Statistical Significance\.

To rigorously assess whether MMG\-PopNet’s improvements over baseline models reflect genuine performance gains rather than sampling variability, we conduct a formal statistical significance analysis on the targetSizefor the final cascade state\. For each dataset, we use the largest available observation window as the input context\. This corresponds to 20 min for Bluesky, 90 min for Gaming, 60 min for AMA, and 180 min for Futurology\. The largest observation window provides each model with the richest possible input signal, making it the most demanding and representative setting in which to compare methods, as any advantage held by MMG\-PopNet cannot be attributed to information asymmetry\. We apply a paired bootstrap significance test\[[17](https://arxiv.org/html/2606.27539#bib.bib89)\]\(5,000 resamples\) in log1p\-MSE space\. The test is paired because all models predict the same set of cascades, matched by cascade identifier, and this pairing removes between\-cascade variance\. Errors are computed in log1p space similar to other experiments\. For each test cascadeG∈𝒢TestG\\in\\mathcal\{G\}^\{\\mathrm\{Test\}\}, the ground\-truth target isY~Gt′=log⁡\(1\+YGt′\)\\widetilde\{\\textbf\{Y\}\}\_\{G\}^\{t^\{\\prime\}\}=\\log\(1\+\\textbf\{Y\}\_\{G\}^\{t^\{\\prime\}\}\), and each model predictsY^Gt′\\widehat\{\\textbf\{Y\}\}\_\{G\}^\{t^\{\\prime\}\}\. We compute the per\-cascade squared error of modelmmaseG\(m\)=\(Y^Gt′,\(m\)−Y~Gt′\)2\.e\_\{G\}^\{\(m\)\}=\(\\widehat\{\\textbf\{Y\}\}\_\{G\}^\{t^\{\\prime\},\(m\)\}\-\\widetilde\{\\textbf\{Y\}\}\_\{G\}^\{t^\{\\prime\}\}\)^\{2\}\.The reported effect size compares a baseline modelbbagainst MMG\-PopNet:Δb=1\|𝒢Test\|​∑G∈𝒢Test\(eG\(b\)−eG\(MMG\)\)\.\\Delta\_\{b\}=\\frac\{1\}\{\|\\mathcal\{G\}^\{\\mathrm\{Test\}\}\|\}\\sum\_\{G\\in\\mathcal\{G\}^\{\\mathrm\{Test\}\}\}\(e\_\{G\}^\{\(b\)\}\-e\_\{G\}^\{\(\\mathrm\{MMG\}\)\}\)\.Thus,Δb\>0\\Delta\_\{b\}\>0indicates that MMG\-PopNet has lower squared error than baselinebb, whileΔb<0\\Delta\_\{b\}<0would indicate that the baseline performs better\. The 95% confidence interval is obtained using the bootstrap percentile method\. Since we conduct simultaneous multiple comparisons, we apply Benjamini–Hochberg FDR correction to control the rate of false discoveries across all tests jointly\.

###### Results\.

Table[11](https://arxiv.org/html/2606.27539#A4.T11)reports the full results\. MMG\-PopNet significantly outperforms all four baselines on all four datasets \(16/16 comparisons,pBH<0\.05p\_\{\\mathrm\{BH\}\}<0\.05\)\. All confidence intervals are strictly positive, with lower bounds above zero\. The largest improvements are against DeepCas, whereΔb\\Delta\_\{b\}exceeds\+1\.0\+1\.0in log1p\-MSE on every dataset, including Gaming \(Δb=\+1\.573\\Delta\_\{b\}=\+1\.573\) and Bluesky \(Δb=\+1\.522\\Delta\_\{b\}=\+1\.522\)\. CasSeqGCN is the closest competitor\. The smallest significant improvement occurs on Futurology \(Δb=\+0\.102\\Delta\_\{b\}=\+0\.102, 95% CI\[\+0\.029,\+0\.183\]\[\+0\.029,\+0\.183\],pBH=0\.004p\_\{\\mathrm\{BH\}\}=0\.004\)\. These results show that MMG\-PopNet’s gains are consistent, statistically reliable, and not driven by any single dataset or baseline\.

Table 11:Statistical significance ofMMG\-PopNet \(Ours\)versus baselines forSizeprediction one the final cascade state in log1p\-MSE space\.Δb\>0\\Delta\_\{b\}\>0indicates that MMG\-PopNet has lower squared error than baselinebb\. 95% confidence intervals andprawp\_\{\\rm raw\}are computed using a paired bootstrap test with 5,000 resamples\.pBHp\_\{\\rm BH\}denotes Benjamini–Hochberg correction across all 16 comparisons atα=0\.05\\alpha=0\.05\.Yesindicates that MMG\-PopNet is significantly better\.DatasetBaselinenn𝚫𝒃\\bm\{\\Delta\_\{b\}\}95% CI𝒑𝐫𝐚𝐰\\bm\{p\_\{\\rm raw\}\}𝒑𝐁𝐇\\bm\{p\_\{\\rm BH\}\}Sig\.BlueskyCasSeqGCN1510\+0\.2213\+0\.2213\[\+0\.1729,\+0\.2717\]\[\+0\.1729,\\ \+0\.2717\]<0\.001<0\.001<0\.001<0\.001YesGraph\-LSTM1510\+0\.1375\+0\.1375\[\+0\.0983,\+0\.1795\]\[\+0\.0983,\\ \+0\.1795\]<0\.001<0\.001<0\.001<0\.001YesDeepHawkes1510\+0\.3601\+0\.3601\[\+0\.3054,\+0\.4171\]\[\+0\.3054,\\ \+0\.4171\]<0\.001<0\.001<0\.001<0\.001YesDeepCas1510\+1\.1522\+1\.1522\[\+1\.0206,\+1\.2909\]\[\+1\.0206,\\ \+1\.2909\]<0\.001<0\.001<0\.001<0\.001Yesr/GamingCasSeqGCN675\+0\.1328\+0\.1328\[\+0\.0572,\+0\.2078\]\[\+0\.0572,\\ \+0\.2078\]<0\.001<0\.001<0\.001<0\.001YesGraph\-LSTM675\+0\.3519\+0\.3519\[\+0\.2695,\+0\.4358\]\[\+0\.2695,\\ \+0\.4358\]<0\.001<0\.001<0\.001<0\.001YesDeepHawkes675\+0\.6019\+0\.6019\[\+0\.4750,\+0\.7306\]\[\+0\.4750,\\ \+0\.7306\]<0\.001<0\.001<0\.001<0\.001YesDeepCas675\+1\.5730\+1\.5730\[\+1\.3016,\+1\.8464\]\[\+1\.3016,\\ \+1\.8464\]<0\.001<0\.001<0\.001<0\.001Yesr/AMACasSeqGCN1420\+0\.0986\+0\.0986\[\+0\.0696,\+0\.1298\]\[\+0\.0696,\\ \+0\.1298\]<0\.001<0\.001<0\.001<0\.001YesGraph\-LSTM1420\+0\.2269\+0\.2269\[\+0\.1923,\+0\.2626\]\[\+0\.1923,\\ \+0\.2626\]<0\.001<0\.001<0\.001<0\.001YesDeepHawkes1420\+0\.2167\+0\.2167\[\+0\.1776,\+0\.2578\]\[\+0\.1776,\\ \+0\.2578\]<0\.001<0\.001<0\.001<0\.001YesDeepCas1420\+1\.0033\+1\.0033\[\+0\.8893,\+1\.1251\]\[\+0\.8893,\\ \+1\.1251\]<0\.001<0\.001<0\.001<0\.001Yesr/FuturologyCasSeqGCN527\+0\.1016\+0\.1016\[\+0\.0293,\+0\.1828\]\[\+0\.0293,\\ \+0\.1828\]0\.0040\.0040\.0040\.004YesGraph\-LSTM527\+0\.2892\+0\.2892\[\+0\.1987,\+0\.3867\]\[\+0\.1987,\\ \+0\.3867\]<0\.001<0\.001<0\.001<0\.001YesDeepHawkes527\+1\.3799\+1\.3799\[\+1\.1493,\+1\.6302\]\[\+1\.1493,\\ \+1\.6302\]<0\.001<0\.001<0\.001<0\.001YesDeepCas527\+1\.3210\+1\.3210\[\+1\.0222,\+1\.6373\]\[\+1\.0222,\\ \+1\.6373\]<0\.001<0\.001<0\.001<0\.001Yes

### Appendix EUnified Training Details and Full Results

Tables[12](https://arxiv.org/html/2606.27539#A5.T12)and[13](https://arxiv.org/html/2606.27539#A5.T13)report the full comparison between dataset\-specific MMG\-PopNet and unified\-trained MMG\-PopNet across datasets, observation windows, and prediction targets\. The dataset\-specific setting trains a separate model for each dataset and observation window\. In contrast, MMG\-PopNet \(unified\-dataset\) is trained once on the combined training split and is evaluated separately on each dataset\-window test split\. For each training example, the unified model receives one\-hot indicators for the dataset and observation window\. These indicators allow the model to condition its predictions on both the community source and the amount of observed cascade history\.

The full results show that the benefit of unified training is largest on the Reddit datasets\. On r/AMA, r/Gaming, and r/Futurology, MMG\-PopNet \(unified\-dataset\) reduces the dataset\-level average MSE for every prediction target\. The reductions are most pronounced forLike Score,Size, andUnique Users\. These targets have the largest dataset\-level errors under dataset\-specific training on r/Gaming and r/Futurology, and they also show the largest absolute reductions after unified training\. For example, on r/Gaming, the dataset\-level average MSE forLike Scoredecreases from 4\.525 to 1\.479\. On r/Futurology, the corresponding average decreases from 3\.890 to 1\.569\.

Bluesky follows a different pattern\. On this dataset, the dataset\-specific model obtains lower dataset\-level average MSE for most targets\. The unified model remains close in absolute MSE, but it does not improve over the dataset\-specific model on Bluesky\. This result indicates that the cross\-dataset benefit is stronger when the evaluation dataset is closer to the Reddit training communities\.

Averaged across all datasets and observation windows, MMG\-PopNet \(unified\-dataset\) obtains lower MSE for all six targets\. The overall average decreases from 0\.678 to 0\.426 forMax Width, from 0\.284 to 0\.232 forMax Depth, from 0\.932 to 0\.586 forSize, from 0\.104 to 0\.084 forStructural Virality, from 0\.752 to 0\.453 forUnique Users, and from 2\.669 to 1\.217 forLike Score\. Thus, unified training improves benchmark\-level performance across targets while maintaining comparable accuracy on the outlying Bluesky platform\.

Table 12:Comparison between dataset\-specific and unified\-dataset MMG\-PopNet training, Part I\.Dataset\-specific models are trained separately for each dataset and observation window, while the unified\-dataset model is trained once on the combined training data across all datasets and windows\. Results are reported as MSE, wherelower is better and marked in bold\. For readability, results are split by dataset\. Each subtable reports observation\-window MSEs, the dataset\-level average, and the same final overall average across all datasets\. This table reports results for Bluesky and r/AMA; Table[13](https://arxiv.org/html/2606.27539#A5.T13)continues with r/Gaming and r/Futurology\.BlueskyTaskModel021020Dataset AvgOverall AvgMax WidthMMG\-PopNet \(specific\)0\.4570\.3690\.2840\.2340\.3360\.678MMG\-PopNet \(unified\)0\.4490\.4010\.3160\.2760\.3610\.426Max DepthMMG\-PopNet \(specific\)0\.3630\.3250\.2730\.2380\.3000\.284MMG\-PopNet \(unified\)0\.3470\.3320\.2800\.2470\.3010\.232SizeMMG\-PopNet \(specific\)0\.7050\.5870\.4700\.3970\.5400\.932MMG\-PopNet \(unified\)0\.6880\.6370\.5170\.4500\.5730\.586StructuralViralityMMG\-PopNet \(specific\)0\.1550\.1350\.1140\.1010\.1260\.104MMG\-PopNet \(unified\)0\.1440\.1380\.1170\.1040\.1260\.084UniqueUsersMMG\-PopNet \(specific\)0\.4670\.3830\.3020\.2550\.3520\.752MMG\-PopNet \(unified\)0\.4720\.4260\.3410\.2980\.3840\.453LikeScoreMMG\-PopNet \(specific\)1\.2601\.0871\.0260\.9501\.0812\.669MMG\-PopNet \(unified\)1\.2381\.1911\.0931\.0431\.1411\.217
r/AMATaskModel0153060Dataset AvgOverall AvgMax WidthMMG\-PopNet \(specific\)0\.6250\.4910\.4290\.2760\.4550\.678MMG\-PopNet \(unified\)0\.4760\.3540\.2390\.1840\.3130\.426Max DepthMMG\-PopNet \(specific\)0\.3140\.2560\.2190\.1720\.2400\.284MMG\-PopNet \(unified\)0\.2840\.2360\.1890\.1580\.2170\.232SizeMMG\-PopNet \(specific\)0\.8630\.6610\.5730\.3660\.6160\.932MMG\-PopNet \(unified\)0\.6810\.5090\.3490\.2570\.4490\.586StructuralViralityMMG\-PopNet \(specific\)0\.1300\.1000\.0870\.0640\.0950\.104MMG\-PopNet \(unified\)0\.1150\.0930\.0750\.0600\.0860\.084UniqueUsersMMG\-PopNet \(specific\)0\.6220\.4940\.4470\.2900\.4630\.752MMG\-PopNet \(unified\)0\.4540\.3430\.2350\.1800\.3030\.453LikeScoreMMG\-PopNet \(specific\)1\.3111\.2121\.1661\.0281\.1792\.669MMG\-PopNet \(unified\)0\.8100\.7370\.5900\.5820\.6801\.217

Table 13:Comparison between dataset\-specific and unified\-dataset MMG\-PopNet training, Part II\.Continuation of Table[12](https://arxiv.org/html/2606.27539#A5.T12)\. Results are reported as MSE, wherelower is better and marked in bold\. For readability, results are split by dataset\. Each subtable reports observation\-window MSEs, the dataset\-level average, and the same final overall average across all datasets\. This table reports results for r/Gaming and r/Futurology\.r/GamingTaskModel0205090Dataset AvgOverall AvgMax WidthMMG\-PopNet \(specific\)1\.5571\.0960\.7140\.5620\.9820\.678MMG\-PopNet \(unified\)0\.9620\.6310\.2990\.2660\.5390\.426Max DepthMMG\-PopNet \(specific\)0\.3350\.2750\.2000\.1600\.2430\.284MMG\-PopNet \(unified\)0\.2480\.2040\.1400\.1210\.1780\.232SizeMMG\-PopNet \(specific\)1\.8291\.3170\.8420\.6531\.1600\.932MMG\-PopNet \(unified\)1\.0920\.7540\.3480\.2940\.6220\.586StructuralViralityMMG\-PopNet \(specific\)0\.0980\.0810\.0570\.0450\.0700\.104MMG\-PopNet \(unified\)0\.0710\.0570\.0380\.0320\.0500\.084UniqueUsersMMG\-PopNet \(specific\)1\.6931\.2180\.7920\.6131\.0790\.752MMG\-PopNet \(unified\)1\.0040\.6840\.3180\.2740\.5700\.453LikeScoreMMG\-PopNet \(specific\)5\.0674\.6544\.3074\.0734\.5252\.669MMG\-PopNet \(unified\)2\.0771\.8090\.9601\.0701\.4791\.217
r/FuturologyTaskModel03090180Dataset AvgOverall AvgMax WidthMMG\-PopNet \(specific\)1\.5811\.1320\.6650\.3720\.9380\.678MMG\-PopNet \(unified\)0\.8860\.6340\.2420\.2020\.4910\.426Max DepthMMG\-PopNet \(specific\)0\.5170\.4180\.2820\.2050\.3560\.284MMG\-PopNet \(unified\)0\.3360\.2820\.1640\.1470\.2320\.232SizeMMG\-PopNet \(specific\)2\.3561\.7430\.9970\.5591\.4140\.932MMG\-PopNet \(unified\)1\.2530\.9320\.3480\.2670\.7000\.586StructuralViralityMMG\-PopNet \(specific\)0\.1820\.1460\.1010\.0730\.1260\.104MMG\-PopNet \(unified\)0\.1100\.0900\.0530\.0470\.0750\.084UniqueUsersMMG\-PopNet \(specific\)1\.8591\.3740\.7850\.4391\.1140\.752MMG\-PopNet \(unified\)0\.9870\.7360\.2770\.2190\.5550\.453LikeScoreMMG\-PopNet \(specific\)4\.9994\.5963\.2142\.7503\.8902\.669MMG\-PopNet \(unified\)2\.2832\.0541\.0410\.8981\.5691\.217

### Appendix FDetailed LLM\-based comparison\.

To evaluate whether LLMs can serve as competitive predictors for multimodal social popularity forecasting, we conduct experiments comparing the performance of various LLM\-based approaches against MMG\-PopNet\. We evaluate three models, namelyQwen3\-VL\-8B\-Instruct,Gemma\-3\-12b\-it, andGPT\-4o\-mini\. To manage computational costs for the Bluesky dataset, we randomly sample 15,000 cascades from the original data and evaluate all models including MMG\-PopNet on this subset\. Furthermore, large cascades can contain excessive content that exceeds standard context windows and prevents fair comparison\. We address this by restricting the maximum number of observed nodes to 25 for all social cascades and apply this exact limit to MMG\-PopNet\. Given the early observation data statistics where the average node count ranges from 1\.62 to 23\.6, this setting provides sufficient context for the vast majority of cascades\.

We evaluate three prompting and training settings\. In the zero\-shot setting, the observed cascade prefix is serialized as a structured JSON input containing the available content and interaction, along with an image, and the model predicts the popularity targets directly\. In the retrieval\-augmented few\-shot setting, each test instance is paired with four training examples selected by root\-post similarity, and their ground\-truth targets are included in the prompt\. In the fine\-tuning setting, the model is trained on early\-observation cascade inputs to predict the same targets as MMG\-PopNet\. Note thatGPT\-4o\-miniis only evaluated under the zero\-shot and few\-shot settings\. Prompts used for the experiment are provided here[F](https://arxiv.org/html/2606.27539#A6.SS0.SSS0.Px1)\.

Tables[14](https://arxiv.org/html/2606.27539#A6.T14)\-[19](https://arxiv.org/html/2606.27539#A6.T19)report results across the LLM models and settings, and datasets\. Across both datasets and all tested models, MMG\-PopNet achieves the lowest MSE for every target and observation window\. Zero\-shot prompting consistently performs worst, confirming that structured input alone is insufficient for calibrated numerical prediction\. Few\-shot prompting substantially reduces error, especially under root\-only and early observation windows where retrieved examples provide useful target\-range calibration for sparse cascades\. Fine\-tuning becomes more effective as the observation window increases forQwen3\-VL\-8B\-Instructand to some extent forGemma\-3\-12b\-it\. This trend is most visible for targets such asMax Width,Structural Virality,Size, andUnique Users, where longer prefixes expose richer interaction and temporal signals that supervised adaptation can exploit\. Here,Qwen3\-VL\-8B\-Instructexhibits a more consistent increase with fine\-tuning whereas, onGemma\-3\-12b\-itfew\-shot setting is better across observation windows for most targets for the r/Futurology dataset and with mixed results for Bluesky dataset\.

ForGPT\-4o\-mini, the few shot setting results as the best setting when compared with the zero\-shot\. However, it still lags far behind the MMG\-PopNet in performance\. Improvement in few\-shot setting with increased observation window is higher and consistent in r/Futurology dataset when compared to Bluesky\.

Overall, the tables show that fine\-tuning does not uniformly dominate few\-shot prompting\. In several early\-window cases, few\-shot prompting remains competitive or stronger, particularly forMax Depthand some engagement targets\. This indicates that retrieval is highly useful when the observed prefix contains limited cascade evidence, while fine\-tuning benefits more from denser prefixes\. Overall, the detailed results support the main finding\. LLM\-based adaptation improves over zero\-shot prompting, but MMG\-PopNet remains consistently stronger because it is optimized for the benchmark formulation and directly models observed cascade dynamics rather than predicting from serialized inputs alone\.

Table 14:LLM Qwen3\-VL\-8B\-Instruct comparison with MMG\-PopNet on Bluesky across observation windows using MSE values\.TaskRoot Only21020ZeroFewFine\-TuneMMG\-PopNetZeroFewFine\-TuneMMG\-PopNetZeroFewFine\-TuneMMG\-PopNetZeroFewFine\-TuneMMG\-PopNetMax Width3\.0451\.6661\.9230\.6662\.6151\.5641\.2950\.5772\.1241\.0380\.8160\.5061\.9890\.8890\.6470\.459Max Depth1\.7150\.7300\.9050\.5231\.5770\.7900\.8350\.4801\.2400\.6660\.6770\.4281\.0480\.5610\.6040\.415StructuralVirality2\.3680\.3630\.4290\.2431\.9470\.4140\.3860\.2251\.4450\.3530\.3030\.1961\.2560\.2980\.2670\.191Size6\.5962\.2932\.6681\.0075\.2942\.2891\.9830\.8703\.5441\.4761\.2840\.7512\.8421\.1941\.0650\.700Unique Users4\.0511\.8672\.0090\.6903\.3891\.7391\.3490\.6002\.5571\.0850\.8370\.5212\.2130\.9040\.6750\.468Like Score16\.2565\.1216\.6251\.54016\.1624\.9805\.1701\.44514\.3834\.0873\.8461\.39413\.4443\.6153\.1641\.327

Table 15:LLM Qwen3\-VL\-8B\-Instruct comparison with MMG\-PopNet on r/Futurology across observation windows using MSE values\.TaskRoot Only30 min90 min180 minZeroFewFine\-TuneMMG\-PopNetZeroFewFine\-TuneMMG\-PopNetZeroFewFine\-TuneMMG\-PopNetZeroFewFine\-TuneMMG\-PopNetMax Width4\.9122\.2742\.4661\.5813\.6612\.0051\.7741\.3212\.9521\.4170\.9710\.7942\.7581\.1550\.7820\.457Max Depth1\.6050\.7851\.0120\.5171\.1510\.6850\.7820\.4960\.7690\.4390\.5170\.2960\.5510\.3570\.4250\.242StructuralVirality2\.1640\.2640\.3430\.1821\.1980\.2650\.2620\.1780\.6670\.1880\.1790\.1040\.4340\.1600\.1410\.091Size9\.2383\.8954\.0802\.3565\.7063\.0962\.9531\.9213\.5721\.9861\.7211\.1412\.6301\.4661\.3580\.700Unique Users6\.6933\.0053\.0281\.8594\.3652\.4402\.1811\.5512\.7981\.6161\.2730\.9172\.0671\.2171\.0380\.552Like Score21\.1358\.8598\.8634\.99920\.0188\.1788\.5725\.03017\.0776\.9336\.9443\.40013\.9817\.0526\.0163\.039

Table 16:LLM Gemma\-3\-12b\-it comparison with MMG\-PopNet on Bluesky across observation windows using MSE values\.TaskRoot Only21020ZeroFewFine\-TuneMMG\-PopNetZeroFewFine\-TuneMMG\-PopNetZeroFewFine\-TuneMMG\-PopNetZeroFewFine\-TuneMMG\-PopNetMax Width5\.3661\.3121\.9930\.6663\.7961\.3201\.3910\.5772\.5450\.9690\.8120\.5062\.1790\.8600\.6570\.459Max Depth2\.2850\.6280\.9610\.5231\.6550\.6590\.8000\.4801\.2690\.5770\.7210\.4281\.1050\.5330\.6030\.415StructuralVirality2\.3620\.2860\.4900\.2431\.8280\.3140\.4050\.2251\.2760\.2800\.3200\.1961\.0670\.2640\.2770\.191Size6\.5771\.6252\.8191\.0075\.1861\.6562\.0380\.8703\.4061\.2001\.3550\.7512\.6751\.0611\.1020\.700Unique Users4\.0411\.4541\.9680\.6903\.1321\.3481\.3950\.6001\.9560\.9600\.8330\.5211\.4910\.8250\.6720\.468Like Score16\.2544\.1497\.5861\.54015\.1353\.9985\.7821\.44514\.3553\.5473\.6141\.39414\.0183\.4623\.2211\.327

Table 17:LLM Gemma\-3\-12b\-it comparison with MMG\-PopNet on r/Futurology across observation windows using MSE values\.TaskRoot Only30 min90 min180 minZeroFewFine\-TuneMMG\-PopNetZeroFewFine\-TuneMMG\-PopNetZeroFewFine\-TuneMMG\-PopNetZeroFewFine\-TuneMMG\-PopNetMax Width7\.8371\.9842\.8311\.5814\.0271\.9481\.9001\.3213\.0581\.3501\.2740\.7942\.6801\.0261\.0170\.457Max Depth2\.4390\.6871\.2990\.5171\.3110\.6500\.8910\.4960\.9740\.4570\.5890\.2960\.7660\.3210\.5090\.242StructuralVirality2\.2040\.2410\.4290\.1820\.9690\.2370\.2960\.1780\.6370\.1770\.1960\.1040\.4490\.1320\.1700\.091Size9\.2383\.2264\.8692\.3565\.5752\.9483\.3051\.9213\.4151\.9302\.2021\.1412\.5101\.3611\.8160\.700Unique Users6\.6932\.4973\.4701\.8594\.3192\.2832\.5761\.5512\.6761\.5071\.7310\.9171\.9511\.0741\.4060\.552Like Score21\.1167\.74712\.9864\.99919\.0778\.05513\.2025\.03018\.9796\.80312\.1233\.40018\.7666\.93810\.6493\.039

Table 18:LLM GPT\-4o\-mini comparison with MMG\-PopNet on Bluesky across observation windows using MSE values\.TaskRoot Only21020ZeroFewMMG\-PopNetZeroFewMMG\-PopNetZeroFewMMG\-PopNetZeroFewMMG\-PopNetMax Width2\.7421\.6660\.6662\.2681\.5620\.5771\.6321\.4860\.5061\.3861\.3950\.459Max Depth1\.5770\.5900\.5231\.3980\.5980\.4801\.0550\.5390\.4280\.9060\.4840\.415StructuralVirality0\.7990\.2750\.2430\.7460\.2810\.2250\.6140\.2480\.1960\.5320\.2230\.191Size5\.5311\.8801\.0074\.5451\.7890\.8703\.0961\.4640\.7512\.4461\.3070\.700Unique Users3\.4721\.6900\.6902\.7671\.5300\.6001\.7961\.2660\.5211\.3921\.1150\.468Like Score12\.5664\.0021\.54014\.1303\.8521\.44513\.8283\.4781\.39413\.4153\.2121\.327

Table 19:LLM GPT\-4o\-mini comparison with MMG\-PopNet on r/Futurology across observation windows using MSE values\.TaskRoot Only30 min90 min180 minZeroFewMMG\-PopNetZeroFewMMG\-PopNetZeroFewMMG\-PopNetZeroFewMMG\-PopNetMax Width4\.8532\.1791\.5813\.3882\.1731\.3212\.4491\.8130\.7942\.1181\.5220\.457Max Depth1\.4760\.6970\.5171\.1260\.6230\.4960\.7490\.4310\.2960\.5510\.3600\.242StructuralVirality0\.8470\.2370\.1820\.6920\.2210\.1780\.5260\.1800\.1040\.4050\.1580\.091Size8\.4183\.4512\.3565\.4443\.1631\.9213\.2262\.2081\.1412\.2381\.6120\.700Unique Users6\.1232\.6481\.8594\.0562\.5111\.5512\.4601\.7580\.9171\.6891\.3180\.552Like Score20\.2917\.3414\.99920\.5467\.9305\.03020\.3856\.9323\.40019\.9587\.0703\.039

###### Zero\-Shot Prompt

The zero\-shot prompt is constructed once per query with no retrieved examples\. The placeholder<WINDOW\>is replaced by the observation window size, and<EARLY\_CONVERSATION\_TREE\_JSON\>is replaced by the serialised conversation tree\.

Prompt Template — Zero\-Shot[⬇](data:text/plain;base64,WW91IGFyZSBhbmFseXppbmcgYSBzb2NpYWwgbWVkaWEgY29udmVyc2F0aW9uIHRyZWUgdG8gcHJlZGljdCBpdHMgZmluYWwgZ3Jvd3RoIG1ldHJpY3MuIFlvdSB3aWxsIGJlIHByb3ZpZGVkIHdpdGggdGhlIGZpcnN0IDxXSU5ET1c+IG9mIGEgc29jaWFsIG1lZGlhIGNvbnZlcnNhdGlvbiB0aHJlYWQuIFRoaXMgY29udmVyc2F0aW9uIHRocmVhZCBjb250aW51ZWQgYWZ0ZXIgdGhpcy4KCkJhc2VkIG9uIHRoZSBjb250ZW50LCB0aW1pbmcsIGFuZCBzdHJ1Y3R1cmFsIHBhdHRlcm5zIG9mIHRoZXNlIGVhcmx5IGludGVyYWN0aW9ucywgcHJlZGljdCB0aGUgRklOQUwgc3RhdGUgb2YgdGhpcyBjb252ZXJzYXRpb24gdHJlZSB3aGVuIGl0IHJlYWNoZXMgc2F0dXJhdGlvbiAodGhhdCBpcywgc3RvcHMgZ3Jvd2luZykuCgpNRVRSSUMgREVGSU5JVElPTlM6Ci0gbWF4X3dpZHRoOiBNYXhpbXVtIG51bWJlciBvZiByZXBsaWVzIGF0IGFueSBzaW5nbGUgZGVwdGggbGV2ZWwuCi0gbWF4X2RlcHRoOiBNYXhpbXVtIGRlcHRoIG9mIHRoZSBjb252ZXJzYXRpb24gdHJlZS4KLSBzdHJ1Y3R1cmFsX3ZpcmFsaXR5OiBBdmVyYWdlIGRpc3RhbmNlIGJldHdlZW4gYWxsIHBhaXJzIG9mIG5vZGVzIChMb3cgPSBCcm9hZGNhc3QsIEhpZ2ggPSBWaXJhbCkuCi0gbnVtX3Bvc3RzOiBUb3RhbCBudW1iZXIgb2YgcG9zdHMgaW4gdGhlIGZpbmFsIGNvbnZlcnNhdGlvbiAocm9vdCArIHJlcGxpZXMpLgotIG51bV91bmlxdWVfdXNlcnM6IFRvdGFsIG51bWJlciBvZiB1bmlxdWUgdXNlcnMgaW4gdGhlIGZpbmFsIGNvbnZlcnNhdGlvbi4KLSByb290X3Njb3JlOiBFbmdhZ2VtZW50IHNjb3JlIG9mIHRoZSBpbnRpYWwgcm9vdCBwb3N0LgoKWW91IHdpbGwgYmUgZ2l2ZW4gYSBORVNURUQgSlNPTiBUUkVFLiBFYWNoIHJlcGx5IG5vZGUgaGFzIHRleHQsIHRpbWVfc2luY2Vfcm9vdCwgdGltZV9zaW5jZV9wYXJlbnQsIGFuZCBjaGlsZHJlbi4gVGhlIHJvb3Qgbm9kZSBoYXMgdGV4dCwgdGltZXN0YW1wLCBhbmQgY2hpbGRyZW4uCgpPVVRQVVQgRk9STUFUOgpSZXR1cm4gYSBzaW5nbGUgSlNPTiBvYmplY3QuIERvIG5vdCBpbmNsdWRlIG1hcmtkb3duIGZvcm1hdHRpbmcuCnsibWF4X3dpZHRoIjogPG51bWJlcj4sICJtYXhfZGVwdGgiOiA8bnVtYmVyPiwgInN0cnVjdHVyYWxfdmlyYWxpdHkiOiA8bnVtYmVyPiwgIm51bV9wb3N0cyI6IDxudW1iZXI+LCAibnVtX3VuaXF1ZV91c2VycyI6IDxudW1iZXI+LCAicm9vdF9zY29yZSI6IDxudW1iZXI+fQoKRUFSTFkgQ09OVkVSU0FUSU9OIFRSRUU6CjxFQVJMWV9DT05WRVJTQVRJT05fVFJFRV9KU09OPgoKSU5TVFJVQ1RJT05TOgpVc2luZyBvbmx5IHRoZSBjb252ZXJzYXRpb24gdHJlZSBhYm92ZSBhcyBldmlkZW5jZSwgcHJlZGljdCB0aGUgRklOQUwgdmFsdWVzLgpEbyBOT1QgaW5jbHVkZSBhbnkgZXhwbGFuYXRpb24gb3IgZXh0cmEgZmllbGRzLgpPdXRwdXQgZXhhY3RseSBvbmUgSlNPTiBvYmplY3Qgd2l0aCB0aGVzZSBrZXlzIG9ubHk6IG1heF93aWR0aCwgbWF4X2RlcHRoLCBzdHJ1Y3R1cmFsX3ZpcmFsaXR5LCBudW1fcG9zdHMsIG51bV91bmlxdWVfdXNlcnMsIHJvb3Rfc2NvcmUuCgpZb3VyIEpTT04gUHJlZGljdGlvbjo=)Youareanalyzingasocialmediaconversationtreetopredictitsfinalgrowthmetrics\.Youwillbeprovidedwiththefirst<WINDOW\>ofasocialmediaconversationthread\.Thisconversationthreadcontinuedafterthis\.Basedonthecontent,timing,andstructuralpatternsoftheseearlyinteractions,predicttheFINALstateofthisconversationtreewhenitreachessaturation\(thatis,stopsgrowing\)\.METRICDEFINITIONS:\-max\_width:Maximumnumberofrepliesatanysingledepthlevel\.\-max\_depth:Maximumdepthoftheconversationtree\.\-structural\_virality:Averagedistancebetweenallpairsofnodes\(Low=Broadcast,High=Viral\)\.\-num\_posts:Totalnumberofpostsinthefinalconversation\(root\+replies\)\.\-num\_unique\_users:Totalnumberofuniqueusersinthefinalconversation\.\-root\_score:Engagementscoreoftheintialrootpost\.YouwillbegivenaNESTEDJSONTREE\.Eachreplynodehastext,time\_since\_root,time\_since\_parent,andchildren\.Therootnodehastext,timestamp,andchildren\.OUTPUTFORMAT:ReturnasingleJSONobject\.Donotincludemarkdownformatting\.\{"max\_width":<number\>,"max\_depth":<number\>,"structural\_virality":<number\>,"num\_posts":<number\>,"num\_unique\_users":<number\>,"root\_score":<number\>\}EARLYCONVERSATIONTREE:<EARLY\_CONVERSATION\_TREE\_JSON\>INSTRUCTIONS:Usingonlytheconversationtreeaboveasevidence,predicttheFINALvalues\.DoNOTincludeanyexplanationorextrafields\.OutputexactlyoneJSONobjectwiththesekeysonly:max\_width,max\_depth,structural\_virality,num\_posts,num\_unique\_users,root\_score\.YourJSONPrediction:

Below is a minimal example of the input that populates<EARLY\_CONVERSATION\_TREE\_JSON\>when only the root post has been observed\.

Example Input — Root\-Only Tree[⬇](data:text/plain;base64,ewogICJ0ZXh0IjogIk52aWRpYSBwYXJ0bmVyIHNheXMgaXQgY2FuIGN1dCBkYXRhIGNlbnRlciBlbmVyZ3kgdXNlIGJ5IDUwJSBhcyBBSSBib29tIHN0cmFpbnMgcG93ZXIgZ3JpZCIsCiAgInRpbWVzdGFtcCI6ICIyMDI0LTA4LTI3IDExOjE3OjE4IiwKICAiY2hpbGRyZW4iOiBbXQp9)\{"text":"Nvidiapartnersaysitcancutdatacenterenergyuseby50%asAIboomstrainspowergrid","timestamp":"2024\-08\-2711:17:18","children":\[\]\}

###### Few\-Shot RAG Prompt

The few\-shot prompt augments the zero\-shot template withkkretrieved examples \(defaultk=4k=4\), selected via embedding\-based nearest\-neighbour search over the training set\. Each placeholder<RETRIEVED\_EXAMPLE\_TREE\_ii\_JSON\>and<RETRIEVED\_EXAMPLE\_ii\_TARGETS\_JSON\>is replaced by the corresponding retrieved tree and its ground\-truth final metrics;<TARGET\_TREE\_JSON\>is the query conversation\.

Prompt Template — Few\-Shot RAG[⬇](data:text/plain;base64,WW91IGFyZSBhbmFseXppbmcgYSBzb2NpYWwgbWVkaWEgY29udmVyc2F0aW9uIHRyZWUgdG8gcHJlZGljdCBpdHMgZmluYWwgZ3Jvd3RoIG1ldHJpY3MuIFlvdSB3aWxsIGJlIHByb3ZpZGVkIHdpdGggdGhlIGZpcnN0IDxXSU5ET1c+IG9mIGEgc29jaWFsIG1lZGlhIGNvbnZlcnNhdGlvbiB0aHJlYWQuIFRoaXMgY29udmVyc2F0aW9uIHRocmVhZCBjb250aW51ZWQgYWZ0ZXIgdGhpcy4KCkJhc2VkIG9uIHRoZSBjb250ZW50LCB0aW1pbmcsIGFuZCBzdHJ1Y3R1cmFsIHBhdHRlcm5zIG9mIHRoZXNlIGVhcmx5IGludGVyYWN0aW9ucywgcHJlZGljdCB0aGUgRklOQUwgc3RhdGUgb2YgdGhpcyBjb252ZXJzYXRpb24gdHJlZSB3aGVuIGl0IHJlYWNoZXMgc2F0dXJhdGlvbiAodGhhdCBpcywgc3RvcHMgZ3Jvd2luZykuCgpNRVRSSUMgREVGSU5JVElPTlM6Ci0gbWF4X3dpZHRoOiBNYXhpbXVtIG51bWJlciBvZiByZXBsaWVzIGF0IGFueSBzaW5nbGUgZGVwdGggbGV2ZWwuCi0gbWF4X2RlcHRoOiBNYXhpbXVtIGRlcHRoIG9mIHRoZSBjb252ZXJzYXRpb24gdHJlZS4KLSBzdHJ1Y3R1cmFsX3ZpcmFsaXR5OiBBdmVyYWdlIGRpc3RhbmNlIGJldHdlZW4gYWxsIHBhaXJzIG9mIG5vZGVzIChMb3cgPSBCcm9hZGNhc3QsIEhpZ2ggPSBWaXJhbCkuCi0gbnVtX3Bvc3RzOiBUb3RhbCBudW1iZXIgb2YgcG9zdHMgaW4gdGhlIGZpbmFsIGNvbnZlcnNhdGlvbiAocm9vdCArIHJlcGxpZXMpLgotIG51bV91bmlxdWVfdXNlcnM6IFRvdGFsIG51bWJlciBvZiB1bmlxdWUgdXNlcnMgaW4gdGhlIGZpbmFsIGNvbnZlcnNhdGlvbi4KLSByb290X3Njb3JlOiBFbmdhZ2VtZW50IHNjb3JlIG9mIHRoZSBpbnRpYWwgcm9vdCBwb3N0LgoKRWFjaCBjb252ZXJzYXRpb24gaXMgcmVwcmVzZW50ZWQgYXMgYSBORVNURUQgSlNPTiBUUkVFLgoKSSB3aWxsIHNob3cgeW91IHNldmVyYWwgZXhhbXBsZXMgb2YgZWFybHkgdHJlZXMgYW5kIHRoZWlyIGZpbmFsIG1ldHJpY3MsIHRoZW4gYXNrIHlvdSB0byBwcmVkaWN0IGZvciBhIG5ldyB0cmVlLgoKPT09PT09PT09PT09PT09PT09PT0KRkVXLVNIT1QgRVhBTVBMRVMKPT09PT09PT09PT09PT09PT09PT0KCi0tLSBFeGFtcGxlIDEgLS0tCkVhcmx5IGNvbnZlcnNhdGlvbiB0cmVlOgo8UkVUUklFVkVEX0VYQU1QTEVfVFJFRV8xX0pTT04+CkZpbmFsIG1ldHJpY3M6CjxSRVRSSUVWRURfRVhBTVBMRV8xX1RBUkdFVFNfSlNPTj4KCi0tLSBFeGFtcGxlIDIgLS0tCkVhcmx5IGNvbnZlcnNhdGlvbiB0cmVlOgo8UkVUUklFVkVEX0VYQU1QTEVfVFJFRV8yX0pTT04+CkZpbmFsIG1ldHJpY3M6CjxSRVRSSUVWRURfRVhBTVBMRV8yX1RBUkdFVFNfSlNPTj4KClsuLi4gayBleGFtcGxlcyB0b3RhbCAuLi5dCgo9PT09PT09PT09PT09PT09PT09PQpORVcgQ09OVkVSU0FUSU9OIChQUkVESUNUIFRISVMpCj09PT09PT09PT09PT09PT09PT09CkVhcmx5IGNvbnZlcnNhdGlvbiB0cmVlOgo8VEFSR0VUX1RSRUVfSlNPTj4KRmluYWwgbWV0cmljczo=)Youareanalyzingasocialmediaconversationtreetopredictitsfinalgrowthmetrics\.Youwillbeprovidedwiththefirst<WINDOW\>ofasocialmediaconversationthread\.Thisconversationthreadcontinuedafterthis\.Basedonthecontent,timing,andstructuralpatternsoftheseearlyinteractions,predicttheFINALstateofthisconversationtreewhenitreachessaturation\(thatis,stopsgrowing\)\.METRICDEFINITIONS:\-max\_width:Maximumnumberofrepliesatanysingledepthlevel\.\-max\_depth:Maximumdepthoftheconversationtree\.\-structural\_virality:Averagedistancebetweenallpairsofnodes\(Low=Broadcast,High=Viral\)\.\-num\_posts:Totalnumberofpostsinthefinalconversation\(root\+replies\)\.\-num\_unique\_users:Totalnumberofuniqueusersinthefinalconversation\.\-root\_score:Engagementscoreoftheintialrootpost\.EachconversationisrepresentedasaNESTEDJSONTREE\.Iwillshowyouseveralexamplesofearlytreesandtheirfinalmetrics,thenaskyoutopredictforanewtree\.====================FEW\-SHOTEXAMPLES====================\-\-\-Example1\-\-\-Earlyconversationtree:<RETRIEVED\_EXAMPLE\_TREE\_1\_JSON\>Finalmetrics:<RETRIEVED\_EXAMPLE\_1\_TARGETS\_JSON\>\-\-\-Example2\-\-\-Earlyconversationtree:<RETRIEVED\_EXAMPLE\_TREE\_2\_JSON\>Finalmetrics:<RETRIEVED\_EXAMPLE\_2\_TARGETS\_JSON\>\[\.\.\.kexamplestotal\.\.\.\]====================NEWCONVERSATION\(PREDICTTHIS\)====================Earlyconversationtree:<TARGET\_TREE\_JSON\>Finalmetrics:

### Appendix GModality Ablation Details

In Section[5\.5](https://arxiv.org/html/2606.27539#S5.SS5), to evaluate the specific contributions of each modality in the MMG\-PopNet model, we conducted comprehensive ablation studies\. By removing one input source at a time, we measured the relative Mean Squared Error \(MSE\) increase over the full model\. The full MMG\-PopNet architecture combines four distinct information sources: text \(node\-level post text\), image \(root\-post visual content\), topology \(reply\-tree structure\), and temporal \(node\-level timing features\)\.

Implementation Details:To isolate each modality, we modified the model architecture and inputs as follows \(w/o means “without”\):

- •w/o Text \(Semantic Ablation\):and its corresponding projection layer are completely disabled\. To maintain architectural compatibility with the downstream Graph Neural Network \(GNN\), the text embedding for every node is replaced with a zero\-initialized vector of the exact same dimensionality\. Transformer parameters are frozen and excluded from optimization\. This isolates the contribution of actual semantic content\. By passing zero\-vectors into an otherwise unchanged text feature slot, any performance drop directly reflects the loss of node\-level language cues\.
- •w/o Image \(Visual Ablation\):The CLIP vision encoder is bypassed and all raw image pixels are ignored\. However, the graph\-level image fusion slot is preserved\. Every cascade is instead assigned an identical, learnable “dummy” image embedding that is optimized alongside the rest of the network parameters\.
- •w/o Topology \(Structural Ablation\):We completely disable the neighbor aggregation mechanism of the bidirectional GraphSAGE network by ignoring the reply edges entirely\. Instead of message passing, the network treats the cascade as a disconnected set of independent nodes\. To ensure any performance drop is strictly due to the loss of structural connectivity and not a reduction in parameter capacity, the GraphSAGE layers are substituted with standard MLPs\. These MLPs are applied independently to each node and strictly match the layer count and hidden feature dimensions of the original GNN\. This isolates the value of explicit cascade interactions\. Because the replacement MLPs retain the exact same per\-node transformation capacity as the original model, this ablation cleanly measures the impact of structural message passing\.
- •w/o Temporal \(Timing Ablation\):The explicit node\-level cascade timing features, specifically time since root and time since parent, are stripped from the input feature matrix before entering the GNN\. The nodes are processed using only their text features and topological connections\. This specifically targets localized reaction speeds and the pace of the cascade\. Global posting metadata may remain, but the precise timing of individual user interactions is entirely removed\.

Tables[20](https://arxiv.org/html/2606.27539#A7.T20),[21](https://arxiv.org/html/2606.27539#A7.T21),[22](https://arxiv.org/html/2606.27539#A7.T22)report the detailed modality\-to\-target sensitivity analysis across the r/Gaming and r/Futurology datasets\. Table[20](https://arxiv.org/html/2606.27539#A7.T20)aggregates the MSE performance for both datasets to highlight the trade\-offs between semantic and structural signals\. Table[21](https://arxiv.org/html/2606.27539#A7.T21)and[22](https://arxiv.org/html/2606.27539#A7.T22)provide the detailed mean and standard deviation across three random seed runs \(42, 1042, and 2042\) for r/Gaming and r/Futurology, respectively\.

The detailed results confirm the main trend in Section[5\.5](https://arxiv.org/html/2606.27539#S5.SS5)\. The full MMG\-PopNet model consistently achieves the best average MSE across every target, confirming that each modality contributes valuable predictive information\. However, the sensitivity to specific modalities varies distinctly between the two communities

Onr/Gaming, temporal information is the most influential signal for most structural and participation targets\. Temporal information is the most influential signal for structural and participation metrics\. Removing temporal features causes the sharpest performance degradation forMax Width\(21\.2%\),Unique Users\(21\.0%\),Size\(18\.7%\),Max Depth, andStructural Virality\. This indicates that the pace and timing of early responses are critical for predicting how gaming discussions grow and branch\. Conversely, removing image features has the smallest overall effect, indicating that root visual content plays a more modest, complementary role in this specific community\.

Onr/Futurology, the ablation effects are more distributed across modalities\. Temporal and topology features remain important for structural targets, but text and image removals also produce visible degradation across several targets\. This suggests that r/Futurology popularity prediction depends on a broader mixture of semantic, temporal, and structural signals, rather than being dominated by a single modality\.

Across both datasets, text is consistently the most important modality forLike Score\. Removing textual semantics increases the averageLike ScoreMSE from 4\.300 to 4\.876 on r/Gaming and from 3\.397 to 4\.088 on r/Futurology\. These results are consistent with the text\-centered nature of the dataset platforms of the benchmark\., where discussion content is primarily text based\. This suggests a broader hypothesis that the dominant modality may shift with platform design\. On platforms where images or videos are the primary medium of interaction, visual features may play a larger role in predicting engagement and cascade growth\.

Table 20:Modality\-to\-Target Sensitivity Analysis\.MSE performance of the full MMG\-PopNet model versus single\-modality ablations across the r/Gaming and r/Futurology datasets\. This table isolates the contribution of each modality across the six distinct predictive targets, highlighting the trade\-offs between semantic and structural signals\. Lower is better\. The best value isbolded; the second\-best isunderlined\.TaskModelr/Gamingr/FuturologyAvg205090Avg3090180AvgMax WidthMMG\-PopNet1\.0540\.6520\.5090\.7391\.0930\.6200\.3480\.6870\.713w/o Text1\.0790\.6740\.4870\.7471\.2020\.6170\.3190\.7130\.730w/o Image1\.0560\.6530\.5090\.7391\.1450\.6500\.3400\.7110\.725w/o Temporal1\.0831\.1840\.6410\.9691\.1270\.7110\.4410\.7590\.864w/o Topology1\.0610\.7060\.5510\.7731\.2020\.6930\.3770\.7570\.765Max DepthMMG\-PopNet0\.2660\.1960\.1520\.2050\.4130\.2710\.2020\.2950\.250w/o Text0\.2690\.1960\.1500\.2050\.4570\.2710\.2050\.3110\.258w/o Image0\.2670\.1940\.1550\.2050\.4220\.2800\.2040\.3020\.254w/o Temporal0\.2690\.2960\.1630\.2430\.4180\.2850\.2110\.3050\.274w/o Topology0\.2770\.2140\.1920\.2280\.4400\.2950\.2180\.3180\.273SizeMMG\-PopNet1\.2720\.7840\.5980\.8851\.7030\.9490\.5241\.0590\.972w/o Text1\.2930\.8130\.5620\.8891\.8990\.9620\.5091\.1231\.006w/o Image1\.2720\.7840\.6030\.8861\.7830\.9950\.5201\.0990\.993w/o Temporal1\.3021\.3910\.7451\.1461\.7491\.0860\.6491\.1611\.154w/o Topology1\.2670\.8070\.6440\.9061\.8541\.0500\.5481\.1511\.028StructuralViralityMMG\-PopNet0\.0760\.0570\.0450\.0590\.1430\.0970\.0720\.1040\.082w/o Text0\.0780\.0580\.0450\.0600\.1580\.0980\.0720\.1090\.085w/o Image0\.0770\.0560\.0440\.0590\.1470\.1000\.0720\.1060\.083w/o Temporal0\.0760\.1170\.0450\.0790\.1460\.1020\.0730\.1070\.093w/o Topology0\.0820\.0680\.0610\.0700\.1520\.1070\.0780\.1120\.091UniqueUsersMMG\-PopNet1\.1700\.7300\.5560\.8191\.3340\.7470\.4070\.8290\.824w/o Text1\.1970\.7640\.5250\.8291\.4850\.7560\.3920\.8780\.853w/o Image1\.1720\.7350\.5610\.8221\.4020\.7890\.4060\.8660\.844w/o Temporal1\.1981\.3310\.6971\.0751\.3800\.8630\.5160\.9200\.997w/o Topology1\.1660\.7710\.5970\.8451\.4670\.8300\.4360\.9110\.878LikeScoreMMG\-PopNet4\.6594\.2374\.0034\.3004\.5273\.0322\.6323\.3973\.848w/o Text5\.2155\.0164\.3974\.8765\.3473\.6803\.2374\.0884\.482w/o Image4\.7144\.3274\.0234\.3554\.7833\.2712\.6743\.5763\.965w/o Temporal4\.7265\.0414\.2844\.6834\.5973\.4543\.1663\.7394\.211w/o Topology4\.7024\.3484\.1134\.3874\.8123\.2922\.7283\.6113\.999

Table 21:Modality\-to\-Target Sensitivity Analysis on r/Gaming\.MSE performance of full MMG\-PopNet model versus single\-modality ablations on the r/Gaming dataset\. Results are averaged across three random seeds: 42, 1042, and 2042, and are reported as mean±\\pmstandard deviation\. Lower is better\. The best value isbolded; the second\-best isunderlined\.TaskModel205090AvgMax WidthMMG\-PopNet1\.054±\\pm0\.0080\.652±\\pm0\.0130\.509±\\pm0\.0140\.739±\\pm0\.009w/o Text1\.079±\\pm0\.0150\.674±\\pm0\.0030\.487±\\pm0\.0300\.747±\\pm0\.013w/o Image1\.056±\\pm0\.0120\.653±\\pm0\.0060\.509±\\pm0\.0220\.739±\\pm0\.006w/o Temporal1\.083±\\pm0\.0081\.184±\\pm0\.7080\.641±\\pm0\.0070\.969±\\pm0\.234w/o Topology1\.061±\\pm0\.0080\.706±\\pm0\.0040\.551±\\pm0\.0120\.773±\\pm0\.007Max DepthMMG\-PopNet0\.266±\\pm0\.0030\.196±\\pm0\.0030\.152±\\pm0\.0060\.205±\\pm0\.001w/o Text0\.269±\\pm0\.0020\.196±\\pm0\.0030\.150±\\pm0\.0010\.205±\\pm0\.000w/o Image0\.267±\\pm0\.0040\.194±\\pm0\.0030\.155±\\pm0\.0050\.205±\\pm0\.002w/o Temporal0\.269±\\pm0\.0020\.296±\\pm0\.1530\.163±\\pm0\.0020\.243±\\pm0\.052w/o Topology0\.277±\\pm0\.0030\.214±\\pm0\.0070\.192±\\pm0\.0050\.228±\\pm0\.001SizeMMG\-PopNet1\.272±\\pm0\.0120\.784±\\pm0\.0040\.598±\\pm0\.0230\.885±\\pm0\.010w/o Text1\.293±\\pm0\.0160\.813±\\pm0\.0080\.562±\\pm0\.0280\.889±\\pm0\.012w/o Image1\.272±\\pm0\.0140\.784±\\pm0\.0050\.603±\\pm0\.0200\.886±\\pm0\.004w/o Temporal1\.302±\\pm0\.0141\.391±\\pm0\.8200\.745±\\pm0\.0041\.146±\\pm0\.277w/o Topology1\.267±\\pm0\.0130\.807±\\pm0\.0110\.644±\\pm0\.0220\.906±\\pm0\.011StructuralViralityMMG\-PopNet0\.076±\\pm0\.0030\.057±\\pm0\.0000\.045±\\pm0\.0030\.059±\\pm0\.001w/o Text0\.078±\\pm0\.0010\.058±\\pm0\.0020\.045±\\pm0\.0030\.060±\\pm0\.001w/o Image0\.077±\\pm0\.0030\.056±\\pm0\.0020\.044±\\pm0\.0030\.059±\\pm0\.002w/o Temporal0\.076±\\pm0\.0010\.117±\\pm0\.0960\.045±\\pm0\.0010\.079±\\pm0\.032w/o Topology0\.082±\\pm0\.0040\.068±\\pm0\.0040\.061±\\pm0\.0030\.070±\\pm0\.001UniqueUsersMMG\-PopNet1\.170±\\pm0\.0100\.730±\\pm0\.0110\.556±\\pm0\.0190\.819±\\pm0\.010w/o Text1\.197±\\pm0\.0160\.764±\\pm0\.0050\.525±\\pm0\.0240\.829±\\pm0\.011w/o Image1\.172±\\pm0\.0160\.735±\\pm0\.0060\.561±\\pm0\.0170\.822±\\pm0\.003w/o Temporal1\.198±\\pm0\.0071\.331±\\pm0\.8200\.697±\\pm0\.0071\.075±\\pm0\.273w/o Topology1\.166±\\pm0\.0120\.771±\\pm0\.0060\.597±\\pm0\.0160\.845±\\pm0\.008LikeScoreMMG\-PopNet4\.659±\\pm0\.0234\.237±\\pm0\.1384\.003±\\pm0\.0484\.300±\\pm0\.064w/o Text5\.215±\\pm0\.0175\.016±\\pm0\.0384\.397±\\pm0\.0094\.876±\\pm0\.013w/o Image4\.714±\\pm0\.0284\.327±\\pm0\.1364\.023±\\pm0\.0554\.355±\\pm0\.044w/o Temporal4\.726±\\pm0\.0205\.041±\\pm0\.9884\.284±\\pm0\.0234\.683±\\pm0\.335w/o Topology4\.702±\\pm0\.0504\.348±\\pm0\.0514\.113±\\pm0\.0424\.387±\\pm0\.047

Table 22:Modality\-to\-Target Sensitivity Analysis on r/Futurology\.MSE performance of full MMG\-PopNet model versus single\-modality ablations on the r/Futurology dataset\. Results are averaged across three random seeds: 42, 1042, and 2042, and are reported as mean±\\pmstandard deviation\. Lower is better\. The best value isbolded; the second\-best isunderlined\.TaskModel3090180AvgMax WidthMMG\-PopNet1\.093±\\pm0\.0130\.620±\\pm0\.0200\.348±\\pm0\.0080\.687±\\pm0\.009w/o Text1\.202±\\pm0\.0200\.617±\\pm0\.0050\.319±\\pm0\.0160\.713±\\pm0\.011w/o Image1\.145±\\pm0\.0150\.650±\\pm0\.0100\.340±\\pm0\.0080\.711±\\pm0\.007w/o Temporal1\.127±\\pm0\.0220\.711±\\pm0\.0120\.441±\\pm0\.0130\.759±\\pm0\.005w/o Topology1\.202±\\pm0\.0090\.693±\\pm0\.0110\.377±\\pm0\.0220\.757±\\pm0\.005Max DepthMMG\-PopNet0\.413±\\pm0\.0010\.271±\\pm0\.0030\.202±\\pm0\.0050\.295±\\pm0\.003w/o Text0\.457±\\pm0\.0120\.271±\\pm0\.0020\.205±\\pm0\.0010\.311±\\pm0\.004w/o Image0\.422±\\pm0\.0050\.280±\\pm0\.0060\.204±\\pm0\.0020\.302±\\pm0\.004w/o Temporal0\.418±\\pm0\.0090\.285±\\pm0\.0060\.211±\\pm0\.0040\.305±\\pm0\.004w/o Topology0\.440±\\pm0\.0050\.295±\\pm0\.0070\.218±\\pm0\.0030\.318±\\pm0\.003SizeMMG\-PopNet1\.703±\\pm0\.0150\.949±\\pm0\.0170\.524±\\pm0\.0101\.059±\\pm0\.012w/o Text1\.899±\\pm0\.0550\.962±\\pm0\.0050\.509±\\pm0\.0131\.123±\\pm0\.021w/o Image1\.783±\\pm0\.0220\.995±\\pm0\.0190\.520±\\pm0\.0061\.099±\\pm0\.010w/o Temporal1\.749±\\pm0\.0391\.086±\\pm0\.0220\.649±\\pm0\.0061\.161±\\pm0\.009w/o Topology1\.854±\\pm0\.0251\.050±\\pm0\.0140\.548±\\pm0\.0191\.151±\\pm0\.007StructuralViralityMMG\-PopNet0\.143±\\pm0\.0020\.097±\\pm0\.0010\.072±\\pm0\.0020\.104±\\pm0\.000w/o Text0\.158±\\pm0\.0030\.098±\\pm0\.0020\.072±\\pm0\.0010\.109±\\pm0\.001w/o Image0\.147±\\pm0\.0020\.100±\\pm0\.0020\.072±\\pm0\.0010\.106±\\pm0\.000w/o Temporal0\.146±\\pm0\.0050\.102±\\pm0\.0030\.073±\\pm0\.0020\.107±\\pm0\.002w/o Topology0\.152±\\pm0\.0050\.107±\\pm0\.0030\.078±\\pm0\.0010\.112±\\pm0\.001UniqueUsersMMG\-PopNet1\.334±\\pm0\.0120\.747±\\pm0\.0140\.407±\\pm0\.0120\.829±\\pm0\.010w/o Text1\.485±\\pm0\.0400\.756±\\pm0\.0040\.392±\\pm0\.0100\.878±\\pm0\.015w/o Image1\.402±\\pm0\.0170\.789±\\pm0\.0170\.406±\\pm0\.0040\.866±\\pm0\.009w/o Temporal1\.380±\\pm0\.0300\.863±\\pm0\.0090\.516±\\pm0\.0060\.920±\\pm0\.007w/o Topology1\.467±\\pm0\.0210\.830±\\pm0\.0130\.436±\\pm0\.0190\.911±\\pm0\.001LikeScoreMMG\-PopNet4\.527±\\pm0\.0033\.032±\\pm0\.0392\.632±\\pm0\.0963\.397±\\pm0\.035w/o Text5\.347±\\pm0\.0503\.680±\\pm0\.0383\.237±\\pm0\.0294\.088±\\pm0\.035w/o Image4\.783±\\pm0\.0663\.271±\\pm0\.0072\.674±\\pm0\.0743\.576±\\pm0\.024w/o Temporal4\.597±\\pm0\.0803\.454±\\pm0\.0753\.166±\\pm0\.0153\.739±\\pm0\.003w/o Topology4\.812±\\pm0\.0553\.292±\\pm0\.0552\.728±\\pm0\.0973\.611±\\pm0\.004

### Appendix HQualitative Case Study

This qualitative case study examines MMG\-PopNet predictions for the r/Futurology dataset under a 180\-minute observation window, focusing on theSizetarget, which measures the final number of nodes in a social cascade\. The scatter plot compares predictedSizeagainst actualSize\. Both axes are shown in log scale and are displayed using powers of ten, such as10110^\{1\},10210^\{2\}, and10310^\{3\}, to make the heavy\-tailed popularity distribution easier to interpret\. In this view, points near the red dashed diagonal indicate more accurate predictions, while points above or below the diagonal indicate over\-prediction or under\-prediction\.

For the qualitative analysis, several cascades are selected from different regions of the scatter plot to provide a wide range of examples, including accurate predictions, over\-predictions, and under\-predictions\. These selected examples are shown as colored nodes on the scatter plot\. For each selected case, interpretability is applied to the root post content\. Specifically, GradSAM\[[5](https://arxiv.org/html/2606.27539#bib.bib87)\]token explanations are used to highlight influential text tokens, and Grad\-CAM\[[48](https://arxiv.org/html/2606.27539#bib.bib86),[24](https://arxiv.org/html/2606.27539#bib.bib88)\]image explanations are used to highlight influential image regions\. This helps provide intuition about how the root post’s text and image content may have contributed to the predicted cascadeSize, while also showing that final popularity depends on additional temporal and interaction dynamics captured by the full model\.

![Refer to caption](https://arxiv.org/html/2606.27539v1/figure/qual2.png)Figure 7:The top example shows a relatively accurate high\-popularity prediction, where the predictedSizeis close to the actualSize\. GradSAM highlights root\-post tokens such as “discover”, “quantum states”, and “longer” suggesting that the model attends to scientific novelty and breakthrough\-oriented language\. Grad\-CAM emphasizes parts of the laboratory image, including the people and equipment, which may provide visual cues of scientific credibility\. The bottom example shows an under\-predicted cascade, where the actualSizeis much larger than the predictedSize\. GradSAM highlights tokens such as “exclusive,” “satellite,” “megacity,” and “well underway,” while Grad\-CAM focuses on several regions of the satellite image\. This case suggests that although the root post contains visually and textually salient signals, the model may underestimate posts whose later popularity is driven by broader public interest in large\-scale infrastructure or geopolitical topics\.![Refer to caption](https://arxiv.org/html/2606.27539v1/figure/qual3.png)Figure 8:The top example shows an over\-predicted cascade, where the predictedSizeis larger than the actualSize\. GradSAM highlights root\-post tokens such as “says,” “facial recognition,” “helped,” “fugitive,” and “match,” suggesting that the model attends to crime, surveillance, and authority\-related language\. Grad\-CAM highlights multiple regions across the facial\-recognition image, which may reinforce the post’s technology and public\-safety framing\. The bottom example shows a relatively accurate low\-popularity prediction, where the predictedSizeis close to the actualSize\. GradSAM highlights tokens such as “floating,” “wind turbine,” and “hour,” while Grad\-CAM focuses on parts of the offshore turbine and surrounding scene\. This comparison suggests that root\-post content can provide meaningful cues forSizeprediction, but attention to salient technology\-related terms does not always translate into high cascade growth\.![Refer to caption](https://arxiv.org/html/2606.27539v1/figure/qual4.png)Figure 9:The highlighted example shows an under\-predicted cascade, where the actualSizeis much larger than the predictedSize\. GradSAM highlights root\-post tokens such as “bezos,” looking,” death,” and science of aging,” suggesting that the model attends to the named entity and the longevity\-related framing of the post\. Grad\-CAM focuses strongly on the face of Jeff Bezos, indicating that the visual explanation is concentrated on him in the image\. This case suggests that the root post contains salient celebrity and science\-related cues, but the model still underestimates the eventual discussion volume, possibly because later cascade growth is driven by broader public debate around wealth, longevity, and aging beyond the root content alone\.
### Appendix ILimitations\.

Despite the unified design of MMG\-Pop and the strong empirical performance of MMG\-PopNet, this work has several limitations\.

First, the benchmark is constructed from Bluesky and Reddit, which provide diverse but still incomplete coverage of social media ecosystems\. Platform\-specific moderation policies, recommendation algorithms, user demographics, and interaction norms can substantially affect popularity dynamics\. Therefore, conclusions drawn from these datasets may not fully generalize to platforms such as X/Twitter, TikTok, Instagram, YouTube, or private messaging communities\.

Second, our formulation represents social cascades primarily as tree\-structured reply or interaction graphs\. This abstraction captures explicit propagation paths, but it may omit broader network exposure effects, algorithmic ranking effects, cross\-platform diffusion \(e\.g\., getting high engagement on one social platform due to a viral event on second social platform\), and unobserved impressions\. A post may become popular not only because of its visible reply tree, but also because of recommendation systems, external sharing, creator reputation, or coordinated amplification that is not directly observable in the collected cascade\.

Third, although MMG\-PopNet jointly models text, image, temporal, and structural signals, the available modalities are uneven across platforms and communities\. Root visual content contributes modestly in our experiments, but this may partly reflect dataset composition rather than the intrinsic value of visual signals\. Similarly, richer video, audio and user\-history features are not fully modeled\. Future extensions should consider broader media types and more complete user\-context features while carefully protecting user privacy\.

Fourth, the work of popularity prediction has the potential of being misused by nefarious parties to identify the signals that provide them the highest social engagement to spread hateful or toxic messages or content on the social media\. So, one has to be careful and mindful in using such techniques to ensure well being of all\.

Similar Articles

Towards Robust Federated Multimodal Graph Learning under Modality Heterogeneity

arXiv cs.LG

This paper proposes FedMPO, a robust federated multimodal graph learning method that addresses modality heterogeneity and missing modalities through topology-aware cross-modal generation, missing-aware expert routing, and reliability-aware aggregation, achieving performance gains on multiple datasets.

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory

arXiv cs.CL

Introduces SMMBench, a benchmark to evaluate multimodal agents' ability to retrieve, align, and compose evidence scattered across independently originated sources like conversations, tables, and documents. Experiments show current systems struggle with this source-distributed memory composition task.