Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance

arXiv cs.AI Papers

Summary

This paper evaluates multimodal LLMs for disaster risk communication, finding that current models lack consistency across text and audio modalities, which undermines accessibility for vulnerable populations.

arXiv:2608.14651v1 Announce Type: new Abstract: Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency infrastructure fails to meet the needs of individuals with access and functional needs, including hard-of-hearing individuals, pregnant women, mothers with toddlers, and elderly individuals with dementia. Recent advancements in Artificial Intelligence (AI), especially Multi-Modal Large Language Models (MM-LLMs), demonstrate powerful capabilities to serve diverse users across text, audio, image, and video modalities within a single unified system, such as a chatbot. However, their suitability for deployment rests on a property that receives limited scrutiny, i.e., whether these systems produce consistent, actionable outputs regardless of the modality through which a user communicates. In this paper, we conduct a comprehensive analysis to understand the status of open-weight MM-LLMs using real emergency alert scenarios across four different vulnerable personas. These state-of-the-art (SOTA) models are evaluated on consistency of responses across text and audio modalities when the same task scenario is given. Findings indicate that no model achieves reliable consistency across modalities, and that performance gaps are heightened for personas with access needs, introducing modality-dependent inequity that undermines the humanitarian value of these systems. These results inform concrete design recommendations for building equitable, trustworthy, and inclusive AI tools for disaster risk communication.
Original Article
View Cached Full Text

Cached at: 08/18/26, 09:56 AM

# Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance
Source: [https://arxiv.org/html/2608.14651](https://arxiv.org/html/2608.14651)
###### Abstract

Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency infrastructure fails to meet the needs of individuals with access and functional needs, including hard\-of\-hearing individuals, pregnant women, mothers with toddlers, and elderly individuals with dementia\. Recent advancements in Artificial Intelligence \(AI\), especially Multi\-Modal Large Language Models \(MM\-LLMs\), demonstrate powerful capabilities to serve diverse users across text, audio, image, and video modalities within a single unified system, such as a chatbot\. However, their suitability for deployment rests on a property that receives limited scrutiny, i\.e\., whether these systems produce consistent, actionable outputs regardless of the modality through which a user communicates\. In this paper, we conduct a comprehensive analysis to understand the status of open\-weight MM\-LLMs using real emergency alert scenarios across four different vulnerable personas\. These state\-of\-the\-art \(SOTA\) models are evaluated on consistency of responses across text and audio modalities when the same task scenario is given\. Findings indicate that no model achieves reliable consistency across modalities, and that performance gaps are heightened for personas with access needs, introducing modality\-dependent inequity that undermines the humanitarian value of these systems\. These results inform concrete design recommendations for building equitable, trustworthy, and inclusive AI tools for disaster risk communication\.

## IIntroduction

Natural disasters represent diverse emergency situations faced by communities worldwide\. For instance, from earthquakes and storms to floods and droughts, the U\.S\. sustained 403 weather and climate disasters from 1980 to 2024 with costs over 1 billion dollars\[[23](https://arxiv.org/html/2608.14651#bib.bib5)\]\. The frequency and intensity of such events continues to rise under shifting climate conditions\[[23](https://arxiv.org/html/2608.14651#bib.bib5)\]\. Thus, to enable the public better prepare, respond to, and recover from disasters, there has been a noticeable emphasis on effective disaster risk communication\[[24](https://arxiv.org/html/2608.14651#bib.bib6)\]\. This includes early warning devices and alerting systems that can monitor, forecast, and communicate risks to enable individuals, communities, governments, businesses and others to take timely action to reduce disaster risks before hazardous events occur\[[14](https://arxiv.org/html/2608.14651#bib.bib18)\]\. Such risk communication systems are increasingly being modernized nowadays, including experimenting with Artificial Intelligence \(AI\) techniques\[[29](https://arxiv.org/html/2608.14651#bib.bib4),[9](https://arxiv.org/html/2608.14651#bib.bib1)\]\.

While the general public faces significant risks during these disasters, there are certain groups of people who are disproportionately affected by them and may require additional accommodation in the design of risk communication systems\. Vulnerability refers to the characteristics of a person or group and their situation that influences their ability to prepare, cope, respond, and recover from a disaster\[[21](https://arxiv.org/html/2608.14651#bib.bib7)\]\. The vulnerable population is less likely to have access to resources and may encounter prejudice, discrimination, and stigma due to their socio\-economic status, race/ethnicity, gender, age, cognitive and/or physical ability, etc\[[3](https://arxiv.org/html/2608.14651#bib.bib8)\]\. According to the U\.S\. Federal Emergency Management Agency \(FEMA\), people with access and functional needs make up to 43% of the U\.S\. population and may increase as a result of a disaster\[[8](https://arxiv.org/html/2608.14651#bib.bib20)\]\. People with ‘access and functional needs’ include individuals with disabilities, individuals with limited English proficiency, individuals with limited access to transportation, individuals with limited access to financial resources, older adults, etc\.\[[5](https://arxiv.org/html/2608.14651#bib.bib19)\]\.

Risk communication tools to assist and support the vulnerable population require provisions of information accessibility that can create and embed effective means of interaction for these individuals with various needs and preferences\[[16](https://arxiv.org/html/2608.14651#bib.bib3),[5](https://arxiv.org/html/2608.14651#bib.bib19)\]\. While AI systems have shown promising applications for disaster risk management, in order to better understand the situations around us for risk communication, AI needs to interpret and reason about the diverse interaction needs of individuals with access and functional needs\[[1](https://arxiv.org/html/2608.14651#bib.bib27)\]\. To bridge this gap between human perception and AI, multi\-modal interfaces seek to leverage natural human capabilities to communicate via speech, gesture, touch, facial expression, and other modalities, resulting in more refined interaction\[[28](https://arxiv.org/html/2608.14651#bib.bib28)\]\. One such advancement is multi\-modal Large Language Models \(MM\-LLMs\)\[[37](https://arxiv.org/html/2608.14651#bib.bib13)\], which are able to take inputs, process, and generate outputs across text, image, audio, and video\.

As we progress towards employing the MM\-LLMs based AI technologies for public communication, their evaluation plays a crucial role in development and deployment\. Despite the claim of achieving a unified multimodal system design, these systems often demonstrate misalignment in understanding and response generation across input modalities\[[40](https://arxiv.org/html/2608.14651#bib.bib36)\]\. Fig\.[1](https://arxiv.org/html/2608.14651#S1.F1)illustrates this problem\. If not carefully designed and deployed, they are prone to hallucinations, are fragile to adversarial input, and fail to produce appropriate responses in high\-stake domains, such as neglecting the urgency of the risk and accessibility constraints in disaster risk communication\. Prior work has suggested that some of these issues correlate with the inconsistency of LLMs, which is generally defined as their tendency to generate low\-confidence responses or conflicting responses when the same input prompt is resampled\[[33](https://arxiv.org/html/2608.14651#bib.bib30)\]\. Accurately estimating MM\-LLM consistency is important during critical applications such as disaster risk communication and affects the user’s level of trust in the AI\-powered systems\[[17](https://arxiv.org/html/2608.14651#bib.bib29)\]\.

To support the diverse communication requirements of people with access and functional needs, MM\-LLMs based systems must provide multiple modes of interaction effectively: voice, text, or video, while maintaining consistent interpretation across these modalities during communication\. This paper proposes an evaluation framework that serves as the first step in the design of responsible AI systems based on consistent MM\-LLMs for disaster risk communication\. Specifically, we make the following contributions in this paper:

1. 1\.This study introduces a consistency evaluation framework for open\-weight MM\-LLMs, incorporating both quantitative and qualitative analyses, with an emphasis on supporting multimodal interaction requirements in real\-world disaster risk communication scenarios\.
2. 2\.In addition to assessing current capabilities, this work evaluates variations in MM\-LLM behavior across a spectrum of functional and access needs, thus measuring inclusivity in model responses if deployed for risk communication\. These findings have broad implications for advancing equitable, trustworthy, and human\-centered AI systems that enhance accessibility and resilience in critical communication contexts\.

The rest of the paper is structured as follows\. Section[II](https://arxiv.org/html/2608.14651#S2)presents the related work, followed by Section[III](https://arxiv.org/html/2608.14651#S3), a description of the methodology\. Section[IV](https://arxiv.org/html/2608.14651#S4)discusses the results of empirical analysis, followed by conclusion in Section[V](https://arxiv.org/html/2608.14651#S5)\.

![Refer to caption](https://arxiv.org/html/2608.14651v1/problem_updated.png)Figure 1:Illustration of cross\-modal inconsistency in AudioFlamingo and SALMONN model where the same emergency query submitted via text and audio modalities receives a detailed, actionable text response and a generic\.
## IIRelated Work

### II\-ARisk Communication and Supporting Technologies

Extensive literature exists for risk communication during disasters that often emphasizes that information needs to be actionable, accessible, and appropriately tailored to the needs of diverse individuals\[[3](https://arxiv.org/html/2608.14651#bib.bib8),[24](https://arxiv.org/html/2608.14651#bib.bib6),[29](https://arxiv.org/html/2608.14651#bib.bib4)\]\. The International Telecommunication Union \(ITU\) has identified telecommunications and early warning systems as foundational infrastructure for all phases of disaster risk reduction and management, emphasizing that warning systems must reach all populations at risk with clear and usable guidance\[[14](https://arxiv.org/html/2608.14651#bib.bib18)\]\. Despite this imperative, traditional broadcast\-based emergency alert systems such as sirens, television crawls, and Wireless Emergency Alerts have well\-documented limitations in reaching individuals with sensory, cognitive, or linguistic barriers\[[30](https://arxiv.org/html/2608.14651#bib.bib9)\]\. With the emergence of conversational AI and LLMs, there are new possibilities for interactive, personalized risk communication\[[29](https://arxiv.org/html/2608.14651#bib.bib4)\]\. Early work explored ChatGPT’s potential for disaster prevention information dissemination, science education, and emergency response support, noting its rapid availability and natural language reasoning as distinctive advantages over static alert systems\[[35](https://arxiv.org/html/2608.14651#bib.bib10)\]\. More recently, LLMs have been applied to classify crisis information from social media streams, monitor infrastructure during active disasters, and generate structured warning messages grounded in official guidelines\[[18](https://arxiv.org/html/2608.14651#bib.bib11),[6](https://arxiv.org/html/2608.14651#bib.bib12)\]\.

In this work, we systematically examine the assumption that an AI system will produce reliable, consistent outputs regardless of how a user interacts with it\. For populations with access and functional needs, who may be unable to use text\-based interfaces under disaster conditions, whether the same system would provide equivalent guidance to them is a gap that still remains\. We address this directly by evaluating current open\-weight MM\-LLMs that serve as backbone of risk communication tools across text and audio modalities\.

### II\-BMM\-LLMs based Systems

MM\-LLMs based systems take a user query as input and generate the response as output in the interaction modality specified by the user\. The system generally comprises three core components: theModality Encoderthat encodes inputs from diverse modalities;LLM Backbonethat does zero\-shot generalization, Chain\-of\-Thought \(CoT\), and instruction following; andModality Generatorthat produces outputs in distinct modalities\[[38](https://arxiv.org/html/2608.14651#bib.bib32)\]\. This architecture allows MM\-LLMs to process and generate content spanning text, image, audio, and video within a single unified system, representing a significant improvement from earlier task\-specific models that operated on single modalities\[[37](https://arxiv.org/html/2608.14651#bib.bib13)\]\. However, a critical design challenge in MM\-LLMs remains modality alignment, which ensures representations from different input modalities are mapped into a shared semantic space that the LLM backbone can reason over uniformly\. Recent advances address this challenge by employing lightweight adapter modules, such as Q\-Formers and linear projections, that bridge pre\-trained modality\-specific encoders to a frozen or partially\-tuned LLM\[[32](https://arxiv.org/html/2608.14651#bib.bib33)\]\. This approach has resulted in a diverse ecosystem of models, including vision\-based models such as BLIP\-2, Flamingo, MiniGPT\-4and LLaVA, as well as audio\-vision\-based models such as Video\-LLaMA\[[39](https://arxiv.org/html/2608.14651#bib.bib34)\]and SALMONN\[[25](https://arxiv.org/html/2608.14651#bib.bib43)\], which pair pre\-trained speech encoders like Whisper with LLM backbones via alignment layers\. More recent MM\-LLMs are now capable of processing any combination of modality without cascaded pipelines\[[38](https://arxiv.org/html/2608.14651#bib.bib32)\]\. These models employ multi\-stage progressive alignment during training by first aligning individual modality pairs before jointly fine\-tuning across modalities\. Some examples include Qwen3\-Omni\[[34](https://arxiv.org/html/2608.14651#bib.bib42)\]and Audio Flamingo\[[11](https://arxiv.org/html/2608.14651#bib.bib41)\]\.

While the above examples of models show promising strategies to improve efficiency, there exists a risk of cross modality interference, where training on one modality degrades performance or consistency in another\[[15](https://arxiv.org/html/2608.14651#bib.bib14)\]\. This is seen in omni\-modal LLMs where model attends to dominant modalities, typically text, while producing degraded or inconsistent outputs for audio inputs\[[15](https://arxiv.org/html/2608.14651#bib.bib14)\]\. This structural asymmetry is the core motivation for our evaluation that despite being presented as unified systems, the underlying training dynamics of MM\-LLMs provide opportunities to expect inconsistency across modalities in response generation, particularly for underrepresented input types such as audio data in safety\-critical contexts like disaster risk communication\.

### II\-CEvaluation of MM\-LLMs

Current evaluation of MM\-LLMs relies primarily on benchmarking MM foundational models\[[17](https://arxiv.org/html/2608.14651#bib.bib29)\]\. Advances in both text\-based and audio/vision\-based MM\-LLMs have led to a new set of benchmarks designed to track and guide their development efficiently\[[31](https://arxiv.org/html/2608.14651#bib.bib40)\]\. These benchmarks can be broadly categorized based on text, vision, and audio\-based LLMs\. Text\-based LLMs are evaluated on their reasoning capabilities, correct answer generation, and possible mitigation of bias\[[10](https://arxiv.org/html/2608.14651#bib.bib45)\]\. Vision LLMs focus on attack, hallucination, ethical, and cultural aspects or input modalities, i\.e\., visual or language perspective\[[27](https://arxiv.org/html/2608.14651#bib.bib46)\]\. However, benchmarks built on modality or task\-specific datasets are increasingly inadequate for capturing general capabilities\[[17](https://arxiv.org/html/2608.14651#bib.bib29)\]\. Recent research explores human\-like any\-to\-any modality conversion as a step toward artificial general intelligence\[[38](https://arxiv.org/html/2608.14651#bib.bib32)\], that would require further comprehensive benchmarking\. However, these existing works still neglect understanding whether MM\-LLMs provide consistent responses across modalities for the same domain\-specific task, such as the facilitation of risk communication for user queries\.

## IIIMethodology: Consistency Evaluation Framework

Our framework \(Fig\.[2](https://arxiv.org/html/2608.14651#S3.F2)\) comprises four components:Stakeholder Identificationbased on FEMA guidelines,Persona Creationgrounded in official alert types,Multimodal LLMsthat take these persona inputs paired as text and audio prompts, andMultifaceted Evaluationdone through manual analysis, semantic similarity metrics, and factual overlap scoring\.

![Refer to caption](https://arxiv.org/html/2608.14651v1/x1.png)Figure 2:Framework overview illustrating the four components of: Stakeholder Identification; Persona Creation; MM\-LLM Inference, and Evaluation\.### III\-AIdentification of Stakeholders

To identify populations vulnerable during disaster events, we conducted a critical review of relevant resources from the U\.S\. Centers for Disease Control and Prevention \(CDC\), identifying a formal group classified as individuals with “access and functional needs\.” According to the CDC, this category refers to individuals who may require additional assistance due to temporary conditions or permanent conditions that may limit their ability to respond effectively in emergencies\[[5](https://arxiv.org/html/2608.14651#bib.bib19)\]\.

Notably, individuals with access and functional needs are not required to have a formal diagnosis or medical evaluation\. The CDC identifies several groups that may be disproportionately affected during emergencies, including children \(with or without disabilities\), pregnant women, older adults, individuals with physical, sensory, intellectual, developmental, cognitive, or mental disabilities, those with chronic health conditions or pharmacological dependencies, people with limited English proficiency, and individuals facing financial, transportation, or legal barriers to emergency preparedness and recovery\. From this list of individuals, we selected four stakeholder groups based on two eligibility criteria: \(1\) the existence of prior literature on the group and \(2\) high vulnerability during disaster scenarios\. The final stakeholder groups considered in this study include:Pregnant Women, Mothers with Toddlers, Hard of Hearing People, Elderly Individuals with Dementia\. Disasters have been linked to potential adverse outcomes and impacts for pregnant women\[[13](https://arxiv.org/html/2608.14651#bib.bib15)\]\. Their increased rates of preterm birth, pregnancy complications, and maternal mortality, with documented cases from Hurricane Katrina and other events showing disproportionate health impacts on this group\[[4](https://arxiv.org/html/2608.14651#bib.bib16)\]\. Their mobility limitations and time\-sensitive medical needs make rapid, actionable emergency communication especially critical\. It is widely recognized that emergency plans should account for the unique needs of mothers and children, with the youngest infants being the most vulnerable to the effects of natural disasters\[[19](https://arxiv.org/html/2608.14651#bib.bib17)\]\. Mothers with toddlers face increased risk of their children getting acute illness, developmental disabilities, being underweight and the mothers developing maternal anxiety\[[36](https://arxiv.org/html/2608.14651#bib.bib21)\]\. The Great East Japan Earthquake and Tsunami has been described as one of the worst natural disasters in Japanese history that affected the physical and socioenvironmental conditions of the local communities including mothers with infants and preschool\-aged children\[[20](https://arxiv.org/html/2608.14651#bib.bib22)\]\. Hard of hearing individuals represent one of the most directly impacted groups in multimodal risk communication research since the standard emergency alerting systems \(sirens, broadcast audio, Wireless Emergency Alerts with sound\) are either inaccessible or partially inaccessible to them, making alternative modalities precisely the focus of this study\. This trend is illustrated by the preemptive evacuation of hard of hearing individuals during Hurricane Rita where hard of hearing individuals evacuated preemptively due to hurricane announcements being exclusively accessible through specific television stations, translators being unavailable at shelters, and information from FEMA and the Red Cross not being conveyed in sign language or any other accessible manner\[[26](https://arxiv.org/html/2608.14651#bib.bib23)\], illustrating how communication gaps could materially alter evacuation behavior and risk exposure\. Finally, research shows adverse impact of disasters on memory and awareness of elderly individuals with dementia\[[7](https://arxiv.org/html/2608.14651#bib.bib25)\]\. Some older individuals impacted by hurricanes experienced a transient decrease in working memory lasting 6 months after the disaster, with a subsequent return to pre\-disaster levels by the 14\-month follow\-up period\[[7](https://arxiv.org/html/2608.14651#bib.bib25)\]\. Similarly, those affected by the earthquake in Turkey experienced significant declines in memory, daily functioning, speech, and overall cognitive abilities\[[12](https://arxiv.org/html/2608.14651#bib.bib24)\]\. Based on this evidence, our study scope includes four vulnerable user groups as stakeholders\.

### III\-BPersona Creation

Based on the identified stakeholders, we create personas and communication scenarios corresponding to each\. These scenarios are derived from real emergency alerts that are being sent out in coordination with the FEMA’s national system for local alerting called the Integrated Public Alert & Warning System \(IPAWS\)111https://www\.fema\.gov/emergency\-managers/practitioners/integrated\-public\-alert\-warning\-system, which delivers verified emergency information to the public via multiple channels, including mobile phones through Wireless Emergency Alerts \(WEA\)222https://www\.weather\.gov/wrn/wea360, broadcast media through the Emergency Alert System, and NOAA Weather Radio\. Using samples of such alerts, we constructed a list of 40 distinct prompts that reflect various alert types and disaster scenarios\.

For instance, considering ‘Flash Flood Warning’, the prompt template is “I am a X, and I just received a flash flood warning\. What should I do during this time to stay safe?” and similarly, when considering ‘Flash Flood Alert’, the prompt template is “I am a X, and I just received a flash flood alert\. What should I do during this time to stay safe?” X represents a persona; we have summarized the four different personas, along with their sample queries, in Table[I](https://arxiv.org/html/2608.14651#S3.T1)\. For the audio prompts one person from the team recorded their voice to generate the query which was then supplied to the MM\-LLM\.\(Note: we will share the full set of prompts and result logs with the accepted, camera\-ready paper\.\)

TABLE I:Illustrative prompts for validating consistency across MM\-LLMs including the Wireless Emergency Alerts sent out by National Weather Service during emergency situations\.For scoping purpose, it is important to distinguish between two layers of accessibility challenge that MM\-LLMs must address for users with access and functional needs\. The first isperceptual accessibilitywhich is whether a user can physically receive and decode the signal delivered by an interface \(e\.g\., whether a hard of hearing individual can perceive an audio alert or whether a person with dementia can retain spoken instructions\)\. The second isinformational consistencywhich is whether the system produces equivalent content regardless of the modality through which a query is submitted\. This study evaluates the latter\. Our methodology treats audio as a clean digital signal passed programmatically to the MM\-LLM, and therefore does not simulate the perceptual barriers such as signal frequency, clarity, or amplification requirements that hard of hearing individuals face when receiving audio in real\-world environments\. Rather, we evaluate whether the model itself introduces asymmetry in the information it provides across modalities, which is also necessary to meet information\-access expectations\. A system that produces perfectly consistent outputs across text and audio modalities is still inaccessible to a hard of hearing individual if the audio channel itself cannot be perceived\. For the current study’s scope, we leave this exploration for future work, including realistic acoustic degradation, hearing aid signal processing simulations, and user studies with hard of hearing participants\.

### III\-CMultimodal LLMs: State\-of\-the\-Art Open\-Weight Models

We evaluate three state\-of\-the\-art open\-weight MM\-LLMs across 40 different prompts for the four personas summarized in Table[I](https://arxiv.org/html/2608.14651#S3.T1)\. The MM\-LLMs that we employed are:

#### III\-C1AudioFlamingo

Audio Flamingo 3 is based on a 7B language model and the LLaVA architecture\. This model is trained on a unified AF\-Whisper audio encoder based on Whisper that handles understanding beyond speech recognition\. Audio Flamingo 3 is able to handle three distinct signal types in audio: sound, music, and speech\[[11](https://arxiv.org/html/2608.14651#bib.bib41)\]\.

#### III\-C2Qwen

Qwen3\-Omni is a natively end\-to\-end, multilingual model capable of processing text, images, audio, and video while delivering real\-time streaming responses in both text and natural speech\. The model incorporated several architectural upgrades that significantly improved performance and efficiency across MM applications spanning audio, image, video, and audio\-visual tasks\[[34](https://arxiv.org/html/2608.14651#bib.bib42)\]\.

#### III\-C3SALMONN

SALMONN \(Speech, Audio, Language, and Music Open Neural Network\) is designed to process and reason over diverse audio inputs, including speech, environmental sounds, and music\. It integrates pretrained audio encoders with a LLM through alignment layers, enabling capabilities such as audio captioning, speech comprehension, and audio\-based reasoning\. SALMONN is particularly notable for its ability to generalize across heterogeneous audio domains, making it suitable for applications that require contextual understanding of real\-world audio signals\[[25](https://arxiv.org/html/2608.14651#bib.bib43)\]\.

### III\-DMultifaceted Evaluation

In order to understand the consistency of responses across text and audio modalities, we employed existing approaches \(1\) Using semantic similarity metrics and \(2\) Factual consistency metrics to evaluate LLM consistency\. For each prompt, 4 personas yielded 160 responses, given the 40 prompts in our study\. This resulted in 160 responses for text and 160 responses for audio\. For these response pairs, the semantic similarity scores were calculated using four existing metrics: Bert \(sBert\), BLEU \(sBLEU\), Rouge \(sRouge\), and USE \(sUSE\)\. We calculated sBert using the BERTScore Python package, sBLEU using the NLTK Python package, sRouge using the rouge\-score Python package, and sUSE using the universal\-sentence\-encoder model on Kaggle\[[33](https://arxiv.org/html/2608.14651#bib.bib30)\]\. Inspired by FactSumm\[[22](https://arxiv.org/html/2608.14651#bib.bib2)\], factual consistency is calculated through a factual overlap score that measures how much content from the text response is preserved in the audio response\. The text response is treated as the source and the audio response as the candidate, extracting named entities and disaster\-relevant action terms \(e\.g\.,evacuate,shelter\) selected based on common disaster terminology\. The factual overlap score is then computed as the harmonic mean of the entity precision score, which is the proportion of candidate keywords and entities also present in the source, with the ROUGE\-L lexical overlap score between the two responses\.

## IVResults and Discussion

As defined in Section[III](https://arxiv.org/html/2608.14651#S3), we evaluated semantic similarities across response pairs using previously published methods and further assessed factual consistency using the factual overlap metric, which combines entity\-level fact overlap with lexical similarity to measure whether safety\-critical facts present in a text response are preserved in the corresponding audio response\. Table[II](https://arxiv.org/html/2608.14651#S4.T2)reports these metrics across the three models\.

### IV\-AModel Performance Analysis

As shown in Table[II](https://arxiv.org/html/2608.14651#S4.T2), when comparing each of the four metrics to evaluate the similarity of response pairs: Bert, BLEU, Rouge, and USE; Qwen achieved the strongest performance across all the semantic similarity metrics, indicating the highest level of consistency between text and audio responses\. The high S\-BERT score \(0\.896\) indicates that the underlying meaning was preserved more effectively across modalities\. SALMONN demonstrated moderate performance, suggesting limited overlap in wording and possible scope for improvement\. Flamingo, on the other hand, demonstrated the lowest performance across all metrics, indicating weaker alignment in meaning across modalities\. Among all models, S\-USE and S\-BERT are consistently higher than S\-BLEU and S\-ROUGE, indicating that while models often preserve core meaning, they vary substantially in wording and structure across modalities\.

TABLE II:Performance of the three different MM\-LLMs across modalities evaluated using semantic consistency metricsModelS\-ROUGES\-BLEUS\-USES\-BERTFlamingo0\.22270\.06520\.61890\.4948SALMONN0\.24860\.13180\.69640\.6253Qwen0\.34080\.34260\.89550\.8961
### IV\-BError Analysis

The examples in Table[III](https://arxiv.org/html/2608.14651#S4.T3)highlight systematic inconsistencies between text and audio responses across modalities, with direct implications for the deployment of MM\-LLMs in disaster\-affected communities\. In both cases, the text responses provide detailed, actionable guidance tailored to the specific hazard \(e\.g\., avoiding dust exposure, staying hydrated, and limiting physical exertion\), whereas the corresponding audio responses are lower quality, generic and, in some cases, fail to address the scenario meaningfully\. For instance, in the dust advisory case, the audio response omits critical protective measures such asfiltration,indoor sheltering, andmask useand instead, offers vague instructions about following safety protocols\. In a humanitarian context, where affected individuals may lack access to follow\-on resources, a caregiver, or reliable connectivity, such vagueness is not merely a quality issue but a safety risk\. Elderly individuals with dementia, in particular, are unlikely to seek clarification or consult secondary sources under crisis situations\[[2](https://arxiv.org/html/2608.14651#bib.bib26)\]and the first response they receive is often the only one they act on\. Another concern observed from these responses is the semantic misalignment across modalities as the core meaning drifts across text and audio\. In the case of a mother with a toddler \(see Fig\.[3](https://arxiv.org/html/2608.14651#S4.F3)\), the audio output produces a response entirely unrelated to the user’s condition, representing a major drift between the persona\-adapting capabilities across both modalities\.

These patterns reflect a broader challenge in deploying AI\-based communication tools for social good that the current open\-weight MM\-LLMs are optimized for general capability benchmarks rather than for the reliability that disaster risk communication contexts of underserved populations demand\. The modality\-dependent degradation observed here, where the audio interface, most likely to be used by individuals with visual impairments, low literacy, or motor disabilities, consistently receives less comprehensive guidance, represents a direct failure to meet that standard\. Humanitarian technology deployments of MM\-LLMs must therefore treat cross\-modal consistency not as a secondary evaluation criterion, but as a core requirement alongside accuracy and factual correctness\.

TABLE III:Examples of responses by AudioFlamingo that demonstrated poor consistency when prompted in different modalities for the same query scenario\. \(Bluehighlights the situation;Greenindicates helpful tips generated only in text modality\.\)![Refer to caption](https://arxiv.org/html/2608.14651v1/x2.png)Figure 3:Cross\-modal inconsistency in AudioFlamingo for a mother with a toddler persona given a flash flood emergency where the text response provides specific, actionable guidance \(stay indoors, move to higher ground\), while the audio response remains vague and omits critical safety instructions\.
### IV\-CModality Type Analysis

For practical usage in risk communication where equitable access to emergency information is fundamental, the consistency gap revealed across all three models is of great concern\. Based on Table[II](https://arxiv.org/html/2608.14651#S4.T2), current open\-weight MM\-LLMs preserve high\-level intent while failing to maintain alignment in actionable detail and semantics\. Qwen’s near parity between S\-USE \(0\.896\) and S\-BERT \(0\.896\) reflects relatively robust semantic and factual grounding, however, its substantially lower S\-BLEU and S\-ROUGE scores indicate that consistency is achieved through paraphrasing rather than faithful content reproduction\. This matters in safety\-critical scenarios where reformulation of key guidance such as evacuation routes or hazard avoidance steps can materially alter the usefulness of a response\. On the other hand, scores from SALMONN suggest partial loss of fine\-grained meaning across modalities\. AudioFlamingo performs worst with both lexical divergence and semantic drift, underscoring the need for evaluation metrics that explicitly capture action\-level agreement and factual omission\. In risk communication, where a missed instruction to stay indoors, seek higher ground, or avoid downed power lines can directly determine survival outcomes, modality\-level paraphrasing is not a stylistic concern but a humanitarian design failure that future MM\-LLMs must explicitly address\.

### IV\-DPersona Type Analysis

Table[IV](https://arxiv.org/html/2608.14651#S4.T4)presents factual overlap scores broken down by persona across all three models\. The hard of hearing persona scores lowest in Qwen \(0\.397\) and AudioFlamingo \(0\.237\), suggesting greatest factual divergence across modalities\. Conversely, this persona scores highest in SALMONN \(0\.306\), revealing that models handle this persona differently across architectures\. A contrasting pattern also appears for the pregnant woman persona where it is strongest persona for Qwen \(0\.452\) and AudioFlamingo \(0\.335\) but the weakest for SALMONN \(0\.296\), indicating that persona\-specific factual consistency is not a stable property across MM\-LLMs\. More broadly, responses involving condition\-specific needs show lower factual overlap scores, introducing a form of modality\-dependent bias where individuals with specialized needs receive incomplete or less actionable information\. These results show that current open\-weight MM\-LLMs lack robustness in handling persona\-specific constraints, required for designing accessible systems in disaster contexts\.

### IV\-EImplications, Limitations, and Future Work

These findings carry direct implications for MM\-LLM based system design and humanitarian deployment\. The multimodal capability of accepting diverse inputs is different from multimodal reliability, i\.e\., a system that produces inconsistent outputs across modalities provides unequal service to users whose access needs determine how they communicate, encoding modality\-dependent inequity into the technology\. Cross\-modal consistency should therefore be treated as a priority evaluation criterion alongside accuracy, with training pipelines incorporating explicit consistency objectives and fine\-tuning on domain knowledge for relevant user personas\[[25](https://arxiv.org/html/2608.14651#bib.bib43),[34](https://arxiv.org/html/2608.14651#bib.bib42)\]\.

At the deployment level, the persona\-level variation in our results highlights the need for inclusive evaluation protocols that assess model behavior across the needs of diverse users rather than aggregate benchmarks alone\. A model that performs well on average but fails for hard of hearing individuals or elderly users with dementia is not fit for humanitarian use\.

The limitations of this study point to at least five future directions\. First, all audio prompts were recorded by a single human speaker, which does not capture natural variation in accent, pace, pitch, or age\-related vocal characteristics that MM\-LLMs encounter in the real world\. Future work should incorporate audios generated by diverse speaker profiles and expand the prompt set to confirm whether observed patterns hold across a broader and more naturalistic range of vocal inputs\. Second, the evaluation scope is relatively small and findings should be treated as indicative rather than definitive\. To address this, future work should develop a framework of semantic, factual, and information\-based metrics that capture safety\-critical omissions in a more fine\-grained manner across modalities, alongside appropriate significance tests to establish whether the reported gaps constitute robust findings\. Third, the persona\-based prompting strategy assumes explicit self\-disclosure of vulnerable status, which may not reflect naturalistic user interactions, and the absence of a neutral baseline condition prevents isolation of whether consistency gaps are specific to vulnerable persona conditioning or reflect general cross\-modal behavior\. We plan to address this through community engagement; conducting surveys and interviews with a broader set of individuals to expand, validate, and refine our persona\-based scenarios with lived experience, while also establishing neutral baseline conditions in future evaluations\. Fourth, fine\-tuning open\-weight multimodal LLMs on disaster risk communication corpora with explicit cross\-modal consistency objectives and persona\-conditioned training data represents a promising path toward more equitable and reliable humanitarian AI systems\. Fifth, disaster\-affected populations span diverse linguistic and cultural contexts that are often overlooked in model evaluation\. Improving the performance of MM\-LLMs in these contexts is essential to design and deploy accessible and equitable systems globally, particularly in low\-resource language settings and multicultural communities\.

TABLE IV:Factual overlap scores by persona across all three MM\-LLMs\.

## VConclusion

The primary goal of this study is to investigate the existing open\-weight MM\-LLMs for accessible disaster risk communication contexts\. We identify stakeholders in disaster scenarios by reviewing existing literature\. We then create personas and scenarios that represent risk communication situations involving the identified stakeholders\. These serve as prompts in audio mode and text mode and the obtained output pairs were evaluated using semantic similarity metrics\. The findings of this work reveal that open\-weight MM\-LLMs are unable to achieve complete cross\-modal consistency, with performance disparities most pronounced for personas with specialized access and functional needs\. These results provide evidence of existing accessibility limitations and motivate the development of more inclusive humanitarian AI systems\.

## VIAcknowledgment

This research was partially supported by the grant \# 2531369 from the National Science Foundation\. During the preparation of this work, the authors used Claude for Fig\. 1 preparation, code assistance, and literature review support\. After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the content of the manuscript\.

## References

- \[1\]T\. Baltrušaitis, C\. Ahuja, and L\. Morency\(2018\)Multimodal machine learning: a survey and taxonomy\.IEEE transactions on pattern analysis and machine intelligence41\(2\),pp\. 423–443\.Cited by:[§I](https://arxiv.org/html/2608.14651#S1.p3.1)\.
- \[2\]S\. A\. Bell, M\. L\. Miranda, J\. P\. Bynum, and M\. A\. Davis\(2023\)Mortality after exposure to a hurricane among older adults living with dementia\.JAMA network open6\(3\),pp\. e232043\.Cited by:[§IV\-B](https://arxiv.org/html/2608.14651#S4.SS2.p1.1)\.
- \[3\]M\. A\. Benevolenza and L\. DeRigne\(2019\)The impact of climate change and natural disasters on vulnerable populations: a systematic review of literature\.Journal of Human Behavior in the Social Environment29\(2\),pp\. 266–281\.Cited by:[§I](https://arxiv.org/html/2608.14651#S1.p2.1),[§II\-A](https://arxiv.org/html/2608.14651#S2.SS1.p1.1)\.
- \[4\]W\. M\. Callaghan, S\. A\. Rasmussen, D\. J\. Jamieson, S\. J\. Ventura, S\. L\. Farr, P\. D\. Sutton, T\. J\. Mathews, B\. E\. Hamilton, K\. R\. Shealy, D\. Brantley,et al\.\(2007\)Health concerns of women and infants in times of natural disasters: lessons learned from hurricane katrina\.Maternal and child health journal11\(4\),pp\. 307–311\.Cited by:[§III\-A](https://arxiv.org/html/2608.14651#S3.SS1.p2.1)\.
- \[5\]Centers for Disease Control and Prevention\(2021\-03\)Access and functional needs toolkit: integrating a community partner network to inform risk communication strategies\.Technical reportCDC Center for Preparedness and Response,Atlanta, GA\.External Links:[Link](https://www.cdc.gov/readiness/media/pdfs/CDC_Access_and_Functional_Needs_Toolkit_March2021.pdf)Cited by:[§I](https://arxiv.org/html/2608.14651#S1.p2.1),[§I](https://arxiv.org/html/2608.14651#S1.p3.1),[§III\-A](https://arxiv.org/html/2608.14651#S3.SS1.p1.1)\.
- \[6\]M\. Chen, Z\. Tao, W\. Tang, T\. Qin, R\. Yang, and C\. Zhu\(2024\)Enhancing emergency decision\-making with knowledge graphs and large language models\.International Journal of Disaster Risk Reduction113,pp\. 104804\.Cited by:[§II\-A](https://arxiv.org/html/2608.14651#S2.SS1.p1.1)\.
- \[7\]C\. Fahmy and M\. N\. Alme\(2026\)Effects of natural disasters on the cognitive state of older adults with dementia: a scoping review\.European Geriatric Medicine,pp\. 1–11\.Cited by:[§III\-A](https://arxiv.org/html/2608.14651#S3.SS1.p2.1)\.
- \[8\]Federal Emergency Management Agency\(2024\-05\)Access and functional needs support\.Fact SheetFEMA Mass Care and Emergency Assistance,Washington, DC\.External Links:[Link](https://www.fema.gov/sites/default/files/documents/fema_access-and-functional-needs-support_fact-sheet.pdf)Cited by:[§I](https://arxiv.org/html/2608.14651#S1.p2.1)\.
- \[9\]S\. Foubert, M\. Park, K\. Pardeshi, T\. Manzini, R\. McDaniel, D\. Merrick, R\. Murphy, A\. Singh, and H\. Heidari\(2026\)Investigating the role of ai in emergency management: use cases, challenges, and opportunities\.InThe 2026 ACM Conference on Fairness, Accountability, and Transparency,pp\. 2055–2091\.Cited by:[§I](https://arxiv.org/html/2608.14651#S1.p1.1)\.
- \[10\]M\. Gao, X\. Hu, X\. Yin, J\. Ruan, X\. Pu, and X\. Wan\(2025\)Llm\-based nlg evaluation: current status and challenges\.Computational Linguistics51\(2\),pp\. 661–687\.Cited by:[§II\-C](https://arxiv.org/html/2608.14651#S2.SS3.p1.1)\.
- \[11\]A\. Goel, S\. Ghosh, J\. Kim, S\. Kumar, Z\. Kong, S\. Lee, C\. H\. Yang, R\. Duraiswami, D\. Manocha, R\. Valle,et al\.\(2025\)Audio flamingo 3: advancing audio intelligence with fully open large audio language models\.arXiv preprint arXiv:2507\.08128\.Cited by:[§II\-B](https://arxiv.org/html/2608.14651#S2.SS2.p1.1),[§III\-C1](https://arxiv.org/html/2608.14651#S3.SS3.SSS1.p1.1)\.
- \[12\]S\. Güney and Ö\. Çiçek Doğan\(2026\)Experiences of people with dementia and their family caregivers after earthquakes: a qualitative study\.Dementia25\(2\),pp\. 315–331\.Cited by:[§III\-A](https://arxiv.org/html/2608.14651#S3.SS1.p2.1)\.
- \[13\]E\. W\. Harville, L\. Beitsch, C\. K\. Uejio, S\. Sherchan, and M\. Y\. Lichtveld\(2021\)Assessing the effects of disasters and their aftermath on pregnancy and infant outcomes: a conceptual model\.International Journal of Disaster Risk Reduction62,pp\. 102415\.Cited by:[§III\-A](https://arxiv.org/html/2608.14651#S3.SS1.p2.1)\.
- \[14\]International Telecommunication Union\(2025\)The use of telecommunications/ICT for disaster risk reduction and management\.Output Report on ITU\-D Question 3/1Technical ReportD\-STG\-SG01\.03\.1\-2025,ITU\-D Study Group 1,Geneva, Switzerland\.Note:Study period 2022–2025External Links:[Link](https://www.itu.int/dms_pub/itu-d/opb/stg/D-STG-SG01.03.1-2025-PDF-E.pdf)Cited by:[§I](https://arxiv.org/html/2608.14651#S1.p1.1),[§II\-A](https://arxiv.org/html/2608.14651#S2.SS1.p1.1)\.
- \[15\]S\. Jiang, J\. Liang, J\. Wang, X\. Dong, H\. Chang, W\. Yu, J\. Du, M\. Liu, and B\. Qin\(2025\)From specific\-mllms to omni\-mllms: a survey on mllms aligned with multi\-modalities\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 8617–8652\.Cited by:[§II\-B](https://arxiv.org/html/2608.14651#S2.SS2.p2.1)\.
- \[16\]H\. Lambert, D\. Doumont, N\. Reviers, M\. Vandenbroucke, and I\. Aujoulat\(2025\)Enhancing accessibility of crisis communication to people in vulnerable circumstances\.Journal of Public Health,pp\. 1–13\.Cited by:[§I](https://arxiv.org/html/2608.14651#S1.p3.1)\.
- \[17\]P\. P\. Liang, A\. Goindani, T\. Chafekar, L\. Mathur, H\. Yu, R\. Salakhutdinov, and L\. Morency\(2024\)Hemm: holistic evaluation of multimodal foundation models\.Advances in Neural Information Processing Systems37,pp\. 42899–42940\.Cited by:[§I](https://arxiv.org/html/2608.14651#S1.p4.1),[§II\-C](https://arxiv.org/html/2608.14651#S2.SS3.p1.1)\.
- \[18\]V\. Linardos, M\. Drakaki, and P\. Tzionas\(2025\)Utilizing llms and ml algorithms in disaster\-related social media content\.GeoHazards6\(3\),pp\. 33\.Cited by:[§II\-A](https://arxiv.org/html/2608.14651#S2.SS1.p1.1)\.
- \[19\]S\. R\. Mudiyanselage, D\. Davis, E\. Kurz, and M\. Atchan\(2022\)Infant and young child feeding during natural disasters: a systematic integrative literature review\.Women and Birth35\(6\),pp\. 524–531\.Cited by:[§III\-A](https://arxiv.org/html/2608.14651#S3.SS1.p2.1)\.
- \[20\]M\. Nishihara, Y\. Nakamura, T\. Fuchimukai, and M\. Ohnishi\(2018\)Factors associated with social support in child\-rearing among mothers in post\-disaster communities\.Environmental health and preventive medicine23\(1\),pp\. 58\.Cited by:[§III\-A](https://arxiv.org/html/2608.14651#S3.SS1.p2.1)\.
- \[21\]J\. D\. Rivera\(2021\)Disaster and emergency management methods: social science approaches in application\.Taylor & Francis\.Cited by:[§I](https://arxiv.org/html/2608.14651#S1.p2.1)\.
- \[22\]H\. Shakil, Z\. Ortiz, G\. C\. Forbes, and J\. Kalita\(2024\)Utilizing gpt to enhance text summarization: a strategy to minimize hallucinations\.Procedia Computer Science244,pp\. 238–247\.Cited by:[§III\-D](https://arxiv.org/html/2608.14651#S3.SS4.p1.1)\.
- \[23\]A\. B\. Smith\(2024\)2023 us billion\-dollar weather and climate disasters in historical context\.In104th Annual AMS Meeting 2024,Vol\.104,pp\. 428624\.Cited by:[§I](https://arxiv.org/html/2608.14651#S1.p1.1)\.
- \[24\]I\. S\. Stewart\(2024\)Advancing disaster risk communications\.Earth\-Science Reviews249,pp\. 104677\.Cited by:[§I](https://arxiv.org/html/2608.14651#S1.p1.1),[§II\-A](https://arxiv.org/html/2608.14651#S2.SS1.p1.1)\.
- \[25\]C\. Tang, W\. Yu, G\. Sun, X\. Chen, T\. Tan, W\. Li, L\. Lu, Z\. Ma, and C\. Zhang\(2023\)Salmonn: towards generic hearing abilities for large language models\.arXiv preprint arXiv:2310\.13289\.Cited by:[§II\-B](https://arxiv.org/html/2608.14651#S2.SS2.p1.1),[§III\-C3](https://arxiv.org/html/2608.14651#S3.SS3.SSS3.p1.1),[§IV\-E](https://arxiv.org/html/2608.14651#S4.SS5.p1.1)\.
- \[26\]C\. Tannenbaum\-Baruchi, I\. Ashkenazi, and C\. Rapaport\(2024\)Risk inclusion of vulnerable people during a climate\-related disaster: a case study of people with hearing loss facing wildfires\.International Journal of Disaster Risk Reduction103,pp\. 104335\.Cited by:[§III\-A](https://arxiv.org/html/2608.14651#S3.SS1.p2.1)\.
- \[27\]H\. Tu, C\. Cui, Z\. Wang, Y\. Zhou, B\. Zhao, J\. Han, W\. Zhou, H\. Yao, and C\. Xie\(2024\)How many are in this image a safety evaluation benchmark for vision llms\.InEuropean Conference on Computer Vision,pp\. 37–55\.Cited by:[§II\-C](https://arxiv.org/html/2608.14651#S2.SS3.p1.1)\.
- \[28\]M\. Turk\(2014\)Multimodal interaction: a review\.Pattern recognition letters36,pp\. 189–195\.Cited by:[§I](https://arxiv.org/html/2608.14651#S1.p3.1)\.
- \[29\]A\. Urbanelli, A\. Frisiello, L\. Bruno, and C\. Rossi\(2024\)The ermes chatbot: a conversational communication tool for improved emergency management and disaster risk reduction\.International Journal of Disaster Risk Reduction112,pp\. 104792\.Cited by:[§I](https://arxiv.org/html/2608.14651#S1.p1.1),[§II\-A](https://arxiv.org/html/2608.14651#S2.SS1.p1.1)\.
- \[30\]M\. Villarreal, C\. MacPherson\-Krutsky, and M\. A\. Painter\(2025\)Barriers and best practices for inclusive emergency alerts and warnings\.International Journal of Disaster Risk Reduction125,pp\. 105581\.Cited by:[§II\-A](https://arxiv.org/html/2608.14651#S2.SS1.p1.1)\.
- \[31\]B\. Wang, X\. Zou, G\. Lin, S\. Sun, Z\. Liu, W\. Zhang, Z\. Liu, A\. Aw, and N\. Chen\(2025\)Audiobench: a universal benchmark for audio large language models\.InProceedings of the NAACL\-2025,pp\. 4297–4316\.Cited by:[§II\-C](https://arxiv.org/html/2608.14651#S2.SS3.p1.1)\.
- \[32\]S\. Wu, H\. Fei, L\. Qu, W\. Ji, and T\. Chua\(2024\)Next\-gpt: any\-to\-any multimodal llm\.InForty\-first International Conference on Machine Learning,Cited by:[§II\-B](https://arxiv.org/html/2608.14651#S2.SS2.p1.1)\.
- \[33\]X\. Wu, W\. Lin, O\. Akgul, and L\. Bauer\(2025\)Estimating llm consistency: a user baseline vs surrogate metrics\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 30518–30532\.Cited by:[§I](https://arxiv.org/html/2608.14651#S1.p4.1),[§III\-D](https://arxiv.org/html/2608.14651#S3.SS4.p1.1)\.
- \[34\]J\. Xu, Z\. Guo, H\. Hu, Y\. Chu, X\. Wang, J\. He, Y\. Wang, X\. Shi, T\. He, X\. Zhu,et al\.\(2025\)Qwen3\-omni technical report\.arXiv preprint arXiv:2509\.17765\.Cited by:[§II\-B](https://arxiv.org/html/2608.14651#S2.SS2.p1.1),[§III\-C2](https://arxiv.org/html/2608.14651#S3.SS3.SSS2.p1.1),[§IV\-E](https://arxiv.org/html/2608.14651#S4.SS5.p1.1)\.
- \[35\]Z\. Xue, C\. Xu, and X\. Xu\(2023\)Application of chatgpt in natural disaster prevention and reduction\.Natural Hazards Research3\(3\),pp\. 556–562\.Cited by:[§II\-A](https://arxiv.org/html/2608.14651#S2.SS1.p1.1)\.
- \[36\]C\. Yamazaki and H\. Nakai\(2023\)Understanding mothers’ worries about the effects of disaster evacuation on their children: a cross\-sectional study\.International journal of environmental research and public health20\(3\),pp\. 1850\.Cited by:[§III\-A](https://arxiv.org/html/2608.14651#S3.SS1.p2.1)\.
- \[37\]S\. Yin, C\. Fu, S\. Zhao, K\. Li, X\. Sun, T\. Xu, and E\. Chen\(2024\)A survey on multimodal large language models\.National Science Review11\(12\),pp\. nwae403\.Cited by:[§I](https://arxiv.org/html/2608.14651#S1.p3.1),[§II\-B](https://arxiv.org/html/2608.14651#S2.SS2.p1.1)\.
- \[38\]D\. Zhang, Y\. Yu, J\. Dong, C\. Li, D\. Su, C\. Chu, and D\. Yu\(2024\)Mm\-llms: recent advances in multimodal large language models\.Findings of the Association for Computational Linguistics: ACL 2024,pp\. 12401–12430\.Cited by:[§II\-B](https://arxiv.org/html/2608.14651#S2.SS2.p1.1),[§II\-C](https://arxiv.org/html/2608.14651#S2.SS3.p1.1)\.
- \[39\]H\. Zhang, X\. Li, and L\. Bing\(2023\)Video\-llama: an instruction\-tuned audio\-visual language model for video understanding\.InProceedings of the 2023 conference on empirical methods in natural language processing: system demonstrations,pp\. 543–553\.Cited by:[§II\-B](https://arxiv.org/html/2608.14651#S2.SS2.p1.1)\.
- \[40\]S\. Zhao, X\. Zhang, J\. Guo, J\. Hu, L\. Duan, M\. Fu, Y\. X\. Chng, G\. Wang, Q\. Chen, Z\. Xu,et al\.\(2025\)Unified multimodal understanding and generation models: advances, challenges, and opportunities\.arXiv preprint arXiv:2505\.02567\.Cited by:[§I](https://arxiv.org/html/2608.14651#S1.p4.1)\.

Similar Articles

What We are Missing in Multimodal LLM Evaluation?

arXiv cs.AI

This paper reviews current multimodal LLM evaluation benchmarks and identifies key gaps such as temporal-spatial coherence, physical world understanding, multimodal consistency, and selective attention, arguing that existing isolated-task benchmarks fail to measure true cross-modal integration.

Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs

Hugging Face Daily Papers

This paper investigates the arithmetic limitations of multimodal LLMs on multi-digit multiplication across text, image, and audio modalities, introducing a controlled benchmark and a novel 'arithmetic load' metric (C) that better predicts model accuracy than traditional step-counting methods. Results show accuracy collapses as C grows, and that performance degradation is primarily computational rather than perceptual.