CaRGo-T: Causal Reasoning Graph-of-Thought improves Multimodal Humor Comprehension
Summary
CaRGo-T introduces a graph-based reasoning framework to model causal relationships for improving multimodal humor comprehension in vision-language models, showing performance gains on humor understanding and detection tasks.
View Cached Full Text
Cached at: 08/28/26, 03:23 AM
Paper page - CaRGo-T: Causal Reasoning Graph-of-Thought improves Multimodal Humor Comprehension
Source: https://huggingface.co/papers/2608.23172
Abstract
CaRGo-T improves multimodal humor understanding by modeling causal relationships as graph-based reasoning structures interpreted by vision-language models.
Large-scalevision-language models(VLMs) have demonstrated remarkable versatility across a wide range of multimodal tasks. However, understanding humor remains challenging because humorous content often depends on subtle interactions among entities, events, context, and implicit relationships across image and text modalities. These interactions can involve complex chains of reasoning that are difficult to capture through conventional prompting or linear chain-of-thought reasoning. In this work, we propose CaRGo-T (Causal ReasoningGraph-of-Thought), a reasoning framework that represents the causal and contextual relationships underlyingmultimodal humoras a lightweight graph-based reasoning structure. The graph is serialized into a code-based representation generated by a VLM, which can subsequently be interpreted by the same or a different VLM to produce the final prediction inzero-shotorin-context learningsettings. We evaluate CaRGo-T on humor understanding and humor detection across four datasets spanning diverse forms of comedic content, including satire, sarcasm, and memes. Experiments with state-of-the-art commercial and open-source VLMs show that CaRGo-T consistently improves performance over existing reasoning-based baselines, achieving gains of approximately 1-20% on humor understanding and 1-3% on humor detection. Further analysis usingmutual informationindicates that the reasoning representations produced by CaRGo-T contain more information relevant to the target output than those generated by baseline reasoning approaches. Code is available at https://github.com/abhi1nandy2/CaRGo-T.
View arXiv pageView PDFGitHub0Add to collection
Get this paper in your agent:
hf papers read 2608\.23172
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.23172 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.23172 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.23172 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes
This paper introduces MAR-12, a framework using Vision-Language Models and multi-angle reasoning to detect and explain harmful humor in memes, achieving state-of-the-art accuracy on PrideMM and Memotion datasets.
Constraint-Anchored Reasoning Traces
Proposes CART, a neuro-symbolic framework that interleaves natural language reasoning steps with symbolic constraint assertions to detect and correct errors early in chain-of-thought traces for multimodal LLMs. Reduces snowball rate from 65% to 14% and improves accuracy on multiple benchmarks.
TCAR-Gen: Temporal Graph Retrieval with Evidence Fusion for Knowledge-Grounded Generation
TCAR-Gen proposes a framework combining query-conditioned graph neural networks, temporal evidence fusion, and chain-of-trees reasoning for temporal graph retrieval in knowledge-grounded generation. It achieves improved recall on the Victorian Crime Diaries benchmark across multiple query types.
Causal Reasoning with Bipartite Graphical Causal Models
The paper proposes bipartite graphical causal models (BGCMs) to resolve ambiguities in causal interventions for systems at equilibrium with cyclic dependencies, generalizing existing frameworks like causal Bayesian networks and structural causal models.
UniCAR-RL: Seeing Better before Thinking Deeper in Visual Mathematics
UniCAR-RL introduces a reinforcement learning framework that decouples perception and reasoning to improve multimodal large language models' visual mathematical reasoning without annotation.