Tag
Introduces CT-ΔBench, a benchmark for longitudinal 3D CT imaging difference reporting with vision-language models, along with change-aware metrics and a baseline model DeltaMed.
Introduces TradeVerse, a benchmark built from WTO meeting minutes that tests LLMs on longitudinal political trade negotiations, including tasks like predicting product categories, identifying responding countries, and generating statements.
This paper proposes a multidimensional text analysis approach combining Japanese NLP metrics and statistical methods to evaluate changes in risk disclosure quality, applied to Japan's 2019 corporate disclosure reforms. The analysis of 19,770 firm-year observations reveals complex shifts such as increased volume accompanied by decreased readability.
This paper introduces transition-aware best-of-N sampling, a training-free method for generating longitudinal chest X-ray reports by encoding changes between prior and current examinations using set-to-set distance metrics.
This paper analyzes longitudinal conversational trajectories of Bing Copilot users and compares them with WildChat data, finding that individual user habits are sticky and that WildChat overrepresents power users, challenging static views of user-LLM interactions.