Tag
This paper studies confidence estimation and selective prediction for financial named entity recognition under domain shift, evaluating BERT and LoRA-tuned Qwen models to enhance reliability across different input distributions like SEC filings and social media.
This paper investigates prompt ensembling and domain shift in agrifood vision-language models, introducing Prompt-based Inconsistency Detection (PID) to enhance reliability by using prompt disagreement as an uncertainty proxy.
This article presents a GitHub repository for reproducing a paper on acoustic UAV detection in battlefield scenarios, addressing challenges like noise, domain shift, and weak labels. The repository provides code for the method and evaluation protocol, with synthetic data for testing.
Introduces DALMA, a probabilistic representation learning framework that uses biological supervision to improve cross-center generalization of MALDI-TOF mass spectrometry models for clinical microbiology tasks like microbial identification and antimicrobial resistance prediction.
This paper introduces SULAND v2, a refined RGB surface landmine detection dataset and benchmark for UAV/UGV-based surveys, addressing annotation errors and domain-shift evaluation in object detection.
This paper proposes a unified framework for continual learning in LLMs, disentangling change along space (new domains) and time (data drift). It evaluates various methods including prompting, supervised learning, reinforcement learning, and context compression under realistic sequential settings.
Proposes BP-TTA, a test-time adaptation method that handles both class imbalance and continual domain shifts by combining batch-balanced sampling with prototype-guided constraints, achieving state-of-the-art performance in dynamic streaming scenarios.
Domain Arithmetic (DART) proposes a one-shot adaptation method for Vision-Language-Action models under environmental shifts using weight vector arithmetic and subspace alignment, requiring only a single demonstration.
This paper introduces Afrispeech Semantics, a benchmark for evaluating audio language models on semantic reasoning tasks including entailment, consistency, plausibility, accent drift, and accent restraint across diverse domains and accents.