Tag
RankShift is a novel in-database method for detecting and explaining categorical shifts in data streams, using Pearson scores to identify responsible categories, and it shows competitive performance against autoencoders on log datasets.
The paper introduces D^CF5, a diagnostic to predict regionwise gains in dynamic ensembling for regression tasks under distribution shift, validated with high correlation across datasets.
The paper introduces DRACP, a conformal prediction method that unifies multiple adaptation mechanisms for reliable economic forecast intervals under distribution shifts, achieving calibration with wider intervals.
This research paper investigates how data perturbations impact the accuracy and robustness of model cascades for efficient AI inference, identifying key failure modes and stressing the need for evaluation under distribution shift.
This paper introduces Regime-Conditional Verification (RCV), a lightweight wrapper that adapts off-the-shelf safety classifiers for large language models by estimating prediction correctness and detecting distribution shift without retraining.
The paper introduces self-diagnosing models that attribute model failures under distribution shift, linking uncertainty estimation with failure attribution.
This paper identifies invisible metadata traces at the pixel level as shortcuts that vision encoders exploit, leading to performance degradation under metadata distribution shifts. Mitigation strategies during and after pretraining reduce sensitivity to both targeted and unseen metadata without sacrificing downstream performance.
This paper proposes Frontier Learning, a framework that combines representations and predictions from multiple black-box and white-box pretrained models to construct a unified target-domain representation, guaranteeing performance no worse than any individual reuse baseline under distribution shift. Evaluations on visual domain adaptation and clinical mortality prediction show consistent gains over strong baselines.
Introduces CALCoDe, a post-hoc reliability layer for frozen medical vision-language models that mitigates class-tail undercoverage under clinical shift, achieving strong worst-class accepted coverage across multiple dermatology shifts and VLM backbones.
OPERA proposes a multi-agent ensemble framework that treats expert weight assignment as an offline policy learning problem for universal biomedical image analysis, enabling test-time adaptation without retraining and consistently improving performance across 9 datasets and 30+ baselines.
Proposes Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm for faithful generation that reframes post-training as token-level correctness prediction, achieving strong out-of-distribution generalization across summarization and machine translation tasks.
GRACE uses a typed semantic graph to represent persistent instructions for LLM agents, enabling scoped verification of updates to improve reliability under distribution shift. Experiments on a telecom agent harness show significant improvements in strict reliability over baselines.
Introduces PARA-PV, a physics-aware retrieval-augmented framework for photovoltaic power forecasting that uses a frozen Chronos time-series foundation model and distribution shift correction to improve accuracy and handle peak, ramping, and low-power conditions.
This paper introduces NEST, a framework using a regime-oriented mixture-of-experts to handle dataset-level distribution shifts in time series forecasting, achieving state-of-the-art on various benchmarks.
This paper presents an empirical study comparing how different neural architectures (MLPs, CNNs, RNNs, pretrained transformers) degrade under temporal distribution shift across image and text domains, finding that models exploiting localized features degrade fastest while pretrained encoders drift more gradually.
This paper investigates temporal out-of-distribution shift in deep-learning-based climate downscaling and proposes a domain-adaptive framework that combines supervised reconstruction with domain alignment to improve high-resolution climate projections under non-stationary conditions.
This paper proposes a method to construct language models that exhibit controllable generalization failures when trained with reinforcement learning, demonstrating that training success can diverge from generalization in structured ways.
Loss smoothing interpolates between source and target objectives during adaptation, preserving useful features while enabling specialization. Experiments across supervised shifts, RL, and language model fine-tuning show consistent improvements.
ComMem proposes complementary memory systems inspired by biological memory to improve test-time adaptation of vision-language models, outperforming state-of-the-art on 15 benchmarks.
This paper characterizes when conformal risk control can certify structured LLM outputs, proving impossibility bounds and analyzing certification hierarchies across different bounds. Empirical validation on six open-weight models shows that hard configurations are uncertifiable at low risk levels but practical certification is achievable at relaxed targets.