Tag
The paper introduces MASA, a method that uses frozen multimodal large language models to break the self-referential loop in wild test-time adaptation by providing structured semantic descriptions for more reliable adaptation.
This paper proposes TASCO, a framework for test-time adaptation in LLMs that incorporates local stability into confidence-based optimization to improve reasoning accuracy and token efficiency without updating model parameters.
CASTER introduces a gradient-free test-time adaptation method for frozen models, using affine statistics transport and a certificate to decide when to apply adaptation for improved performance without model updates.
Safin-1 presents a memory-routing architecture that embeds safety as an internal, evolving state in foundation models, enabling adaptive safety improvements without external constraints.
PETA proposes a parameter-efficient framework for test-time adaptation in virtual screening, improving performance by updating only LayerNorm parameters during inference.
This paper proposes ReNC, a method for open-world test-time adaptation that leverages neural collapse as a structural prior to filter out-of-distribution samples and refine prototypes for reliable model adaptation.
PRISM is a training-free test-time adaptation method that reverses low-rank affine noise distortions in audio-text models using frozen text prototypes, showing significant improvements under severe acoustic noise.
CoAdapt-GUI is a test-time adaptation framework for mobile GUI agents that jointly adapts workflow context and policy, improving performance on unseen-app benchmarks like AndroidWorld-Generalization and AndroidWorld Plus.
Presents Self-Geometry, a plug-and-play test-time adaptation pipeline that enforces explicit multi-view geometric constraints using 2D pixel correspondences to improve geometrically consistent 3D vision foundation models.
This paper introduces a montage-agnostic encoder for surface-EMG gesture decoding that maintains recognition accuracy across recording sessions without recalibration, and shows that feature-statistic alignment at test time improves adaptation on NinaPro DB6.
This paper presents a theoretical framework for self-poisoning in adaptive out-of-distribution detection, proving a sharp threshold for collapse and a certified label-free calibration method that severs the feedback loop. It also provides impossibility results for distinguishing drift from contamination without labels.
OPERA proposes a multi-agent ensemble framework that treats expert weight assignment as an offline policy learning problem for universal biomedical image analysis, enabling test-time adaptation without retraining and consistently improving performance across 9 datasets and 30+ baselines.
Black-Mamba introduces a test-time adaptive forecasting architecture that uses accumulated surprisal to selectively update memory only upon evidence of distribution drift, achieving efficient adaptation on non-stationary time series.
AdaJEPA introduces an adaptive latent world model that continuously updates during test-time via closed-loop model predictive control, significantly improving planning success under distribution shift.
Introduces Reward-Gated Test-Time Adaptation (RG-TTA), a reinforcement learning framework that selectively applies debiasing to CLIP models based on input bias sensitivity, resolving the fairness-utility trade-off.
Proposes BP-TTA, a test-time adaptation method that handles both class imbalance and continual domain shifts by combining batch-balanced sampling with prototype-guided constraints, achieving state-of-the-art performance in dynamic streaming scenarios.
ComMem proposes complementary memory systems inspired by biological memory to improve test-time adaptation of vision-language models, outperforming state-of-the-art on 15 benchmarks.
This paper proposes a test-time adaptation approach using semi-supervised learning for AI text detection that adapts to continual distribution shifts from new LLMs, adversarial humanization, and temporal drift, outperforming state-of-the-art supervised detectors.
QGF is an RL algorithm that improves policies at test time by using a value gradient to guide a pre-trained flow policy, avoiding training-time instability while maintaining competitive performance.
Proposes Demo2Reward, a test-time prompt optimization technique for VLM reward models using a few expert demonstrations, significantly reducing false positives and improving policy learning in robotics without additional model training.