Tag
The article explores the philosophical problem of Plato's Cave in the context of LLMs, proposing an experiment to compare how different conversational regimes—reconstructive versus perturbation-sensitive—might yield measurable differences in interaction behavior.
This paper introduces PRISM, a perturbation-based method for spatially resolved interpretability of large language models, adapting neuroimaging subtraction analysis to transformers and applying it in parallel to post-stroke aphasia patients to recover shared phonemic-favoring dissociations.
Introduces DECAF, a method that decomposes perturbation responses into evidence, contradiction, and fragility components, improving interpretability over raw response magnitude and achieving strong results across vision benchmarks.
Introduces AndroidReality, a perturbation-based framework for evaluating and improving the robustness of mobile agents, with a taxonomy of real-world interface perturbations and a training-free Test-Time Introspective Recovery (TTIR) mechanism.
This paper systematically studies perturbation-based continued pre-training (CPT) for improving zero-shot dialect robustness in multilingual LLMs, comparing six training conditions across German, Italian, and Arabic. It finds that character-noised CPT is the most effective general strategy and reveals that different perturbation methods induce distinct robustness mechanisms.
This paper proposes a method to identify spuriously correlated samples after model convergence by measuring prediction fragility under input perturbation, requiring no group labels or early-stopping epochs. Rebalancing training with detected samples improves worst-group accuracy on Waterbirds from 57.3% to 80.8%.
This paper proposes using generative inpainting to create photorealistic perturbations for LIME, improving the quality of explanations by avoiding out-of-distribution artifacts common in traditional occlusion methods.
Proposes JaiLIP, a method that jailbreaks vision-language models by generating imperceptible adversarial images using loss-guided perturbation, achieving high toxicity and outperforming existing methods.
Arc Institute's PerturbSpace enables high-throughput single-cell profiling of transcriptome, location, CRISPR guides, clonal relationships, and surface proteins from many samples in one day, using standard single-cell sequencing.
This paper identifies Footprint Bias in document layout analysis robustness evaluation and proposes a structure-aware auditing framework that decouples probe construction and pathway attribution, showing that small structurally targeted probes cause comparable downstream degradation to larger perturbations.
This paper introduces Shesha, a geometric stability metric that quantifies directional coherence of single-cell CRISPR perturbation responses using mean cosine similarity, revealing regulatory architecture and predicting cellular stress across 2,200+ perturbations in five CRISPR datasets.