robustness

Tag

Cards List
#robustness

Learning When Not to Listen: Selective Anti-Interference Pretraining for Language Models

arXiv cs.CL · 5h ago Cached

SPAR is a pretraining objective that uses a gated KL loss to stabilize language model predictions against irrelevant prefix text, improving robustness in long-context scenarios as demonstrated on multiple benchmarks.

0 favorites 0 likes
#robustness

CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments

arXiv cs.AI · 5h ago Cached

This paper introduces CAVEAT, a benchmark for evaluating computer-use agents in incentive-misaligned environments, revealing that agents often fail to preserve user objectives and proposes interventions to improve robustness.

0 favorites 0 likes
#robustness

Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency

arXiv cs.CL · 3d ago Cached

This paper introduces Hallucination-R1, a robustness-oriented paraphrase generation framework designed to identify and improve factual consistency in large language models by creating semantically equivalent yet challenging paraphrases.

0 favorites 0 likes
#robustness

MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs

arXiv cs.CL · 3d ago Cached

MME-Safety is a rigorously verified benchmark for evaluating the safety of Multimodal Large Language Models, featuring a four-dimensional annotation schema and a hierarchical framework to assess risk scenarios, harm severity, and modality-specific stealth levels.

0 favorites 0 likes
#robustness

Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems

arXiv cs.AI · 6d ago Cached

This paper introduces the Generic Multi-Agent Trading System (GMATS) framework to study adversarial attacks on LLM-based trading systems, showing that even simple input attacks can degrade performance but robust system designs can improve robustness.

0 favorites 0 likes
#robustness

SAM-on-the-Curve: Sharpness-Aware Mode Connectivity for Robust Weight-Space Interpolation

arXiv cs.LG · 2026-09-17 Cached

The paper introduces Sharp Mode Connectivity (SMC) to enhance the robustness of weight-space interpolation in neural networks by optimizing for flatness along entire paths, achieving significant accuracy improvements under distribution shifts.

0 favorites 0 likes
#robustness

Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation

Hugging Face Daily Papers · 2026-09-17 Cached

This paper proposes a training-adaptive convolutional sparse coding framework that leverages information bottleneck principles for robust visual representation, achieving improved performance on CIFAR and ImageNet under input perturbations.

0 favorites 0 likes
#robustness

RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue

arXiv cs.CL · 2026-09-16 Cached

RoleBreak is an open benchmark for evaluating long-horizon role-playing robustness in spoken dialogue systems, revealing gaps in current models' ability to maintain role consistency and vocal emotion over extended interactions.

0 favorites 0 likes
#robustness

Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act

arXiv cs.CL · 2026-09-16 Cached

This paper studies how reinforcement learning can lead LLM agents to learn spurious tool-use policies based on superficial cues rather than task requirements, and introduces a dense reward method to mitigate this issue.

0 favorites 0 likes
#robustness

Latent Undertow: How Ordinary Typos Break Probes

arXiv cs.CL · 2026-09-16 Cached

This paper demonstrates that ordinary typos in user inputs can significantly reduce the effectiveness of activation probes designed to detect prompt injection in LLMs, but introduces a KV-cache fork method that recovers most of the performance loss.

0 favorites 0 likes
#robustness

Fed-Equilibrium Framework for Topological Pareto Control in Robust and Fair Clinical Federated Learning

arXiv cs.LG · 2026-09-14 Cached

This paper proposes Fed-Equilibrium, a federated learning framework that balances robustness and fairness in clinical networks using topological Pareto control to ensure minority nodes achieve convergence comparable to dominant hubs.

0 favorites 0 likes
#robustness

DART: Distributional Adversarial Recurrent Training for Algorithm Learning

arXiv cs.LG · 2026-09-10 Cached

DART proposes a distributional adversarial training framework for recurrent reasoning models, improving their robustness and performance on structured problems by using a local target distribution instead of single-point supervision.

0 favorites 0 likes
#robustness

Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations

arXiv cs.LG · 2026-09-10 Cached

This paper evaluates the robustness of LLM-generated SystemVerilog assertions under semantics-preserving RTL transformations, finding that point accuracy can hide substantial instability and advocating for robustness-aware evaluation in AI-assisted hardware verification.

0 favorites 0 likes
#robustness

Which Medical Questions Deserve Rationales? Perturbation-Sensitive Selection for Robust QA

arXiv cs.CL · 2026-09-10 Cached

The paper proposes RMS-RSP, a perturbation-sensitive method for selecting medical questions to receive rationale supervision, improving robust accuracy and semantic consistency in QA systems.

0 favorites 0 likes
#robustness

Mind the Gap: Robustness Risks in PII Detection Systems

arXiv cs.LG · 2026-09-04 Cached

The article evaluates the robustness of PII detection systems under distribution shifts, identifies failure modes in different architectures, and proposes a hybrid detection pipeline with a QA-driven feedback loop for improved privacy protection.

0 favorites 0 likes
#robustness

Lost in Reordering: Structural Sensitivity of Multilingual LLMs under Semantics-Preserving Perturbations

arXiv cs.CL · 2026-09-04 Cached

This paper investigates the structural sensitivity of multilingual large language models to semantics-preserving perturbations in Hindi and Malayalam, showing significant degradation in mathematical reasoning performance and introducing the IndicReStruct benchmark for evaluation.

0 favorites 0 likes
#robustness

How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models

arXiv cs.CL · 2026-09-04 Cached

This paper conducts a multi-level analysis of how input perturbations propagate through decoder-only language models, assessing robustness via output behavior, hidden-state geometry, and attention-head function across models like GPT-2 and Qwen2.5.

0 favorites 0 likes
#robustness

Topological Steering

arXiv cs.LG · 2026-09-02 Cached

Topological Steering is a new framework for controlling large language model behavior using topological data analysis to capture global structures in activation spaces, enabling more robust behavioral control.

0 favorites 0 likes
#robustness

Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?

arXiv cs.CL · 2026-09-02 Cached

This research investigates whether near-tied rankings of large language models remain robust to changes in benchmark item composition using psychometric methods, finding that while overall rankings are stable, individual close orderings can reverse based on item selection.

0 favorites 0 likes
#robustness

RL-FAT: Reinforcement Learning for Fair Adversarial Training

arXiv cs.LG · 2026-09-01 Cached

RL-FAT is a reinforcement learning framework for fair adversarial training that improves robustness while reducing class-wise disparities. Experiments demonstrate competitive accuracy and better fairness compared to standard methods.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback