security-assessment

Tag

Cards List
#security-assessment

Validity-Aware Jailbreak Evaluation for Large Language Models

arXiv cs.AI · yesterday Cached

This paper proposes SEAV, a verification-centric framework for evaluating jailbreak robustness in large language models by assessing response validity and correctness, significantly reducing false-positive rates in safety assessments.

0 favorites 0 likes
#security-assessment

Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection

Hugging Face Daily Papers · 2026-08-18 Cached

This paper evaluates indirect prompt injection risks in DeepSeek Harness using AI-Infra-Guard for controlled testing, finding notable attack success rates and recommending security controls.

0 favorites 0 likes
← Back to home

Submit Feedback