Tag
This paper introduces Semantic Overlays, a technique using learned adapters to annotate input spans for language models, effectively mitigating prompt injection attacks while maintaining utility. It demonstrates strong defense results on benchmarks like SEP and TensorTrust.