confidence-manipulation

Tag

Cards List
#confidence-manipulation

Forced Deferral: Manipulating Routing Decisions in Multimodal LLM Cascades

arXiv cs.AI · 2026-06-16 Cached

This paper introduces the Forced Deferral Attack (FDA), an adversarial image attack that manipulates confidence scores in multimodal LLM cascades, causing queries to be unnecessarily routed to stronger (more expensive) models, thereby shifting compute costs to the provider without degrading answer correctness.

0 favorites 0 likes
← Back to home

Submit Feedback