Tag
This paper presents an empirical study of inference-time gating in a production-scale clinical NLP pipeline using Llama and MMed-Llama models, showing that learning filtering rules from verifier rejections fails at scale, while ontology-based and evidence-testing filters are effective.