Tag
This paper proposes a diagnostic framework to assess whether an LLM-judge can effectively evaluate candidate skills in optimization tasks without a reference verifier, focusing on competence and discriminability metrics.
This paper introduces a fixed-contract diagnostic tool to analyze why KV cache compression methods succeed or fail in long-context LLM inference. It identifies three failure modes—missing evidence, scoring irrelevant tokens, and breaking related evidence—and evaluates them on LongBench and NeedleBench.