Tag
Discussion about the significant gap between Llama model benchmark scores and actual real-world performance, with the author seeking assistance.
This paper proposes GESD, a procedural-oriented fairness metric that measures disparities in explanation stability across subgroups, and integrates it into a multi-objective optimization framework for jointly optimizing utility, outcome fairness, and explanation fairness.