Interpretable and Fair Generalized Additive Neural Networks via Multi-objective Learning
Summary
This paper proposes a multi-objective neural basis model (MONBM) framework to simultaneously optimize accuracy, interpretability, and fairness in generalized additive neural networks, revealing complex trade-offs between these trustworthiness dimensions.
View Cached Full Text
Cached at: 09/10/26, 08:28 AM
# Interpretable and Fair Generalized Additive Neural Networks via Multi-objective Learning
Source: [https://arxiv.org/html/2609.05946](https://arxiv.org/html/2609.05946)
Journal:Neural NetworksZiming WangAffiliation:Guangdong Provincial Key Laboratory of Brain\-inspired Intelligent Computation, Department of Computer Science and Engineering, Southern University of Science and Technology, Shenzhen, 518055, ChinaChangwu HuangCorresponding author:Corresponding authors\.Affiliation:School of AI and Liberal Arts, Beijing Normal\-Hong Kong Baptist University, Zhuhai, 519087, ChinaKe TangCorresponding author:Corresponding authors\.Affiliation:Guangdong Provincial Key Laboratory of Brain\-inspired Intelligent Computation, Department of Computer Science and Engineering, Southern University of Science and Technology, Shenzhen, 518055, China Yew\-Soon OngAffiliation:Agency for Science, Technology and Research \(A\*STAR\), 138632, SingaporeAffiliation:College of Computing and Data Science, Nanyang Technological University, 639798, Singapore
###### Abstract
Interpretability and fairness are two of the most emphasized dimensions in trustworthy artificial intelligence \(AI\)\. Various explainable AI methods have been introduced to improve interpretability\. This paper focuses on neural network \(NN\)\-based generalized additive models \(GAMs\), a class of self\-interpretable models\. While most existing research has prioritized improving the accuracy of NN\-based GAMs, their interpretability remains largely underexplored\. To address this gap, this paper introduces explicit quantitative metrics for evaluating the interpretability of NN\-based GAMs, empirically examines their effectiveness, and explores strategies for improving interpretability within these models\. In addition, the simultaneous and explicit optimization of both interpretability and fairness, along with their trade\-offs and the underlying reasons, remains underexplored\. To address this, we propose a multi\-objective neural basis model \(MONBM\) framework based on multi\-objective evolutionary learning to consider accuracy, interpretability, and fairness simultaneously\. A partial retraining strategy is further developed to facilitate the practical application of evolutionary multi\-objective optimization to deep model architectures\. Based onMONBM, this paper reveals the complex relationships between these dimensions and the reasons behind these intricate relationships\. This analysis demonstrates how multi\-objective optimization can be combined with self\-interpretable models to reveal relationships among trustworthiness objectives\. In addition,MONBMobtains a set of models with different trade\-offs between dimensions, and the competitiveness of the approach is validated by comparing it with state\-of\-the\-art methods\.
###### Keywords:
Explainable AI , Fairness in machine learning , Generalized additive model , Multi\-objective learning , Neural networks
## 1Introduction
With the widespread adoption of artificial intelligence \(AI\), there is growing concern and emphasis on developing trustworthy AI systems\. Trustworthy AI is a multi\-dimensional concept that includes interpretability, fairness, safety, and security, among other aspects\[[18](https://arxiv.org/html/2609.05946#bib.bib3),[28](https://arxiv.org/html/2609.05946#bib.bib2)\]\. Among them, interpretability and fairness are the most frequently emphasized and discussed\[[18](https://arxiv.org/html/2609.05946#bib.bib3)\]\.
To improve interpretability, various explainable AI \(XAI\) methods have been proposed\[[48](https://arxiv.org/html/2609.05946#bib.bib48)\], among which self\-interpretable models are widely favored due to their high fidelity\. In this paper, we focus on a class of self\-interpretable models called generalized additive models \(GAMs\), which are popular for their high accuracy and ability to provide both global and local explanations\. GAMs can be further classified based on the methods used to learn shape functions, including spline\-based\[[17](https://arxiv.org/html/2609.05946#bib.bib12)\], tree\-based\[[40](https://arxiv.org/html/2609.05946#bib.bib21)\], and neural network \(NN\)\-based methods\[[1](https://arxiv.org/html/2609.05946#bib.bib10),[42](https://arxiv.org/html/2609.05946#bib.bib11)\], among others\. This paper focuses on NN\-based GAMs due to their superior performance and scalability\[[1](https://arxiv.org/html/2609.05946#bib.bib10)\]\.
However, existing research on GAMs often assumes that they inherently provide satisfactory explanations, with much of the effort devoted to architectural design for improving predictive accuracy\[[1](https://arxiv.org/html/2609.05946#bib.bib10),[42](https://arxiv.org/html/2609.05946#bib.bib11)\]\. Although GAMs do provide global and local explanations, the quality of their explanations, i\.e\., whether they are easy to understand, is not actually guaranteed\. To the best of our knowledge, only limited research has been dedicated to improving the interpretability of GAMs, especially NN\-based GAMs\. Traditional regularization techniques, including weight decay and Bayesian approaches, can indirectly encourage smoother and simpler models by controlling model complexity\. However, these approaches usually influence interpretability implicitly and do not provide explicit quantitative criteria for evaluating or directly optimizing interpretability as an independent objective\. To address this limitation, we introduce explicitly defined quantitative metrics—smoothness and monotonicity—to evaluate and directly enhance the interpretability of NN\-based GAMs from a functional perspective\.
Concurrently, fairness in machine learning \(ML\) aims to ensure that model decisions are not biased towards any individual or group\[[37](https://arxiv.org/html/2609.05946#bib.bib6),[41](https://arxiv.org/html/2609.05946#bib.bib61)\]\. It is generally categorized into individual fairness and group fairness\. This paper focuses on group fairness, which requires that different groups defined by sensitive attributes should be treated equally\[[37](https://arxiv.org/html/2609.05946#bib.bib6)\]\. To evaluate and improve fairness in ML models, various metrics\[[2](https://arxiv.org/html/2609.05946#bib.bib13)\]and methods\[[58](https://arxiv.org/html/2609.05946#bib.bib7),[57](https://arxiv.org/html/2609.05946#bib.bib52)\]have been proposed\. However, existing methods primarily focus on black\-box models and rarely consider fairness in self\-interpretable models, especially GAMs\.
While interpretability and fairness have each received considerable attention in the development of trustworthy AI, they are often studied in isolation\. Prior work has highlighted the conflict between accuracy and fairness\[[58](https://arxiv.org/html/2609.05946#bib.bib7),[57](https://arxiv.org/html/2609.05946#bib.bib52)\], and interpretability has been framed as a multi\-objective optimization problem\[[47](https://arxiv.org/html/2609.05946#bib.bib1),[61](https://arxiv.org/html/2609.05946#bib.bib51),[39](https://arxiv.org/html/2609.05946#bib.bib50)\]\. However, research examining the relationships among accuracy, interpretability, and fairness remains relatively limited, particularly in terms of systematic empirical analysis\[[21](https://arxiv.org/html/2609.05946#bib.bib14)\]\. More importantly, the underlying mechanisms driving these dimensional relationships are still unclear\.
Additionally, different stakeholders often have divergent needs and expectations regarding fairness and interpretability, depending on their roles and contexts\. For example, critical domains like justice and healthcare both require explainability and fairness, but with differing emphases\. Justice typically emphasizes fairness, whereas healthcare places greater importance on explainability\.
These motivate us to formulate the problem as a multi\-objective learning problem while simultaneously considering the model’s accuracy, interpretability, and fairness—three often conflicting objectives—and propose an evolutionary multi\-objective learning\-based method to solve this problem\. Leveraging the ability of multi\-objective learning to obtain a set of diverse solutions, our method not only has the potential to provide tailored models for different stakeholders, but also enables us to analyze the complex relationship between optimization objectives and the underlying causes by combining with the self\-explanatory capabilities of GAMs\. Additionally, we validate the effectiveness of the proposed metrics and methods by comparing them with state\-of\-the\-art methods\.
Our research can be broken down into three more specific questions:
1. RQ1\.Are the global interpretability metrics we have devised reasonable, and can our method obtain models with the desired global interpretability?
2. RQ2\.What are the relationships among the accuracy, interpretability, and fairness of ML models, and why do these relationships exist?
3. RQ3\.How well can our proposed method address the identified problem, and what are its advantages compared to existing self\-interpretable and black\-box models?
To answer these research questions, this paper proposes a multi\-objective neural basis model \(MONBM\) framework111The code used in this paper is available at https://github\.com/oddwang/MONBM/\.that simultaneously considers accuracy, interpretability, and fairness\. The main contributions of this paper are as follows:
1. \(1\)We introduce explicit quantitative metrics for smoothness and monotonicity that are designed to directly measure and improve the global interpretability of NN\-based GAMs\. Unlike traditional implicit regularization methods, these metrics allow interpretability properties to be explicitly quantified, monitored, and incorporated as independent optimization objectives within a multi\-objective framework\. We experimentally verify both the effectiveness of the metrics and the ability of our method to obtain NBMs with desirable global interpretability\.
2. \(2\)We formulate the problem of constructing trustworthy models as a multi\-objective learning problem and propose theMONBMframework to address it\. As the first to apply multi\-objective learning to GAMs, the framework considers accuracy, interpretability, and fairness simultaneously, and we implement multiple instantiations of it by optimizing different objectives\. Notably, our instantiations adopt an efficient strategy by retraining only the “feature specific” output layer of the pre\-trained NBM, enabling evolutionary multi\-objective optimization \(EMO\) to be efficiently applied to large\-scale networks\.
3. \(3\)We useMONBMto obtain a set of NBMs with different trade\-offs among accuracy, interpretability, and fairness, and leverage the explanations provided by NBM to reveal the relationships between multiple dimensions and the reasons behind them\. This analysis combines EMO with self\-interpretable models, providing a new perspective for revealing the interrelationships among various trustworthiness objectives by leveraging the diversity of the non\-dominated solution set to delve into the “why” of objective interplay\.
4. \(4\)We verify thatMONBMcan solve the proposed problem well, obtaining a set of NBMs with desirable performance and trade\-offs in multiple dimensions, and are competitive with state\-of\-the\-art self\-interpretable and black\-box models\.
The rest of this paper is organized as follows\. Section[2](https://arxiv.org/html/2609.05946#S2)introduces GAM and reviews related work\. Section[3](https://arxiv.org/html/2609.05946#S3)proposes global interpretability metrics and ourMONBMframework and its instantiations\. Section[4](https://arxiv.org/html/2609.05946#S4)answers the three research questions posed through experimental studies\. Sections[5](https://arxiv.org/html/2609.05946#S5)and[6](https://arxiv.org/html/2609.05946#S6)provide further discussion and summarize this paper, respectively\.
## 2Background
In this section, we introduce GAM, focusing on NN\-based GAM, and discuss its interpretability\. Then, we discuss metrics and methods for evaluating and improving group fairness in ML\. Finally, we summarize representative studies on GAM to clarify the positioning of this research\.
### 2\.1Generalized Additive Model \(GAM\)
GAM is an extension of the linear model that establishes dependencies between input features and the target variable by summing multiple independent nonlinear mappings, known as shape functions\[[17](https://arxiv.org/html/2609.05946#bib.bib12)\]\. Formally, GAM can be expressed as:
g\(𝔼\(y\)\)=β0\+f1\(x1\)\+f2\(x2\)\+…\+fd\(xd\),g\(\\mathbb\{E\}\(y\)\)=\\beta\_\{0\}\+f\_\{1\}\(x\_\{1\}\)\+f\_\{2\}\(x\_\{2\}\)\+\.\.\.\+f\_\{d\}\(x\_\{d\}\),\(1\)wherex=\{x1,x2,…,xd\}∈ℝd\\textnormal\{\{x\}\}=\\\{x\_\{1\},x\_\{2\},\.\.\.,x\_\{d\}\\\}\\in\\mathbb\{R\}^\{d\}is the input withddfeatures,yyis the target variable,g\(⋅\)g\(\\cdot\)is the link function \(e\.g\., the logistic function\), and eachfif\_\{i\}is a univariate shape function\.
Since the effect of each feature in a GAM is modeled independently without considering feature interactions, it clearly decomposes the influence of each feature on the ML model’s decisions\. This decomposition allows GAMs to provide two main types of explanations\[[23](https://arxiv.org/html/2609.05946#bib.bib20)\]:
1. \(1\)Global Explanation\[[23](https://arxiv.org/html/2609.05946#bib.bib20)\]: According to Eq\. \([1](https://arxiv.org/html/2609.05946#S2.E1)\), each shape functionfif\_\{i\}depends only on theii\-th featurexix\_\{i\}of the input vectorx\. A global explanation is generated by visualizing each shape functionfif\_\{i\}, as illustrated in Fig\.[1\(a\)](https://arxiv.org/html/2609.05946#S2.F1.sf1)\. The figure illustrates how changes in the input feature affect the predicted output\. By visualizing the shape functions of all features, we can capture the overall decision logic of the ML model\.
2. \(2\)Local Explanation\[[23](https://arxiv.org/html/2609.05946#bib.bib20)\]: As shown in Eq\. \([1](https://arxiv.org/html/2609.05946#S2.E1)\), the final prediction of a GAM is the sum of all shape functions and the intercept term\. For a given data pointx\(j\)\\textnormal\{\{x\}\}^\{\(j\)\}, the shape function valuefi\(xi\(j\)\)f\_\{i\}\(x\_\{i\}^\{\(j\)\}\)for each featurexi\(j\)x\_\{i\}^\{\(j\)\}can be visualized, as illustrated in Fig\.[1\(b\)](https://arxiv.org/html/2609.05946#S2.F1.sf2)\. This visualization provides insight into the contribution of each feature to the ML model’s decision for that particular data point, offering a local explanation\.
\(a\)Global Explanation
\(b\)Local Explanation
Figure 1:Examples of explanations provided by GAM\. For the global explanation, as shown in Fig\.[1\(a\)](https://arxiv.org/html/2609.05946#S2.F1.sf1), the x\-axis represents the input feature valuesxix\_\{i\}, while the y\-axis corresponds to the values of the shape functionfi\(xi\)f\_\{i\}\(x\_\{i\}\)\. For the local explanation, as shown in Fig\.[1\(b\)](https://arxiv.org/html/2609.05946#S2.F1.sf2), the y\-axis represents the featuresxix\_\{i\}, and the x\-axis represents the shape function valuefi\(xi\(j\)\)f\_\{i\}\(x\_\{i\}^\{\(j\)\}\)for the corresponding featurexi\(j\)x\_\{i\}^\{\(j\)\}of data pointx\(j\)\\textnormal\{\{x\}\}^\{\(j\)\}\.Although the explanation provided by GAM appears to be heuristic, it actually provides a precise description of how GAM makes decisions\. The key to GAM is how to learn its shape functions\[[42](https://arxiv.org/html/2609.05946#bib.bib11)\]; its methods can be classified into spline\-based\[[17](https://arxiv.org/html/2609.05946#bib.bib12)\], tree\-based\[[40](https://arxiv.org/html/2609.05946#bib.bib21)\], and NN\-based\[[1](https://arxiv.org/html/2609.05946#bib.bib10),[42](https://arxiv.org/html/2609.05946#bib.bib11)\], among others\. This paper focuses on NN\-based GAMs because of their superior performance and scalability\[[1](https://arxiv.org/html/2609.05946#bib.bib10)\]\.
NN\-based GAMs:Agarwalet al\.\[[1](https://arxiv.org/html/2609.05946#bib.bib10)\]proposed the first NN\-based GAM, known as the neural additive model \(NAM\), in which each shape function is represented by a deep neural network \(DNN\)\. However, NAM’s design, which requires a separate DNN for each feature, limits its scalability when applied to datasets with many features\. To address this, Radenovicet al\.\[[42](https://arxiv.org/html/2609.05946#bib.bib11)\]proposed the neural basis model \(NBM\), as illustrated in Fig\.[2](https://arxiv.org/html/2609.05946#S2.F2)\. NBM employs a single DNN, specifically a multi\-layer perceptron \(MLP\) withone\-input andB\-outputs, as a shared basis function, replacing the need for a separate DNN for each feature\. The shape function of each feature is then obtained through a linear combination of the “feature specific” layer, which has parameters of sizeD×\(B\+1\)D\\times\(B\+1\), significantly reducing the ML model’s parameters, particularly when handling datasets with high dimensions\.
Figure 2:NBM architecture for the binary classification task, the figure taken from\[[42](https://arxiv.org/html/2609.05946#bib.bib11)\]with slight modifications\.It is worth mentioning that, to improve predictive performance, standard GAMs have been extended to model pairwise feature interactions, resulting in the GAM with pairwise interactions \(GA2M\\text\{GA\}^\{2\}\\text\{M\}\)\[[33](https://arxiv.org/html/2609.05946#bib.bib49)\]\. However, this extension often reduces interpretability and introduces substantial computational complexity\. In particular, modeling bivariate interactions leads to a combinatorial increase in candidate interactions \(O\(d2\)O\(d^\{2\}\)\), making effective structure discovery and feature selection essential\. To address this issue, recent studies have introduced dedicated regularization and structure\-learning mechanisms\. For example, Changet al\.\[[11](https://arxiv.org/html/2609.05946#bib.bib63)\]employ differentiable trees to improve the scalability of interaction discovery, while Yanget al\.\[[53](https://arxiv.org/html/2609.05946#bib.bib15)\]introduce structural constraints to identify important main effects and interactions\. Similarly, Xuet al\.\[[52](https://arxiv.org/html/2609.05946#bib.bib16)\]incorporate group sparsity to enable automatic feature selection, and Zhuet al\.\[[62](https://arxiv.org/html/2609.05946#bib.bib60)\]propose a learnable gating mechanism to distinguish different feature types\. Since the primary focus of this paper is not to pursue extreme predictive performance at the expense of model transparency, both our method and the baseline methods center on standard GAM formulations\. Nevertheless, our proposed method can be naturally extended toGA2M\\text\{GA\}^\{2\}\\text\{M\}and integrated with existing structure discovery techniques\. Specifically, each individual in the population can be configured as aGA2M\\text\{GA\}^\{2\}\\text\{M\}\(e\.g\.,NB2M\\text\{NB\}^\{2\}\\text\{M\}\) to natively accommodate pairwise interactions\. Concurrently, existing sparsity constraints or gating techniques can be applied directly to the pre\-trained NBM to effectively filter out redundant interactions, thereby reducing the number of candidate features while retaining key structural components\.
Parallel to these structural advancements within the tabular domain, recent research has also expanded the scope of GAMs to specialized structured modalities, such as time series and images\. For example, GATSM\[[22](https://arxiv.org/html/2609.05946#bib.bib64)\]adapts additive modeling to multivariate time\-series forecasting by combining an NBM\-based feature representation with an interpretable temporal modeling module\. In computer vision settings, the original NBM framework\[[42](https://arxiv.org/html/2609.05946#bib.bib11)\]further demonstrates that additive reasoning can operate on human\-interpretable concepts extracted from images rather than directly on raw pixels\. SinceMONBMis designed as a model\-agnostic optimization framework that separates optimization objectives from the underlying self\-interpretable architecture, and notably because these specialized frameworks\[[22](https://arxiv.org/html/2609.05946#bib.bib64),[42](https://arxiv.org/html/2609.05946#bib.bib11)\]share the exact same NBM backbone architecture as our instantiation, these developments suggest a feasible extension path\. Specifically, temporal or concept\-level additive models could serve as individuals withinMONBM, while existing optimization objectives could be adapted to evaluate task\-specific notions of accuracy and interpretability\. Such extensions may enable systematic analysis of trade\-offs among multiple trustworthiness dimensions in broader application scenarios, which we leave for future work\.
### 2\.2Interpretability of NN\-based GAM
The interpretability of NN\-based GAMs can be further categorized into local interpretability and global interpretability based on the type of explanation they provide, as described below\.
#### 2\.2\.1Local Interpretability \(LI\)
LIemphasizes the ability to interpret a given data point at a local level\. Bhattet al\.\[[7](https://arxiv.org/html/2609.05946#bib.bib22)\]and Chalasaniet al\.\[[9](https://arxiv.org/html/2609.05946#bib.bib23)\]proposed theμC\\mu\_\{C\}andGGmetrics, respectively, to assess the complexity of local explanations, which refers to the ability to explain model decisions with a small number of features\[[47](https://arxiv.org/html/2609.05946#bib.bib1)\]\. TheμC\\mu\_\{C\}andGGmetrics are defined as follows:
μC\(g;x\(j\)\)=−∑i=1dℙ\(i\)ln\(ℙ\(i\)\),\\mu\_\{C\}\(g;\\textnormal\{\{x\}\}^\{\(j\)\}\)=\-\\sum\_\{i=1\}^\{d\}\\mathbb\{P\}\(i\)\\textnormal\{ln\}\(\\mathbb\{P\}\(i\)\),\(2\)ℙ\(i\)=\|fi\(xi\(j\)\)\|∑k=1d\|fk\(xk\(j\)\)\|\.\\mathbb\{P\}\(i\)=\\frac\{\|f\_\{i\}\(x\_\{i\}^\{\(j\)\}\)\|\}\{\\sum\_\{k=1\}^\{d\}\|f\_\{k\}\(x\_\{k\}^\{\(j\)\}\)\|\}\.\(3\)G\(g,x\(j\)\)=1−2∑i=1dvi\(j\)‖𝒗\(j\)‖1\(d−i\+0\.5d\),G\(g;\\textnormal\{\{x\}\}^\{\(j\)\}\)=1\-2\\sum\_\{i=1\}^\{d\}\\frac\{v\_\{i\}^\{\(j\)\}\}\{\|\|\\bm\{v\}^\{\(j\)\}\|\|\_\{1\}\}\\left\(\\frac\{d\-i\+0\.5\}\{d\}\\right\),\(4\)𝒗\(j\)=sortnon\-decreasing\(\[\|f1\(x1\(j\)\)\|,\|f2\(x2\(j\)\)\|,…,\|fd\(xd\(j\)\)\|\]\)\.\\bm\{v\}^\{\(j\)\}=\\textnormal\{sort\}\_\{\\textnormal\{non\-decreasing\}\}\(\[\|f\_\{1\}\(x\_\{1\}^\{\(j\)\}\)\|,\|f\_\{2\}\(x\_\{2\}^\{\(j\)\}\)\|,\.\.\.,\|f\_\{d\}\(x\_\{d\}^\{\(j\)\}\)\|\]\)\.\(5\)
The complexity metric is based on the idea that the local explanations are more complex when feature importance scores \(derived from the shape function\) are evenly distributed across features, and simpler when concentrated on a few features\[[47](https://arxiv.org/html/2609.05946#bib.bib1)\]\. In other words, explanations involving fewer dominant features are generally easier to understand\. Although prior work has assessed\[[7](https://arxiv.org/html/2609.05946#bib.bib22),[9](https://arxiv.org/html/2609.05946#bib.bib23)\]or reduced\[[47](https://arxiv.org/html/2609.05946#bib.bib1)\]the complexity of local explanations from post\-hoc methods, no studies have applied these approaches to self\-interpretable models, particularly GAMs\. Some NN\-based GAMs, such as those incorporating feature selection\[[52](https://arxiv.org/html/2609.05946#bib.bib16),[53](https://arxiv.org/html/2609.05946#bib.bib15),[62](https://arxiv.org/html/2609.05946#bib.bib60)\], indirectly simplify local explanations by retaining only the most important features\. However, this approach is suboptimal, and methods specifically aimed at reducing the complexity of local explanations in GAMs remain lacking\.
#### 2\.2\.2Global Interpretability \(GI\)
GIconcerns the ML model’s ability to explain the relationships between input features and the target variable at a global level\. Existing work indicates that the global interpretability provided by GAM should exhibit monotonicity\[[60](https://arxiv.org/html/2609.05946#bib.bib24)\]and smoothness\[[36](https://arxiv.org/html/2609.05946#bib.bib25)\]\. These properties help ensure that the learned feature–target relationships are stable, intuitive, and consistent with domain knowledge\.
For smoothness, it is worth noting that most existing techniques are not designed to explicitly enhance the smoothness of global explanations\. Some approaches rely on implicit regularization to control overall model complexity\. Classical techniques such as weight decay\[[24](https://arxiv.org/html/2609.05946#bib.bib55)\]and Bayesian priors\[[35](https://arxiv.org/html/2609.05946#bib.bib56)\]constrain parameter magnitudes and reduce overfitting, which may incidentally lead to smoother shape functions\. However, their primary objective is to improve predictive generalization rather than to enforce interpretable smoothness, and they do not provide explicit quantitative measures of the smoothness of the learned functions\. Similarly, while early spline\-based GAMs allowed smoothness adjustment through dedicated smoothing parameters\[[36](https://arxiv.org/html/2609.05946#bib.bib25)\], such mechanisms are intrinsic to spline formulations and are not applicable to modern tree\-based\[[40](https://arxiv.org/html/2609.05946#bib.bib21)\]and NN\-based GAMs\[[1](https://arxiv.org/html/2609.05946#bib.bib10),[42](https://arxiv.org/html/2609.05946#bib.bib11)\]\. Recent tree\-based and NN\-based GAM variants have largely focused on predictive performance, typically neglecting the smoothness of global explanations and lacking explicit metrics for their evaluation and improvement\.
For monotonicity, Zhonget al\.\[[60](https://arxiv.org/html/2609.05946#bib.bib24)\]emphasized that certain features should satisfy monotonicity to align with user intuition\. Recent research\[[12](https://arxiv.org/html/2609.05946#bib.bib57)\]encourages monotonic behaviour through penalty terms, particularly in regulated domains like credit scoring\. However, existing NN\-based GAM approaches still lack explicit quantitative metrics that both measure monotonicity and allow it to be optimized directly as an independent objective within multi\-objective learning frameworks\.
Overall, there remains a significant gap in metrics and methods to explicitly evaluate and improve the interpretability of NN\-based GAMs\. This gap motivates this paper to propose explicit smoothness and monotonicity metrics and to obtain NN\-based GAMs with high interpretability by applying these metrics\.
### 2\.3Group Fairness in Machine Learning
Group fairness refers to ML models that treat different groups equally\[[37](https://arxiv.org/html/2609.05946#bib.bib6)\]\. To assess group fairness, various metrics have been proposed\[[2](https://arxiv.org/html/2609.05946#bib.bib13)\], with demographic parity \(DP\) being one of the most widely used\[[15](https://arxiv.org/html/2609.05946#bib.bib17),[8](https://arxiv.org/html/2609.05946#bib.bib18)\], which is defined as:
DP=\|P\(y^=1\|s=s1\)−P\(y^=1\|s=s2\)\|,DP=\\left\|P\\left\(\\hat\{y\}=1\|s=s\_\{1\}\\right\)\-P\\left\(\\hat\{y\}=1\|s=s\_\{2\}\\right\)\\right\|,\(6\)wherey^\\hat\{y\}is the predicted label,ssrepresents the sensitive attribute \(e\.g\., sex\), ands1s\_\{1\}ands2s\_\{2\}denote two different groups associated with the sensitive attribute\. In this paper, we use theDPmetric to assess the fairness of GAMs\.
Based on the defined fairness metrics, researchers use three types of methods to improve fairness: \(1\) pre\-processing methods, which adjust data before training to reduce bias\[[56](https://arxiv.org/html/2609.05946#bib.bib8),[19](https://arxiv.org/html/2609.05946#bib.bib9)\]; \(2\) in\-processing methods, which incorporate fairness considerations during training\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]; and \(3\) post\-processing methods, which adjust the model’s output after training to improve fairness\[[20](https://arxiv.org/html/2609.05946#bib.bib19),[50](https://arxiv.org/html/2609.05946#bib.bib53),[51](https://arxiv.org/html/2609.05946#bib.bib54)\]\.
However, the majority of existing in\-processing research is tailored specifically for standard black\-box architectures, such as deep neural networks\. These techniques are often closely tied to specific model structures \(e\.g\., customized loss regularization or adversarial training strategies for unconstrained multilayer perceptrons\), making them difficult to directly adapt to the additive structure of inherently interpretable models such as GAMs\. To the best of our knowledge, Krauset al\.\[[23](https://arxiv.org/html/2609.05946#bib.bib20)\]are among the few who have identified fairness issues in NN\-based GAMs\. They proposed a strategy to mitigate unfairness by minimizing the influence of the sensitive attribute on decision\-making\. However, our experiments find limited improvement in fairness with their method\. Therefore, this paper addresses fairness in GAMs by incorporating fairness metrics explicitly as one of the optimization objectives within theMONBMframework\.
### 2\.4Summary and Research Positioning
To clearly present the current state of research on GAMs and clarify the positioning of this study, we summarize representative GAM approaches in Table[1](https://arxiv.org/html/2609.05946#S2.T1)to provide a structured overview of the literature\. The table highlights their model categories \(e\.g\., NN\-based or traditional GAMs\) and indicates which key trustworthiness dimensions they explicitly consider, including fairness, local interpretability \(model complexity\), and global interpretability \(smoothness and monotonicity\)\.
Table 1:Comparison of the proposedMONBMframework with existing GAM\-based approaches\.As shown in Table[1](https://arxiv.org/html/2609.05946#S2.T1), although substantial progress has been made across individual trustworthiness dimensions, existing GAM studies typically address these aspects in isolation rather than in an integrated manner\. Most prior work focuses on at most one trustworthiness dimension at a time—for example, improving model complexity for local interpretability, enforcing monotonicity for global interpretability, or mitigating bias for fairness—without jointly considering their interactions\. First, many NN\-based GAMs \(such as NAM\[[1](https://arxiv.org/html/2609.05946#bib.bib10)\]and NBM\[[42](https://arxiv.org/html/2609.05946#bib.bib11)\]\) primarily emphasize predictive performance while neglecting interpretability and fairness\. Second, although some recent studies have incorporated specific requirements such as complexity\[[52](https://arxiv.org/html/2609.05946#bib.bib16),[53](https://arxiv.org/html/2609.05946#bib.bib15),[62](https://arxiv.org/html/2609.05946#bib.bib60)\], monotonicity\[[12](https://arxiv.org/html/2609.05946#bib.bib57)\], or fairness\[[23](https://arxiv.org/html/2609.05946#bib.bib20)\], these efforts generally introduce a single additional constraint through regularization or penalty terms within a single\-objective optimization framework\. As a result, they neither explicitly model multiple trustworthiness objectives simultaneously nor systematically explore the trade\-offs among them\.
Within this context,MONBMformulates accuracy, fairness, and both local and global interpretability as explicit objectives in a unified multi\-objective evolutionary learning framework\. By jointly optimizing these four dimensions rather than introducing them as auxiliary regularizers, the proposed approach enables systematic analysis of their interactions\. Instead of producing a single optimized solution,MONBMgenerates a set of Pareto\-optimal models, allowing practitioners to directly examine and select among different trade\-offs\. This perspective shifts GAM research from implicit or single\-objective constraints toward an explicit multi\-objective framework for trustworthy model design\.
## 3Multi\-objective Evolutionary Learning to Construct Accurate, Interpretable, and Fair Models
In this section, we introduce explicit quantitative metrics designed to measure and optimize the global interpretability of NN\-based GAMs\. Then, theMONBMframework is proposed to solve the considered multi\-objective learning problem\. The framework uses multi\-objective learning methods\[[47](https://arxiv.org/html/2609.05946#bib.bib1),[58](https://arxiv.org/html/2609.05946#bib.bib7),[10](https://arxiv.org/html/2609.05946#bib.bib26)\]to obtain a set of accurate, interpretable, and fair models with different trade\-offs\. Finally, multiple instantiations of our framework are provided by specifying the self\-interpretable models used, the optimizer, and considering different optimization objectives\.
### 3\.1Smoothness and Monotonicity Metrics
In global explanation, smoothness means that shape functions should change gradually over the range of feature values without abrupt fluctuations, while monotonicity requires that a feature’s shape function consistently increases or decreases, aligning with user intuition or domain knowledge\. However, monotonicity should not be considered a general requirement for interpretability\. In many practical scenarios, non\-monotonic relationships \(e\.g\., U\-shaped patterns or periodic variations\) are natural and necessary for accurately reflecting real data patterns\. Therefore, monotonicity is treated in this work as an optional constraint that can be applied when supported by domain knowledge or stakeholder expectations, rather than as a default assumption\. Furthermore, smoothness and monotonicity are generally only required for continuous features, not discrete ones\.
Traditional regularization approaches, including weight decay and Bayesian methods, can indirectly promote smoother models by penalizing model complexity\. However, they generally do not provide explicit quantitative measures of interpretability properties from the perspective of shape functions\. To complement these approaches, we introduce two explicit quantitative metrics designed to measure and optimize global interpretability properties:ℳs\\mathcal\{M\}\_\{s\}for smoothness andℳm\\mathcal\{M\}\_\{m\}for monotonicity, with the latter divided intoℳm↑\\mathcal\{M\}\_\{m\\uparrow\}\(monotonic increasing\) andℳm↓\\mathcal\{M\}\_\{m\\downarrow\}\(monotonic decreasing\)\. Given a featurexix\_\{i\}, its shape functionfif\_\{i\}, andnnpoints taken uniformly between the minimum and maximum values ofxix\_\{i\}\(denoted asxi,minx\_\{i,\\textnormal\{min\}\}andxi,maxx\_\{i,\\textnormal\{max\}\}, respectively\), represented as\{xi,1,xi,2,…,xi,n\}\\\{x\_\{i,1\},x\_\{i,2\},\.\.\.,x\_\{i,n\}\\\}, the metricsℳs\\mathcal\{M\}\_\{s\},ℳm↑\\mathcal\{M\}\_\{m\\uparrow\}, andℳm↓\\mathcal\{M\}\_\{m\\downarrow\}are defined as follows:
ℳs\(fi\)=∑j=1n−1wi,jmax\{\|fi\(xi,j\+1\)−fi\(xi,j\)\|fi,max−fi,min−k,0\},\\mathcal\{M\}\_\{s\}\(f\_\{i\}\)=\\sum\_\{j=1\}^\{n\-1\}w\_\{i,j\}\\;\\textnormal\{max\}\\left\\\{\\frac\{\|f\_\{i\}\(x\_\{i,j\+1\}\)\-f\_\{i\}\(x\_\{i,j\}\)\|\}\{f\_\{i,\\textnormal\{max\}\}\-f\_\{i,\\textnormal\{min\}\}\}\-k,0\\right\\\},\(7\)
ℳm↑\(fi\)=∑j=1n−1max\{fi\(xi,j\)−fi\(xi,j\+1\)fi,max−fi,min,0\},\\mathcal\{M\}\_\{m\\uparrow\}\(f\_\{i\}\)=\\sum\_\{j=1\}^\{n\-1\}\\textnormal\{max\}\\left\\\{\\frac\{f\_\{i\}\(x\_\{i,j\}\)\-f\_\{i\}\(x\_\{i,j\+1\}\)\}\{f\_\{i,\\textnormal\{max\}\}\-f\_\{i,\\textnormal\{min\}\}\},0\\right\\\},\(8\)
ℳm↓\(fi\)=∑j=1n−1max\{fi\(xi,j\+1\)−fi\(xi,j\)fi,max−fi,min,0\},\\mathcal\{M\}\_\{m\\downarrow\}\(f\_\{i\}\)=\\sum\_\{j=1\}^\{n\-1\}\\textnormal\{max\}\\left\\\{\\frac\{f\_\{i\}\(x\_\{i,j\+1\}\)\-f\_\{i\}\(x\_\{i,j\}\)\}\{f\_\{i,\\textnormal\{max\}\}\-f\_\{i,\\textnormal\{min\}\}\},0\\right\\\},\(9\)wherefi,maxf\_\{i,\\textnormal\{max\}\}andfi,minf\_\{i,\\textnormal\{min\}\}denote the maximum and minimum values of the shape functionfif\_\{i\}over the feature’s range, and the hyper\-parameterkkcontrols the smoothing degree\. The termwi,jw\_\{i,j\}is an optional interval\-specific weight that allows different levels of smoothness to be imposed across different regions of the input space\. This design enables users to incorporate domain knowledge when certain regions require stronger or weaker smoothness constraints\. In this work, we use the default settingwi,j=1w\_\{i,j\}=1for all intervals, which provides a uniform penalty when no prior knowledge about local smoothness is available\. Bothℳs\\mathcal\{M\}\_\{s\}andℳm\\mathcal\{M\}\_\{m\}metrics are the smaller, the better\.
The proposed metrics have several desirable properties\. First, they are intuitive and easy to understand: smoothness is defined by measuring the degree of abrupt change in function values, while monotonicity directly captures the continuous upward or downward trend of a function\. These definitions are highly compatible with human intuition and domain knowledge\. Second, the smoothness metric balances robustness and sensitivity by using a threshold\-based approach that allows users to ignore small changes on demand while clearly identifying significant fluctuations\. In addition, both metrics can be easily converted into differentiable forms for integration into existing ML or multi\-objective optimization frameworks to guide the search for generating GAMs with higher global interpretability\.
### 3\.2MONBM Framework
We first formally define the considered multi\-objective learning problem to obtain a set of accurate, interpretable, and fair models\. Given a dataset𝓓\\bm\{\\mathcal\{D\}\}, a self\-interpretable modelgg, and its parameter matrix to be optimizedw, the multi\-objective learning problem can be defined as:
Minimizew𝓜\(g\(w\),𝓓\)=\[ℳ1\(g\(w\),𝓓\),…,ℳk\(g\(w\),𝓓\)\],\\mathop\{\\rm Minimize\}\_\{\\textnormal\{\{w\}\}\}\\\>\\bm\{\\mathcal\{M\}\}\(g\(\\textnormal\{\{w\}\}\);\\bm\{\\mathcal\{D\}\}\)=\[\\mathcal\{M\}\_\{1\}\(g\(\\textnormal\{\{w\}\}\);\\bm\{\\mathcal\{D\}\}\),\.\.\.,\\mathcal\{M\}\_\{k\}\(g\(\\textnormal\{\{w\}\}\);\\bm\{\\mathcal\{D\}\}\)\],\(10\)wherewis optimized to minimize a selected set of metrics𝓜\(⋅\)\\bm\{\\mathcal\{M\}\}\(\\cdot\)\. Eachℳi\\mathcal\{M\}\_\{i\}fori=1,…,ki=1,\.\.\.,krepresents a distinct objective, such as accuracy, interpretability, or fairness\.
To solve this problem, in theMONBMframework, we use EMO\[[10](https://arxiv.org/html/2609.05946#bib.bib26)\]to train self\-interpretable models, aiming at obtaining a set of models with different trade\-offs in accuracy, interpretability, and fairness\. The overall process of theMONBMframework is shown in Algorithm[1](https://arxiv.org/html/2609.05946#alg1)\.
Algorithm 1Multi\-objective neural basis model \(MONBM\) framework1:A training dataset
𝓓train\\bm\{\\mathcal\{D\}\}\_\{\\textnormal\{\{train\}\}\}, an initialized model set
G=\{g1,…,gτ\}G=\\\{g\_\{1\},\.\.\.,g\_\{\\tau\}\\\}, a set of evaluation metrics
𝓜=\{ℳ1,⋯,ℳk\}\\bm\{\\mathcal\{M\}\}=\\\{\\mathcal\{M\}\_\{1\},\\cdots,\\mathcal\{M\}\_\{k\}\\\}, an environment selection strategy
π\\pi, and a reproduction strategy
ξ\{\\xi\}\.
2:A set of optimized models
G∗G^\{\*\}\.
3:
𝒎\\bm\{m\}←\\leftarrowEvaluate the value of metrics
𝓜\\bm\{\\mathcal\{M\}\}on
GGusing
𝓓train\\bm\{\\mathcal\{D\}\}\_\{\\textnormal\{\{train\}\}\}\.
4:whileiteration termination conditions not metdo
5:G′G^\{\\prime\}←\{\\leftarrow\}Selectnnmodels fromGGaccording to𝒎\\bm\{m\}and the environment selection strategyπ\\pi\.
6:G′′G^\{\\prime\\prime\}←\{\\leftarrow\}Generateμ\\mumodels fromG′G^\{\\prime\}according to the reproduction strategyξ\\xi\.
7:G′′G^\{\\prime\\prime\}←\{\\leftarrow\}Partially train\[[54](https://arxiv.org/html/2609.05946#bib.bib27),[55](https://arxiv.org/html/2609.05946#bib.bib28)\]G′′G^\{\\prime\\prime\}using𝓓train\\bm\{\\mathcal\{D\}\}\_\{\\textnormal\{\{train\}\}\}\.
8:𝒎′′\\bm\{m\}^\{\\prime\\prime\}←\\leftarrowEvaluate the value of the metrics𝓜\\bm\{\\mathcal\{M\}\}onG′′G^\{\\prime\\prime\}using𝓓train\\bm\{\\mathcal\{D\}\}\_\{\\textnormal\{\{train\}\}\}\.
9:<g1,𝒎1\>,…,<gτ,𝒎τ\><g\_\{1\},\\bm\{m\}\_\{1\}\>,\\dots,<g\_\{\\tau\},\\bm\{m\}\_\{\\tau\}\>←\\leftarrowSelectτ\\taumodels fromG∪G′′G\\cup G^\{\\prime\\prime\}based on𝒎\\bm\{m\}and𝒎′′\\bm\{m\}^\{\\prime\\prime\}through the environment selection strategyπ\\piand then updateG=\{g1,…,gτ\}G=\\\{g\_\{1\},\\dots,g\_\{\\tau\}\\\}and𝒎\\bm\{m\}, accordingly\.
10:endwhile
11:returnthe non\-dominated solutions in the set of optimized models
GGas
G∗G^\{\*\}\.
Given a training dataset𝓓train\\bm\{\\mathcal\{D\}\}\_\{\\textnormal\{\{train\}\}\}, an initialized model setGG, a set of evaluation metrics𝓜\\bm\{\\mathcal\{M\}\}, and the corresponding environment selection and reproduction strategiesπ\\piandξ\\xi, the general process of theMONBMframework is as follows: First, the initialized model set is evaluated \(line 1\)\. Then, the evolutionary loop begins \(lines 2\-8\)\. In each generation,nnmodels are first selected asG′G^\{\\prime\}according to the environment selection strategyπ\\pi\(line 3\)\. An offspringG′′G^\{\\prime\\prime\}is then generated according to the reproduction strategyξ\\xiand partial training\[[54](https://arxiv.org/html/2609.05946#bib.bib27),[55](https://arxiv.org/html/2609.05946#bib.bib28)\]\(lines 4\-5\)\. Next, the offspringG′′G^\{\\prime\\prime\}is evaluated \(line 6\), and a new populationGGis selected according to the environment selection strategyπ\\pi\(line 7\)\. The above steps are repeated until the iteration termination condition is reached\. Finally, an optimized model setG∗G^\{\*\}is returned \(line 9\)\.
The key ofMONBMis to generate ideal offspring self\-interpretable models and environment selection based on the given metrics\. Multi\-objective evolutionary algorithms \(MOEAs\)\[[26](https://arxiv.org/html/2609.05946#bib.bib29)\]offer ideal reproduction and environment selection strategies that help our framework generate a set of models with good convergence and diversity\.
### 3\.3Instantiation of the MONBM Framework
In the proposedMONBMframework, the choice of self\-interpretable model, initialization and optimization schemes, evaluation metrics, and optimizer can be adapted based on the specific tasks and requirements\. Fig\.[3](https://arxiv.org/html/2609.05946#S3.F3)illustrates the overall computational pipeline of our instantiation\.
Figure 3:The computational pipeline of theMONBMinstantiation\.The workflow is organized into three primary stages:
1. \(1\)Stage 1 performs data input and pre\-training of an NBM to obtain a high\-accuracy base model\. The shared basis functions are then frozen, which stabilizes the learned representation and reduces the search space for subsequent optimization\.
2. \(2\)Stage 2 executes the evolutionary optimization loop described in Algorithm[2](https://arxiv.org/html/2609.05946#alg2)\. During this stage, only the unfrozen feature\-specific layer is evolved and partially retrained, while the shared basis functions remain fixed\. This stage combines evolutionary search with partial gradient\-based training and SRA\-based environmental selection to jointly optimize multiple objectives\.
3. \(3\)Stage 3 outputs the final set of non\-dominated models, which approximate the Pareto front and represent different trade\-offs among the considered objectives\.
Below, we provide several instantiations of the framework, which will be experimentally evaluated in Section[4](https://arxiv.org/html/2609.05946#S4)\.
#### 3\.3\.1Model
The set of models in theMONBMframework can be various self\-interpretable models\. In our instantiations, a set of NBMs\[[42](https://arxiv.org/html/2609.05946#bib.bib11)\]is used as individuals\. Each NBM’s weights and biases are encoded as a real\-valued vector, representing an individual\. Due to the large architecture of NBMs, training a set of NBMs via multi\-objective learning is complex and time\-consuming\. To address this, we pre\-trained an NBM, denoted asfNBMf\_\{\\textnormal\{\{NBM\}\}\}, to achieve high accuracy\. Subsequently, we re\-trained only its last layer, i\.e\., the “feature specific” layer in Fig\.[2](https://arxiv.org/html/2609.05946#S2.F2), using theMONBMframework\. In our instantiations, each individualgig\_\{i\}in the initialized populationG=\{g1,…,gτ\}G=\\\{g\_\{1\},\.\.\.,g\_\{\\tau\}\\\}is the pre\-trained NBMfNBMf\_\{\\textnormal\{\{NBM\}\}\}, and the corresponding last layer matrix ofGGto be optimized is denoted asW=\{w1,…,wτ\}\\textnormal\{\{W\}\}=\\\{\\textnormal\{\{w\}\}\_\{1\},\.\.\.,\\textnormal\{\{w\}\}\_\{\\tau\}\\\}\. This partial\-retraining strategy facilitates the practical application of EMO to large\-scale neural models by reducing optimization complexity, making EMO more applicable in modern deep learning settings\.
#### 3\.3\.2Evaluation Metrics as Optimization Objectives
In our instantiations, four evaluation metrics are explicitly treated as optimization objectives \(corresponding to theℳi\\mathcal\{M\}\_\{i\}components in Eq\. \([10](https://arxiv.org/html/2609.05946#S3.E10)\)\), addressing accuracy, fairness, local interpretability, and global interpretability:
1. \(1\)Accuracy: Cross\-entropy \(CE\) was used during optimization\. The accuracy metric \(ACC\) was used for experimental analysis and comparisons\.
2. \(2\)Fairness: The widely used group fairness metricDP\[[15](https://arxiv.org/html/2609.05946#bib.bib17),[8](https://arxiv.org/html/2609.05946#bib.bib18)\]was used to evaluate fairness\. Obviously, other fairness metrics\[[2](https://arxiv.org/html/2609.05946#bib.bib13)\]can also be considered\.
3. \(3\)Local Interpretability \(LI\): Local interpretability was measured using the G\-mean\[[13](https://arxiv.org/html/2609.05946#bib.bib34)\]of the complexity metricsμC\\mu\_\{C\}\[[7](https://arxiv.org/html/2609.05946#bib.bib22)\]andGG\[[9](https://arxiv.org/html/2609.05946#bib.bib23)\], which is defined as: LI=μC×G\.LI=\\sqrt\{\\mu\_\{C\}\\times G\}\.\(11\)
4. \(4\)Global Interpretability \(GI\): Global interpretability is evaluated based on the proposed smoothness and monotonicity metrics\. For continuous features, smoothness is always required, and monotonicity is imposed based on prior knowledge\. No requirements are imposed on discrete features\. TheGImetric is defined as: GI=∑i=1d\{ℳs\(fi\),iffirequires smooth,ℳs\(fi\)\+ℳm↑\(fi\),iffirequires smooth and↑,ℳs\(fi\)\+ℳm↓\(fi\),iffirequires smooth and↓,0,iffiis a discrete feature\.\\footnotesize GI=\\sum\_\{i=1\}^\{d\}\\begin\{cases\}\\mathcal\{M\}\_\{s\}\(f\_\{i\}\),&\\text\{if \}f\_\{i\}\\text\{ requires smooth\},\\\\ \\mathcal\{M\}\_\{s\}\(f\_\{i\}\)\+\\mathcal\{M\}\_\{m\\uparrow\}\(f\_\{i\}\),&\\text\{if \}f\_\{i\}\\text\{ requires smooth and\}\\uparrow,\\\\ \\mathcal\{M\}\_\{s\}\(f\_\{i\}\)\+\\mathcal\{M\}\_\{m\\downarrow\}\(f\_\{i\}\),&\\text\{if \}f\_\{i\}\\text\{ requires smooth and\}\\downarrow,\\\\ 0,&\\text\{if \}f\_\{i\}\\text\{ is a discrete feature\}\.\\end\{cases\}\(12\) Here,↑\\uparrowand↓\\downarrowindicate monotonically increasing and decreasing, respectively\. It is worth noting that in this paper, to simplify optimization and align withLI, we combine smoothness and monotonicity into a singleGIobjective, despite their potential conflict\. They remain separable within theMONBMframework if required\.
The optimal values forCE,DP,LI, andGImetrics are 0\.0, while the optimalACCis 100%\\%\.
#### 3\.3\.3Optimizer
In our instantiations, we used the stochastic ranking algorithm \(SRA\)\[[27](https://arxiv.org/html/2609.05946#bib.bib31)\]as the environment selection strategyπSRA\\pi\_\{\\textnormal\{\{SRA\}\}\}\.SRAuses stochastic ranking to balance different metrics\[[43](https://arxiv.org/html/2609.05946#bib.bib30)\]and can select individuals with good convergence and diversity\[[27](https://arxiv.org/html/2609.05946#bib.bib31)\]\. For the reproduction strategyξ\\xi, we used the variant of weight crossover\[[16](https://arxiv.org/html/2609.05946#bib.bib32)\]and isotropic Gaussian perturbation\[[38](https://arxiv.org/html/2609.05946#bib.bib33)\], denoted asξc\\xi\_\{c\}andξp\\xi\_\{p\}, respectively\. Since all four objectives considered in our instantiations are differentiable, partial training\[[54](https://arxiv.org/html/2609.05946#bib.bib27),[55](https://arxiv.org/html/2609.05946#bib.bib28)\]was applied to assist in exploration\.
It is worth noting that although all objective functions are differentiable and could in principle be combined into a scalarized loss, we adopt the EMO framework because the goal of this work is not only to obtain a single optimized model but also to systematically explore the trade\-offs among accuracy, fairness, local, and global interpretability\. Scalarization typically requires careful manual adjustment of multiple regularization weights, and the relationship between these weights and the resulting model behavior can be complex when several competing objectives are involved\. By contrast, EMO naturally produces a diverse set of Pareto\-optimal solutions in a single optimization process\. This allows stakeholders to select models according to different practical priorities and provides a structured basis for analyzing relationships among multiple trustworthiness dimensions\. In addition, using an EMO framework maintains flexibility for future extensions where certain fairness or interpretability metrics may not be differentiable, thereby improving the general applicability of the proposed framework\.
Algorithm 2Instantiation of theMONBMframework1:A training dataset
𝓓train\\bm\{\\mathcal\{D\}\}\_\{\\textnormal\{\{train\}\}\}, an initialized NBM set
G=\{g1,…,gτ\}G=\\\{g\_\{1\},\.\.\.,g\_\{\\tau\}\\\}and its last layer
W=\{w1,…,wτ\}\\textnormal\{\{W\}\}=\\\{\\textnormal\{\{w\}\}\_\{1\},\.\.\.,\\textnormal\{\{w\}\}\_\{\\tau\}\\\}, a set of evaluation metrics
𝓜=\{ℳ1,⋯,ℳk\}\\bm\{\\mathcal\{M\}\}=\\\{\\mathcal\{M\}\_\{1\},\\cdots,\\mathcal\{M\}\_\{k\}\\\}, an environment selection strategy
πSRA\\pi\_\{\\textnormal\{\{SRA\}\}\}, reproduction strategies
ξc\{\\xi\_\{c\}\}and
ξp\\xi\_\{p\}, and parameter
NNfor partial training\.
2:The set of optimized NBMs
G∗G^\{\*\}\.
3:
𝒎\\bm\{m\}←\\leftarrowEvaluate the value of metrics
𝓜\\bm\{\\mathcal\{M\}\}on
GGusing
𝓓train\\bm\{\\mathcal\{D\}\}\_\{\\textnormal\{\{train\}\}\}\.
4:whileiteration termination conditions not metdo
5:
Gbest′′G\_\{best\}^\{\\prime\\prime\}←\{\\leftarrow\}∅\\emptyset\.
6:for
i∈\{1,…,k\}i\\in\\\{1,\.\.\.,k\\\}do
7:Gbest\_i′G\_\{best\\\_i\}^\{\\prime\}←\{\\leftarrow\}Select the topNNbest models onℳi\\mathcal\{M\}\_\{i\}fromGG\.
8:Gbest\_i′′G\_\{best\\\_i\}^\{\\prime\\prime\}←\{\\leftarrow\}GenerateNNmodels fromGbest\_i′G\_\{best\\\_i\}^\{\\prime\}according to the reproduction strategyξp\\xi\_\{p\}\.
9:Gbest\_i′′G\_\{best\\\_i\}^\{\\prime\\prime\}←\{\\leftarrow\}Partially train\[[54](https://arxiv.org/html/2609.05946#bib.bib27),[55](https://arxiv.org/html/2609.05946#bib.bib28)\]Gbest\_i′′G\_\{best\\\_i\}^\{\\prime\\prime\}using𝓓train\\bm\{\\mathcal\{D\}\}\_\{\\textnormal\{\{train\}\}\}on the metricℳi\\mathcal\{M\}\_\{i\}\.
10:
Gbest′′G\_\{best\}^\{\\prime\\prime\}←\{\\leftarrow\}Gbest′′∪Gbest\_i′′G\_\{best\}^\{\\prime\\prime\}\\cup G\_\{best\\\_i\}^\{\\prime\\prime\}\.
11:endfor
12:
λ=τ−k×N\\lambda=\\tau\-k\\times N\.
13:G′G^\{\\prime\}←\{\\leftarrow\}Selectλ\\lambdamodels fromGGaccording to𝒎\\bm\{m\}and environment selection strategyπSRA\\pi\_\{SRA\}\.
14:G′′G^\{\\prime\\prime\}←\{\\leftarrow\}Generateλ\\lambdamodels fromG′G^\{\\prime\}according to the reproduction strategiesξc\\xi\_\{c\}andξp\\xi\_\{p\}\.
15:
G′′=G′′∪Gbest′′G^\{\\prime\\prime\}=G^\{\\prime\\prime\}\\cup G\_\{best\}^\{\\prime\\prime\}\.
16:𝒎′′\\bm\{m\}^\{\\prime\\prime\}←\\leftarrowEvaluate the value of the metrics𝓜\\bm\{\\mathcal\{M\}\}onG′′G^\{\\prime\\prime\}using𝓓train\\bm\{\\mathcal\{D\}\}\_\{\\textnormal\{\{train\}\}\}\.
17:<g1,𝒎1\>,…,<gτ,𝒎τ\><g\_\{1\},\\bm\{m\}\_\{1\}\>,\\dots,<g\_\{\\tau\},\\bm\{m\}\_\{\\tau\}\>←\\leftarrowSelectτ\\taumodels fromG∪G′′G\\cup G^\{\\prime\\prime\}based on𝒎\\bm\{m\}and𝒎′′\\bm\{m\}^\{\\prime\\prime\}through the environment selection strategyπSRA\\pi\_\{\\textnormal\{\{SRA\}\}\}and then updateG=\{g1,…,gτ\}G=\\\{g\_\{1\},\\dots,g\_\{\\tau\}\\\}and𝒎\\bm\{m\}, accordingly\.
18:endwhile
19:returnthe non\-dominated solutions in the set of optimized NBMs
GGas
G∗G^\{\*\}\.
The detailed process of the instantiation based on theMONBMframework is presented in Algorithm[2](https://arxiv.org/html/2609.05946#alg2)\. In the evolutionary loop \(lines 2\-16\), for each evaluation metricℳi\\mathcal\{M\}\_\{i\}, theNNbest\-performing individuals are selected and trained according to the reproduction strategyξp\\xi\_\{p\}andℳi\\mathcal\{M\}\_\{i\}\-based partial training, resulting ink×Nk\\times Noffspring \(lines 4\-9\)\. Subsequently,λ\\lambdaoffspring are generated based on the environmental selection strategyπSRA\\pi\_\{\\textnormal\{\{SRA\}\}\}and reproduction strategiesξc\\xi\_\{c\}andξp\\xi\_\{p\}\(lines 11\-12\)\. Then, all offspring are evaluated \(line 14\), and a new populationGGis produced according to the environmental selection strategyπSRA\\pi\_\{\\textnormal\{\{SRA\}\}\}\(line 15\)\. This iterative cycle continues until the convergence criterion of 200 generations is met\.
The overall time complexity per generation isO\(kτ2\+τ⋅n⋅DlogD\)O\(k\\tau^\{2\}\+\\tau\\cdot n\\cdot D\\log D\), wherekkdenotes the number of optimization objectives,τ\\taurepresents the population size,DDis the feature dimension, andnnis the data size\. A more detailed derivation of both time and space complexity is provided in Section S\-II\.2 of theSupplementary Material\. In addition, further analysis of the framework’s generalization behavior and robustness is presented in Section S\-II\.3 of theSupplementary Material\.
To analyze the relationships between dimensions and compare them with state\-of\-the\-art methods, we considered four instantiations, each with different optimization objectives: \(1\) optimizing theCEandDPmetrics; \(2\) optimizing theCEandLImetrics; \(3\) optimizing theCEandGImetrics; and \(4\) optimizing theCE,DP,LI, andGImetrics\. These four instantiations are referred to asMONBM\-CD,MONBM\-CL,MONBM\-CG, andMONBM\-CDLG, respectively\.
## 4Experimental Studies
In this section, we first give the experimental setup and then answer the three research questions posed through experiments\.
### 4\.1Experimental Setup
This subsection describes the experimental setup, including the datasets used, baseline methods for comparison, parameter settings, and evaluation criteria\.
#### 4\.1\.1Datasets
We consider six datasets that are widely used in the field of fair ML\[[25](https://arxiv.org/html/2609.05946#bib.bib39)\], namelyAdult\[[14](https://arxiv.org/html/2609.05946#bib.bib35)\],Bank\[[14](https://arxiv.org/html/2609.05946#bib.bib35)\],COMPAS\[[3](https://arxiv.org/html/2609.05946#bib.bib36)\],Default\[[14](https://arxiv.org/html/2609.05946#bib.bib35)\],Dutch\[[46](https://arxiv.org/html/2609.05946#bib.bib38)\], andLSAT\[[49](https://arxiv.org/html/2609.05946#bib.bib37)\], whose specific information can be found in Tables S\-I of theSupplementary Material\. The same pre\-processing was applied to each dataset: label encoding of discrete features, followed by normalization of all features with Z\-score\. Then, they were randomly divided into training set𝓓train\\bm\{\\mathcal\{D\}\}\_\{\\textnormal\{\{train\}\}\}and test set𝓓test\\bm\{\\mathcal\{D\}\}\_\{\\textnormal\{\{test\}\}\}in a ratio of 4:1\.
#### 4\.1\.2Baseline Methods
To validate the effectiveness of theMONBMframework to answerRQ3, a comparison was made with the following baseline methods:
1. \(1\)ForMONBM\-CDinstantiation, we considered the following baselines: \(a\) applying the method proposed in\[[23](https://arxiv.org/html/2609.05946#bib.bib20)\]to specifically improve the fairness of GAM to NBM, and referring to it asFair\-NBM; \(b\) combining the existing approaches to improve the fairness of ML models at the pre\-processing stage, i\.e\., learning fair representation \(LFR\)\[[56](https://arxiv.org/html/2609.05946#bib.bib8)\]and reweighting\[[19](https://arxiv.org/html/2609.05946#bib.bib9)\], as well as at the post\-processing stage, i\.e\., reject option classification \(ROC\)\[[20](https://arxiv.org/html/2609.05946#bib.bib19)\], linear post\-processing \(LPP\)\[[50](https://arxiv.org/html/2609.05946#bib.bib53)\], Oracle\[[51](https://arxiv.org/html/2609.05946#bib.bib54)\], and FRAPPÉ\[[45](https://arxiv.org/html/2609.05946#bib.bib59)\]with NBM, calledLFR\-NBM,Re\-NBM,ROC\-NBM,LPP\-NBM,Oracle\-NBM, andFRAPPÉ\-NBM, respectively; and \(c\) a method based on multi\-objective optimization to improve the fairness of MLP in the in\-processing stage, calledFairEMOL\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]\.
2. \(2\)ForMONBM\-CLinstantiation, we compared it with three approaches to improveLIbased on feature selection, i\.e\.,SNAM\[[52](https://arxiv.org/html/2609.05946#bib.bib16)\],GAMI\-Net\[[53](https://arxiv.org/html/2609.05946#bib.bib15)\], andNPLAM\[[62](https://arxiv.org/html/2609.05946#bib.bib60)\]\.
The instantiationsMONBM\-CGandMONBM\-CDLGwere not compared to baseline methods due to the absence of existing work that considers global interpretability\.
#### 4\.1\.3Parameter Settings
For the NBM, the parameter settings and hyper\-parameter search approach were consistent with the original configuration\[[42](https://arxiv.org/html/2609.05946#bib.bib11)\], as detailed in Section S\-III\.1 of theSupplementary Material\. The optimal hyper\-parameter configurations obtained from the search were applied to the pre\-trained NBMfNBMf\_\{\\textnormal\{\{NBM\}\}\}as well as to the baseline methodsFair\-NBM,LFR\-NBM,Re\-NBM,ROC\-NBM,LPP\-NBM,Oracle\-NBM, andFRAPPÉ\-NBM, except for differences in the number of training epochs\. The parameter settings for all baseline methods are provided in Section S\-III\.2 of theSupplementary Material\. Additionally, the computational budget comparisons and wall\-clock time comparisons are presented in Sections S\-II\.1 and S\-II\.4 of theSupplementary Material, respectively\.
In our instantiations of theMONBMframework, the population sizeτ\\tauwas set to 100, the mutation strength of the reproduction strategyξp\\xi\_\{p\}was set to 0\.01, and the iteration termination condition of the evolutionary loop was set to a maximum of 200 generations\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]\. This iteration limit was established through preliminary experiments to ensure sufficient convergence across all benchmark datasets\. Exploring adaptive early stopping mechanisms based on convergence or stagnation is identified as a potential direction for future work\. The hyper\-parameterNNfor partial training was set to 3\. We divided the dataset using 15 different random seeds and conducted independent trials\.
The parameterkkin the smoothness metricℳs\\mathcal\{M\}\_\{s\}serves as a tolerance threshold that controls the permissible degree of localized fluctuations in shape functions\. We definekkrelative to the average smoothness of the pre\-trained NBMfNBMf\_\{NBM\}, i\.e\.,k=η⋅1\(n−1\)∑j=1n−1\|fNBM,i\(xi,j\+1\)−fNBM,i\(xi,j\)\|fNBM,i,max−fNBM,i,mink=\\eta\\cdot\\frac\{1\}\{\(n\-1\)\}\\sum\_\{j=1\}^\{n\-1\}\\frac\{\|f\_\{\\textnormal\{\{NBM\}\},i\}\(x\_\{i,j\+1\}\)\-f\_\{\\textnormal\{\{NBM\}\},i\}\(x\_\{i,j\}\)\|\}\{f\_\{\\textnormal\{\{NBM\}\},i,\\textnormal\{max\}\}\-f\_\{\\textnormal\{\{NBM\}\},i,\\textnormal\{min\}\}\}\. Here,η\\etais a scaling factor that modulates the stringency of the smoothness constraint\. In our primary experiments, we setη=1\.0\\eta=1\.0, implying that the optimized model is expected to maintain a local smoothness at every interval that is at least as consistent as the global average smoothness of the pre\-trained NBM\. Furthermore, a detailed discussion regarding the impact of different values for the scaling factorη\\etais provided in Section S\-IV of theSupplementary Material\. Consistent with intuition, smallerη\\etavalues correspond to stricter smoothness requirements \(effectively reducing the tolerance thresholdkk\), resulting in more regularized and smoother shape functions\. This design provides practical flexibility, allowing stakeholders to adjustη\\etaaccording to domain knowledge or interpretability requirements\.
For the monotonicity metricℳm\\mathcal\{M\}\_\{m\}, we specified a subset of features required to satisfy monotonicity, as detailed in Section S\-V of theSupplementary Material\. This setting is intended to evaluate whether the proposed metrics and training framework can guide shape functions toward user\-specified trends, rather than to imply that monotonicity is universally required\. In practical applications, the choice of whether a feature should satisfy monotonicity depends on task requirements, domain knowledge, and stakeholder preferences\.
#### 4\.1\.4Evaluation Criteria
To answerRQ3, we compared all non\-dominated solutions obtained byMONBMwith baseline methods on the corresponding objectives \(e\.g\.,ACCandDPmetrics on theMONBM\-CDinstantiation\)\. In comparison with population\-based methods \(i\.e\.,FairEMOL\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]\), we used three commonly used metrics: hypervolume \(HV\)\[[64](https://arxiv.org/html/2609.05946#bib.bib42)\], maximum spread \(MS\)\[[63](https://arxiv.org/html/2609.05946#bib.bib44)\], and pure diversity \(PD\)\[[44](https://arxiv.org/html/2609.05946#bib.bib43)\]to assess the convergence and diversity of the solution set\. Among them,HVis the only Pareto\-compliant metric that is used to comprehensively assess the performance of the solution set\[[29](https://arxiv.org/html/2609.05946#bib.bib45)\]\.
To calculate theHVvalue, we first identify the non\-dominated solutions from all solutions obtained by all methods in all generations under the current data split\. These solutions are used as the pseudo\-Pareto front to normalize all objective values\. Then, we used\(1\.1,…,1\.1\)\(1\.1,\.\.\.,1\.1\)as the reference point\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]to compute the final hypervolume, measuring the dominated space relative to the normalized front\. TheMSandPDmetrics, on the other hand, focus on evaluating the diversity of the solution set\.MSevaluates the maximum expansion of each objective, whilePDmeasures the degree to which each solution differs from the others\[[29](https://arxiv.org/html/2609.05946#bib.bib45)\]\. All three metrics are the larger, the better\.
Additionally, we comparedMONBMwith population\-based and single solution\-based approaches to assess mutual dominance relationships, following the comparison methodologies outlined in the literature\[[47](https://arxiv.org/html/2609.05946#bib.bib1)\]and\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\], respectively\. The comparison methods in\[[47](https://arxiv.org/html/2609.05946#bib.bib1)\]and\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]are used to evaluate the mutual dominance relationship between two solution sets and between a solution set and a single solution, respectively\. A more detailed description of the comparative methods is given in Section S\-VI of theSupplementary Material\.
### 4\.2Experimental Results
In this subsection, we address the three research questions introduced earlier through empirical evaluation\.
#### 4\.2\.1Validation of the Proposed Metrics \(RQ1\)
To validate the proposed metrics, we present the global explanations of continuous features generated by models corresponding to two extreme points and one “middle point” obtained fromMONBM\-CG\. In this paper, the “middle point” refers to a model that performs moderately on the trustworthy dimension, selected from all the knee points\. The knee point represents the most concave region of the Pareto front and is considered the best trade\-off for its neighborhood in the absence of specific preferences\[[59](https://arxiv.org/html/2609.05946#bib.bib46)\]\.
In this paper, we referred to the method of Zhanget al\.\[[59](https://arxiv.org/html/2609.05946#bib.bib46)\]for selecting knee points, where the parameterrgr\_\{g\}was set to 0\.02\. Taking theMONBM\-CGon theAdultdataset as an example, the extreme points ofCEandGIare at values of 0\.1713 and 0\.0303, respectively, on theGImetric of the training set\. Therefore, the “middle point” is the NBM where theGIvalue is closest to the average of the two extremes on theGImetric among all knee points, i\.e\.,\(0\.1713\+0\.0303\)/2=0\.1008\(0\.1713\+0\.0303\)/2=0\.1008\.
Due to space limitations, we provided results for theAdultandBankdatasets, as shown in Fig\.[4](https://arxiv.org/html/2609.05946#S4.F4)\. The corresponding results for the remaining datasets are presented in Fig\. S\-3 of theSupplementary Material\. To further illustrate the continuity of the trade\-off surface, Fig\. S\-4 in theSupplementary Materialpresents the global explanations for the complete set of non\-dominated knee\-point solutions across all datasets\.
\(a\)Global explanation results for continuous features ofMONBM\-CGinstantiation on theAdultdataset
\(b\)Global explanation results for continuous features ofMONBM\-CGinstantiation on theBankdataset
Figure 4:Comparison of global explanations of continuous features to validate the proposed smoothness and monotonicity metrics\. For each dataset, the blue, red, and green lines represent the global explanation results of the continuous features by the model corresponding to the bestCE, “middleGI,” and bestGIpoints obtained by theMONBM\-CGinstantiation in one arbitrary trial, respectively\. Each subfigure corresponds to the global explanation of a feature, with the x\-axis representing feature values and the y\-axis showing the outputs of the shape functions\. Each continuous feature has a smoothness requirement, and features highlighted in red also have a monotonicity requirement\.Across all datasets, a consistent pattern is observed: asGIimproves, the learned shape functions become progressively smoother while preserving their primary functional trends\. Unnecessary local oscillations and abrupt variations are gradually reduced, without collapsing the overall feature–target relationships\. This demonstrates that the proposed framework provides a spectrum of balanced solutions that mitigate excessive fluctuations while maintaining meaningful structure in the explanations\. In addition, improvements inGIare accompanied by clearer monotonic tendencies in relevant features\. Although strict smoothness or perfect monotonicity cannot be guaranteed for continuous functions learned by NN\-based GAMs, the empirical results show that promoting approximate smoothness and monotonic trends is both feasible and effective in practice\.
Overall, the explanation results demonstrate that our method can yield NBMs with high global interpretability\. Moreover, these improvements in visual interpretability are consistently aligned with the observed changes in theGImetric values across all datasets, as will be demonstrated in Table[S\-VII](https://arxiv.org/html/2609.05946#S7.T7), further validating the effectiveness of the proposed metrics\.
#### 4\.2\.2Analyzing the Relationship Between Dimensions \(RQ2\)
To answerRQ2, we analyzed each of the four instantiations in two steps: \(1\) selecting representative points to examine the relationship between dimensions based on the performance of the final generation on the training set; \(2\) exploring the experimental results using the global and local explanations provided by the obtained NBMs to further reveal the reasons behind these relationships\.
For the two\-objective instantiations, i\.e\.,MONBM\-CD,MONBM\-CL, andMONBM\-CG, we chose two extreme points and one “middle point” for analysis in each case\. For the four\-objective instantiationMONBM\-CDLG, the analysis first focuses on the four extreme points to characterize the primary multidimensional trade\-offs, and is then further extended to the full set of non\-dominated solutions to provide a more comprehensive validation of these relationships\.
##### Relationship Between Accuracy and Fairness \(MONBM\-CD\)
On theMONBM\-CDinstantiation, Table[2](https://arxiv.org/html/2609.05946#S4.T2)shows the performance of the models corresponding to the extreme and “middleDP” points on the test set \(the performance on the training set is shown in Table S\-V of theSupplementary Material\)\. From Table[2](https://arxiv.org/html/2609.05946#S4.T2), we can draw the following two conclusions: \(1\) Moderate fairness improvement \(defined as a 50% reduction in theDPmetric relative to its extremes\) has little effect on accuracy\. Across six datasets, the impact of reducing theDPmetrics by half on theACCis only0\.48%0\.48\\%on average; \(2\) Pursuing extreme fairness significantly affects accuracy, with the degree of impact varying across datasets\. For theAdult,COMPAS, andDutchdatasets, reducing theDPmetric to zero leads to an averageACCdecrease of17\.06%17\.06\\%, while the remaining three datasets show a much smaller average reduction of2\.15%2\.15\\%\.
Table 2:Comparison of the mean values of theACCandDPon the test set of the models corresponding to the three points chosen from the final solution set obtained by theMONBM\-CDinstantiation over 15 trials\.ACCvalues are expressed as percentages \(%\\%\)\. The best average is highlighted with a gray background\. The difference between it and the best value is shown in parentheses\.To further analyze this phenomenon, we analyzed the global explanation results of the three models\. Due to space limitations, we chose theCOMPASandLSATdatasets as representative of the significant and moderate effects of pursuing extreme fairness on accuracy, respectively, and their global explanation results are shown in Fig\.[5](https://arxiv.org/html/2609.05946#S4.F5)\. The corresponding results for the remaining datasets are shown in Fig\. S\-5 of theSupplementary Material\. As can be seen in these figures, for each dataset, compared to the extreme point of accuracy \(the blue line\), moderate fairness enhancement \(the red line\) primarily influences the sensitive attribute, thus having a low impact on accuracy\. However, the impact of further pursuing extreme fairness \(the green line\) on the features varies across datasets\. For example,COMPAS\(Fig\.[5\(a\)](https://arxiv.org/html/2609.05946#S4.F5.sf1)\) has the vast majority of features affected in the pursuit of extreme fairness, thus significantly affecting accuracy\. In contrast,LSAT\(Fig\.[5\(b\)](https://arxiv.org/html/2609.05946#S4.F5.sf2)\) changes only a few features and, therefore, has little impact on accuracy\. Furthermore, the figures indicate that to enhance fairness, many features are adjusted to contribute minimally to the decision outcome, effectively removing their influence and improving fairness\.
\(a\)Global explanation results ofMONBM\-CDinstantiation on theCOMPASdataset
\(b\)Global explanation results ofMONBM\-CDinstantiation on theLSATdataset
Figure 5:Comparison of global explanations to assess the impact of improved fairness on each feature\. For each dataset, the blue, red, and green lines represent the global explanation results of the model corresponding to the bestCE, “middleDP”, and bestDPpoints obtained by theMONBM\-CDinstantiation in one arbitrary trial, respectively\. Each subfigure corresponds to the global explanation of a feature, with the x\-axis representing feature values and the y\-axis showing the outputs of the shape functions\. The sensitive attribute is highlighted in red\. The Pearson correlation coefficient between each attribute and the sensitive attribute is annotated in the respective subplots\.To examine the influence of training sample variability, we additionally report the mean global explanations together with their standard deviations across 15 independent trials, provided in Fig\. S\-6 of theSupplementary Material\. Overall, the functional shapes observed in the single\-run results \(Figs\.[5](https://arxiv.org/html/2609.05946#S4.F5)and S\-5\) remain highly consistent with the averaged statistics \(Fig\. S\-6\), indicating that the main explanatory patterns are stable across different training samples\. However, while the overall trends are preserved, the magnitude of the shape functions exhibits variability at moderate and extreme fairness operating points\. A more detailed analysis of this variability is provided in Section S\-IX of theSupplementary Material\.
To further investigate the robustness of these findings, we extended our analysis to a label\-aware fairness metric, equalized odds \(EOD\), in Section S\-VIII of theSupplementary Material\. As shown in Table S\-IX and Fig\. S\-2, the experimental results underEODexhibit highly consistent trade\-off trends and global explanation patterns with those ofDP\. This demonstrates that the decision logic adjustments identified byMONBMare intrinsic to the accuracy\-fairness conflict in these datasets, rather than being an artifact of a specific fairness definition\.
Overall, consistent with findings in the existing literature, a trade\-off exists between accuracy and fairness\[[58](https://arxiv.org/html/2609.05946#bib.bib7),[57](https://arxiv.org/html/2609.05946#bib.bib52)\]\. We further revealed the underlying reasons\. Specifically, moderate improvement primarily necessitates adjusting the impact of the sensitive attribute on decision results, with minimal effects on accuracy\. Moreover, to achieve extreme fairness, there are differences across datasets; some require substantial adjustments to multiple features, leading to significant accuracy losses, while others necessitate fewer modifications, resulting in relatively minor accuracy declines\.
##### Relationship Between Accuracy and Local Interpretability \(MONBM\-CL\)
On theMONBM\-CLinstantiation, Table[S\-VI](https://arxiv.org/html/2609.05946#S7.T6)shows the performance of the models corresponding to the extreme and “middleLI” points on the test set \(the performance on the training set is shown in Table S\-VI of theSupplementary Material\)\. From Table[S\-VI](https://arxiv.org/html/2609.05946#S7.T6), we can draw the following conclusions: \(1\) similar to the fairness metricDP, a moderate improvement inLI\(defined as a 50% reduction in theLImetric relative to its extremes\) has a relatively low impact onACC\(2\.03%2\.03\\%decrease on average\), but pursuing extremeLIsignificantly affectsACC\(12\.59%12\.59\\%decrease on average\); \(2\) the impact of improvingLIonACCis significantly larger than that of improving the fairness metricDP\(0\.48%0\.48\\%and9\.61%9\.61\\%decrease on average, respectively\)\.
Table 3:Comparison of the mean values of theACCandLIon the test set of the models corresponding to the three points chosen from the final solution set obtained by theMONBM\-CLinstantiation over 15 trials\.ACCvalues are expressed as percentages \(%\\%\)\. The best average is highlighted with a gray background\. The difference between it and the best value is shown in parentheses\.\(a\)Average local explanation ofMONBM\-CLinstantiation on theBankdataset
\(b\)Average local explanation ofMONBM\-CDinstantiation on theBankdataset
Figure 6:Comparison of the average local explanations of the test set to assess the impact of improvedLIon each feature\. The three columns respectively display the average local explanation results of the model corresponding to: \(1\) the bestCEpoint, \(2\) the “middleLI” or “middleDP” point, and \(3\) the bestLIor bestDPpoint, obtained by theMONBM\-CLorMONBM\-CDinstantiation on theBankdataset in one arbitrary trial\. For each local explanation, the y\-axis represents the features and the x\-axis represents the average of the absolute values of the shape function outputs corresponding to the features for all data points in the test set\.To further analyze this phenomenon, we analyzed the local explanation results of the three models and compared them with the corresponding results from theMONBM\-CDinstantiation\. To ensure that the explanation results are representative, we present the average absolute values of local explanation results for all data points in the test set𝓓test\\bm\{\\mathcal\{D\}\}\_\{\\textnormal\{\{test\}\}\}\. Due to space limitations, we selected theBankdataset as a representative, and the results are shown in Fig\.[6](https://arxiv.org/html/2609.05946#S4.F6), while the results for other datasets are shown in Fig\. S\-7 in theSupplementary Material\. From Fig\.[6\(a\)](https://arxiv.org/html/2609.05946#S4.F6.sf1)and the left column of Fig\. S\-7, it can be seen that compared with pursuing extremeACC, moderateLIenhancement still allows most features to play a role in model decision\-making, thus maintaining the model’s predictive ability\. In contrast, extremeLIdrastically weakens or removes the influence of most features, sharply reducing performance\.
Moreover, comparing with Fig\.[6\(b\)](https://arxiv.org/html/2609.05946#S4.F6.sf2)and the right column of Fig\. S\-7, theLIobjective induces more pronounced shifts in relative feature importance thanDP\. This pronounced disparity stems from the fundamental distinction between their underlying optimization mechanisms: whereas fairness constraints primarily refine the contributions of the sensitive attribute and its proxies, theLIobjective enforces structural sparsity across the entire feature set\. As the model approaches the extremeLIfrontier, the framework increasingly favors a small number of highly informative features that can maintain predictive capability with minimal complexity\. A representative example of this transition is observed in theAdultdataset \(Fig\. S\-7\(a\)\), where the feature “fnlwgt” transitions from a relatively minor contributor in a multi\-feature setting to a dominant predictor under strong sparsity constraints\. Rather than indicating a limitation, this behavior helps reveal the hierarchy of feature importance and clarifies the performance limits of highly simplified models\.
Furthermore, to evaluate the stability of these local explanations across different data partitions, we provide the mean local importance scores and their standard deviations over 15 independent trials in Fig\. S\-8 of theSupplementary Material\. While the overall feature rankings remain consistent with our single\-run observations, we notice a non\-negligible variance at the extreme points of fairness and local interpretability, which is further analyzed in Section S\-IX of theSupplementary Material\.
Overall, there is a significant trade\-off between accuracy and local interpretability\. Moderate improvement preserves the model’s reliance on most features, thereby preserving predictive power\. However, pursuing extremeLIrequires removing most of the features’ contribution to the model’s decisions, significantly reducing the model’s predictive power\. Moreover, Fig\.[6\(a\)](https://arxiv.org/html/2609.05946#S4.F6.sf1)and the left column of Fig\. S\-7 also visually illustrate that our method successfully produces NBMs with varying degrees ofLI\.
##### Relationship Between Accuracy and Global Interpretability \(MONBM\-CG\)
On theMONBM\-CGinstantiation, Table[S\-VII](https://arxiv.org/html/2609.05946#S7.T7)shows the performance of the models corresponding to the extreme and “middleGI” points on the test set \(the performance on the training set is shown in Table S\-VII of theSupplementary Material\)\. From Table[S\-VII](https://arxiv.org/html/2609.05946#S7.T7), we observe that, unlike fairness and local interpretability, improving global interpretability has very little impact on accuracy\. ModerateGIimprovement \(defined as a 50% reduction in theGImetric relative to its extreme values\) results in an averageACCchange of only 0\.12% across datasets, while pursuing extremeGIleads to an averageACCchange of 0\.85%\. Notably, on theBankandCOMPASdatasets, moderately or pursuing extremeGIeven slightly improves theACC\. However, this does not imply that there is no trade\-off between accuracy and global interpretability\. We found that increasingGIon both theBankandCOMPASdatasets leads to worseCEmetrics, even though it may bring about a slight improvement inACC\. The reason behind this is more likely to be the non\-strict linear relationships betweenCEandACC, but it also suggests that improvingGIhas a smaller impact on accuracy\.
Table 4:Comparison of the mean values of theACCandGIon the test set of the models corresponding to the three points chosen from the final solution set obtained by theMONBM\-CGinstantiation over 15 trials\.ACCvalues are expressed as percentages \(%\\%\)\. The best average is highlighted with a gray background\. The difference between it and the best value is shown in parentheses\.To further analyze this phenomenon, we presented the global and average local explanation results of the three corresponding models\. To facilitate comparison with theMONBM\-CLandMONBM\-CDinstantiations, we also selected theBankdataset as a representative, with its global and local explanation results shown in Figs\.[7](https://arxiv.org/html/2609.05946#S4.F7)and[8](https://arxiv.org/html/2609.05946#S4.F8), respectively, while the corresponding results for the other datasets are shown in Fig\. S\-9 and Fig\. S\-11 in theSupplementary Material\.
Figure 7:Comparison of global explanations to assess the impact of improvedGIon each feature\. The blue, red, and green lines represent the global explanation results of the model corresponding to the bestCE, “middleGI,” and bestGIpoints obtained by theMONBM\-CGinstantiation on theBankdataset in one arbitrary trial, respectively\. Each subfigure corresponds to the global explanation of a feature, with the x\-axis representing feature values and the y\-axis showing the outputs of the shape functions\. The continuous features are highlighted in red\.Figure 8:Comparison of the average local explanations of the test set to assess the impact of improvedGIon each feature\. The three columns represent the average local explanation results of the model corresponding to the bestCE, “middleGI,” and bestGIpoints obtained by theMONBM\-CGinstantiation on theBankdataset in one arbitrary trial, respectively\. For each local explanation, the y\-axis represents the features and the x\-axis represents the average of the absolute values of the shape function outputs corresponding to the features for all data points in the test set\.From Fig\.[8](https://arxiv.org/html/2609.05946#S4.F8)and Fig\. S\-11, it can be seen that, compared with the extreme points ofACC\(first column\), moderateGIimprovements \(second column\) and pursuit of extremeGI\(third column\) have a significantly smaller impact on features than the corresponding results ofMONBM\-CLandMONBM\-CDinstantiations in Fig\.[6](https://arxiv.org/html/2609.05946#S4.F6)and Fig\. S\-7\. Similar observations can also be seen in Fig\.[7](https://arxiv.org/html/2609.05946#S4.F7)and Fig\. S\-9, where the variation of the global explanation results is very small in each column\. This difference can be attributed to the fact that, on the one hand,GIhas little effect on discrete features \(i\.e\., features not highlighted in red\)\. On the other hand, for continuous features \(i\.e\., features highlighted in red\),GIonly adjusts regions that are unsmoothed, mutated, or partially violate monotonicity\. Consequently, the overall impact of global interpretability on accuracy is small\. The stability of both global and local explanations forMONBM\-CGis verified in Figs\. S\-10 and S\-12 of theSupplementary Material, respectively\. Compared with theMONBM\-CDandMONBM\-CLinstantiations,MONBM\-CGexhibits lower variance in both global shape functions and local feature importance, even near extremeGIoperating points\. Additional discussion of this phenomenon is provided in Section S\-IX of theSupplementary Material\.
Overall, the trade\-off between accuracy and global interpretability is weak\. The reason behind this phenomenon is thatGImainly affects the continuous features, especially by adjusting the regions that are not smooth or violate monotonicity, and these adjustments usually do not significantly change the overall predictive performance of the model\. As a result, the performance trade\-offs faced by the model in improving global interpretability are less pronounced than those faced in improving fairness or local interpretability\.
##### Relationship Between Accuracy, Fairness, Local and Global Interpretability \(MONBM\-CDLG\)
On theMONBM\-CDLGinstantiation, Table[S\-VIII](https://arxiv.org/html/2609.05946#S7.T8)shows the performance of the models corresponding to the extreme points on the test set \(the performance on the training set is shown in Table S\-VIII of theSupplementary Material\)\. This table confirms the prior conclusion that global interpretability has the least impact on accuracy, followed by fairness, while local interpretability has the greatest impact\. More importantly, Table[S\-VIII](https://arxiv.org/html/2609.05946#S7.T8)reveals that on the vast majority of datasets \(especiallyAdult,Bank,COMPAS, andLSAT\), improvements in fairness, local interpretability, or global interpretability individually are accompanied by improvements in fairness and local interpretability, especially fairness\. For example, compared to the extreme points ofACC\(the first row of each dataset in Table[S\-VIII](https://arxiv.org/html/2609.05946#S7.T8)\), the pursuit ofLI\(third row\) also positively impacts theDPmetric\. However, theDutchdataset is a notable exception; specifically, the pursuit ofDPinstead worsensLI, and the pursuit ofGIworsens bothDPandLI\. Furthermore, the pursuit ofLIsignificantly worsens theDPmetric by up to 0\.928, indicating extreme unfairness\.
Table 5:Comparison of the mean values of theACC,DP,LI, andGIon the test set of the models corresponding to the four extreme points chosen from the final solution set obtained by theMONBM\-CDLGinstantiation over 15 trials\.ACCvalues are expressed as percentages \(%\\%\)\. The best average is highlighted with a gray background\. The difference between it and the best value is shown in parentheses\. The “Majority Baseline” row represents the accuracy of a constant classifier that consistently predicts the majority class for each dataset\.\(a\)Average local explanation ofMONBM\-CDLGinstantiation on theAdultdataset
\(b\)Average local explanation ofMONBM\-CDLGinstantiation on theDutchdataset
Figure 9:Comparison of the average local explanations of the test set to explore the interrelationships among the four objectives\. For each dataset, the four columns represent the average local explanation results of the model corresponding to the bestCE, bestDP, bestLI, and bestGIpoints obtained by theMONBM\-CDLGinstantiation in one arbitrary trial, respectively\. For each local explanation, the y\-axis represents the features and the x\-axis represents the average of the absolute values of the shape function outputs corresponding to the features for all data points in the test set\.\(a\)Distribution ofMONBM\-CDLGsolutions on theAdultdataset
\(b\)Distribution ofMONBM\-CDLGsolutions on theDutchdataset
Figure 10:Distribution of allMONBM\-CDLGnon\-dominated solutions across pairs of trustworthy objectives to examine their interrelationships\. Each subplot visualizes the solutions obtained from one arbitrary run of theMONBM\-CDLGinstantiation on a given dataset, showing pairwise relationships betweenDP,LI, andGI\.To further analyze this phenomenon, we compared the average local explanation results of the models corresponding to the four extreme points\. Due to space limitations, we used theAdultdataset as a general representative to compare with theDutchdataset, as shown in Fig\.[9](https://arxiv.org/html/2609.05946#S4.F9)\. The corresponding results for the remaining datasets are shown in Fig\. S\-13 of theSupplementary Material\.
Beyond the extreme\-point analysis, we further visualized the distribution of all non\-dominatedMONBM\-CDLGsolutions across pairs of trustworthy dimensions \(e\.g\.,DPvs\.LI,DPvs\.GI, andLIvs\.GI\), as shown in Fig\.[10](https://arxiv.org/html/2609.05946#S4.F10)forAdultandDutch\(with additional datasets provided in Fig\. S\-15 of theSupplementary Material\)\. These distribution plots allow us to examine whether the relationships observed at the extreme points also hold for intermediate trade\-off solutions and reduce the risk of conclusions driven by isolated extreme cases\. From both the extreme\-point analysis and the full\-solution distributions, we observe the following:
- \(1\)In general, improvingDPreduces the impact of many features on the decision result, consistent with the phenomenon observed onMONBM\-CDinstantiation\. This leads to an improvement inLIsince only a few features play an important role in the decision\. Importantly, the distribution plots of all solutions confirm that this trend is not restricted to extreme solutions: on datasets such asAdult, most Pareto solutions form a clear monotonic trend where lowerDPdisparity is generally associated with betterLI\. This suggests that the observed fairness–interpretability synergy reflects a systematic structural relationship rather than an artifact of isolated extreme models\. In contrast, theDutchdataset retains many important features, resulting in no improvement inLI\.
- \(2\)Additionally, improving local interpretability reduces the impact of the vast majority of features on the decision and retains only a few important features\. Thus, in most cases, the influence of the sensitive attribute \(and its potential proxy attributes\) on the decision is eliminated, thereby improving fairness\. However, on theDutchdataset, the sensitive attribute “sex” happens to be retained as the only important feature in the pursuit ofLI, leading to decisions that significantly depend on this attribute and resulting in extreme unfairness\. The full\-solution distribution further supports this observation: unlikeAdult, theDutchsolutions exhibit a clear negative trend betweenDPandLI, indicating a persistent trade\-off across the Pareto set rather than only at extreme points\. This highlights that fairness–interpretability relationships are strongly dataset\-dependent and can reverse depending on feature–label correlations and sensitive attribute structure\.
- \(3\)Improving global interpretability—particularly smoothness—can moderately reduce the influence of certain features by removing irregular fluctuations in shape functions \(as illustrated in Fig\.[7](https://arxiv.org/html/2609.05946#S4.F7)and Fig\. S\-9\)\. This may indirectly support fairness andLI\. However, the distribution of all solutions shows that the improvement inLIassociated with betterGIis generally weak and scattered across datasets, indicating only partial alignment between the two interpretability notions\. Conversely, improvements inLIrarely lead to noticeable gains inGI, confirming an asymmetric interaction\. Overall, there is still a clear trade\-off between the two\. Similarly, improving global interpretability also helps fairness by reducing the influence of certain features, but enhancing fairness does not improve global interpretability\. In addition, as can be seen in Fig\.[9\(b\)](https://arxiv.org/html/2609.05946#S4.F9.sf2), on theDutchdataset,GIalso encounters a similar situation where the sensitive attribute “sex” is taken as the most important feature, leading to more unfair decision\-making results\. To better understand which component ofGIcontributes to this phenomenon, we further decompose global interpretability into smoothness and monotonicity and conduct an ablation study in Section S\-XI of theSupplementary Material\. The results suggest that the observed fairness–global interpretability conflict on theDutchdataset is primarily associated with the smoothness component, whereas the conflict between monotonicity and fairness is significantly weaker\. This finding indicates that the relationship between global interpretability and other trustworthy dimensions is not uniform but depends on the specific interpretability mechanism being optimized\.
Overall, our experimental results reveal the complex trade\-offs between accuracy, fairness, local interpretability, and global interpretability\. First, local interpretability has the greatest impact on accuracy, with extreme local interpretability often leading to a significant decrease in model performance, while moderate improvements in local interpretability have a minimal effect on accuracy\. Second, there is a clear trade\-off between fairness and accuracy\. Moderate improvements in fairness have a negligible impact on accuracy, but extreme fairness often results in significant accuracy drops, especially on certain datasets, where fairness adjustments can cause substantial variations in performance\. Global interpretability has the least impact on accuracy because it only affects the parts of the continuous features that do not satisfy smoothness and monotonicity\.
Regarding the relationships between fairness, local interpretability, and global interpretability, our experiments show that these dimensions interact differently depending on the dataset characteristics\. In most cases, increases in fairness, local interpretability, and global interpretability reduce the impact of many features on decision making, thus indirectly improving fairness and local interpretability\. In contrast, theDutchdataset serves as a valuable counterexample that improving local and global interpretability leads to extremely unfair results, and improving fairness and global interpretability also does not benefit local interpretability, highlighting the context\-dependent nature of these relationships\. This complex interplay necessitates careful consideration by researchers aiming to optimize models while ensuring fairness and interpretability\. To assess the stability of the explanations, the aggregate local explanation results across 15 trials are presented in Fig\. S\-14 of theSupplementary Material, offering additional empirical support for the observed trade\-off patterns\.
We also note that four\-objective instantiation \(MONBM\-CDLG\) generally performs worse than two\-objective instantiations on the corresponding optimization objectives\. This is a known challenge in many\-objective evolutionary optimization: as the number of objectives increases, selection pressure weakens, and the search becomes less effective, often resulting in a lower\-quality solution set—even with dedicated many\-objective selection schemes\.
The reason is that maintaining a proper balance between exploration and convergence becomes increasingly difficult as more objectives are introduced\. Therefore, introducing adaptive pressure mechanisms into the optimization process may provide a feasible way to improve search effectiveness\. To explore this possibility, we conduct a preliminary investigation by incorporating an adaptive selection pressure strategy into the underlying stochastic ranking procedure, with detailed descriptions and empirical results provided in Section S\-XII of theSupplementary Material\. Although the results indicate that adaptive pressure adjustment can partially mitigate performance degradation in high\-dimensional objective spaces, further investigation of more specialized many\-objective evolutionary algorithms and adaptive selection operators remains an important direction for future work\.
#### 4\.2\.3Comparison with Baseline Methods \(RQ3\)
To validate thatMONBMsolves the proposed problem well and is competitive with existing methods, we compared all the non\-dominated solutions obtained byMONBM\-CDandMONBM\-CLinstantiations with their respective baseline methods on the corresponding optimization objectives\. Due to the absence of baseline methods for global interpretability, no comparisons are made for theMONBM\-CGinstantiation\. Furthermore, to comprehensively evaluate the framework’s capability in addressing complex high\-dimensional trade\-offs, we present a synthetic evaluation of theMONBM\-CDLGinstantiation against all baseline methods across all four objectives in Section S\-XIII of theSupplementary Material\. This joint analysis, supplemented by visualizations of the complete Pareto fronts for theMONBM\-CGandMONBM\-CDLGinstantiations in Figs\. S\-16 and S\-17 of theSupplementary Material, indicates that our framework produces non\-dominated solution sets with strong diversity and competitive performance\.
##### MONBM\-CD
We compared the performance ofMONBM\-CDand baseline methods on the test set𝓓test\\bm\{\\mathcal\{D\}\}\_\{\\textnormal\{\{test\}\}\}in terms ofACCandDPmetrics, specifically including: \(1\) comparing performance onHV,MS, andPDmetrics withFairEMOL\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]; \(2\) comparing the dominance relationships with all baseline methods; and \(3\) visualizing the performance of the models obtained by each method\.
We first compared the performance ofMONBM\-CDandFairEMOL\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]on theHV,MS, andPDmetrics, and the results are shown in Table[6](https://arxiv.org/html/2609.05946#S4.T6)\. It is obvious from the table that our method has better convergence and diversity\. More importantly, the models obtained by our method are interpretable\.
Table 6:The average of 15 trials ofHV,MS, andPDvalues for the final model set\. “\+/≈\\approx/\-” indicates that the averageFairEMOLmetric value is better/approximate/worse than the correspondingMONBMvalue on the Wilcoxon rank\-sum test, with a significance level of 0\.05\. The best average value is marked in gray\.Moreover, we compared the dominance relationships ofMONBM\-CDwith each baseline method in terms ofACCandDPmetrics, as shown in Table[7](https://arxiv.org/html/2609.05946#S4.T7)\. We also compared with the baseline methods in all four objectives, and the results are shown in Section S\-XIII of theSupplementary Material\. It is worth noting that Table[7](https://arxiv.org/html/2609.05946#S4.T7)comparesMONBM\-CDwith single\-solution\-based \(i\.e\.,Fair\-,LFR\-,Re\-,ROC\-,LPP\-,Oracle\-, andFRAPPÉ\-NBM\) and population\-based baseline methods \(i\.e\.,FairEMOL\) using different dominant metrics\. These follow the approaches in\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]and\[[47](https://arxiv.org/html/2609.05946#bib.bib1)\]for set\-to\-point and set\-to\-set dominance evaluation, respectively\. Details are given in Section S\-VI of theSupplementary Material\. However, they have similar meanings\. “Do\.”, “Non\-Do\.”, and “Be\-Do\.” being larger means that our method has a higher probability of dominating, not dominating each other, and being dominated by the corresponding method, respectively\.
Table 7:Dominance relationship betweenMONBM\-CDand other methods in terms ofACCandDPmetrics\. Results are averaged over 15 trials\. “Do\.”, “Non\-Do\.”, and “Be\-Do\.” mean the probability thatMONBM\-CDdominates, does not dominate each other, and is dominated by the corresponding method\. Dominance metrics for single\-point and population\-based baseline methods follow\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]and\[[47](https://arxiv.org/html/2609.05946#bib.bib1)\], respectively\.From Table[7](https://arxiv.org/html/2609.05946#S4.T7), we observe that \(1\)MONBM\-CDhas quite a few solutions that can dominate other methods, achieving better performance in terms of both accuracy and fairness; \(2\) it also yields a wide range of mutually non\-dominating solutions, serving as a valuable complement to existing methods; \(3\) apart fromLPP\-NBM, the probability ofMONBM\-CDbeing dominated by other methods remains very low\. AlthoughLPP\-NBMmay dominate some solutions generated byMONBM\-CD, our method shows a higher probability of producing solutions that dominate those ofLPP\-NBM, further demonstrating its competitive advantage\.
Finally, we visualized the performance of the models obtained by each method, as shown in Fig\.[11](https://arxiv.org/html/2609.05946#S4.F11)\. Notably, the performance ofROC\-NBMon theBankandDefaultdatasets is too poor to be displayed\. Fig\.[11](https://arxiv.org/html/2609.05946#S4.F11)reveals thatFair\-NBMhas limited improvement in fairness, whileROC\-NBMhas obvious compromises in accuracy\.LFR\-NBMandRe\-NBMcan achieve a compromise solution but lack competitiveness\.LPP\-NBM,Oracle\-NBM, andFRAPPÉ\-NBMcan obtain a single solution with both good accuracy and fairness\. In contrast,MONBM\-CDandFairEMOLobtain a set of models with different trade\-offs\. Moreover, our method is interpretable and much closer to the bottom right corner of the plot, indicating better accuracy and fairness\. This demonstrates the advantages of our approach\.
Figure 11:Comparison of the performance ofMONBM\-CDand baseline methods on model accuracy and fairness metrics on one arbitrary trial\.
##### MONBM\-CL
Similarly, we compared the performance ofMONBM\-CLand baseline methods on the test set𝓓test\\bm\{\\mathcal\{D\}\}\_\{\\textnormal\{\{test\}\}\}in terms ofACCandLImetrics, specifically including: \(1\) comparing the dominance relationship with all baseline methods and \(2\) visualizing the performance of the models obtained by each method\. We also compared the mutual dominance relationships with the baseline methods for all four objectives in Section S\-XIII of theSupplementary Material\.
ForSNAMandGAMI\-Net, we controlled the degree of feature selection by adjusting their hyper\-parameters to generate ten models with different degrees ofLI\. The dominance relationships betweenMONBM\-CLand the baseline methods in terms ofACCandLImetrics are shown in Table[8](https://arxiv.org/html/2609.05946#S4.T8)\. From Table[8](https://arxiv.org/html/2609.05946#S4.T8), it can be seen thatMONBM\-CLsignificantly outperformsSNAMandNPLAM\. However, the comparison withGAMI\-Netvaries across datasets:MONBM\-CLoutperformsGAMI\-Neton theAdult,Bank,COMPAS, andDutchdatasets while underperforming on theDefaultandLSATdatasets\. Overall,MONBM\-CLis more advantageous\. Fig\.[12](https://arxiv.org/html/2609.05946#S4.F12)further corroborates the above conclusion\.
Table 8:Dominance relationship betweenMONBM\-CLand other methods in terms ofACCandLImetrics\. Results are averaged over 15 trials\. “Do\.”, “Non\-Do\.”, and “Be\-Do\.” mean the probability thatMONBM\-CLdominates, does not dominate each other, and is dominated by the corresponding method\. Dominance metrics for all baseline methods follow\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]\.Figure 12:Comparison of the performance ofMONBM\-CLand baseline methods on model accuracy andLImetrics on one arbitrary trial\.More importantly, since the three baseline methods are not specifically designed to improveLI,MONBM\-CLpresents two major additional advantages: \(1\) The baseline methods exhibit high stochasticity, making it challenging to controlLIperformance by adjusting the degree of feature selection\. For example, stronger feature selection may instead result in a model that performs worse onLI\. In contrast,MONBM\-CLgenerates a set of models with different trade\-offs between accuracy andLIin a single run to choose from; \(2\) Compared to our method, the three baseline methods have difficulty in obtaining solutions with lowLI, which can also be seen in Fig\.[12](https://arxiv.org/html/2609.05946#S4.F12)\.
Overall, the comparison with existing methods validates the effectiveness of our method in producing models with competitive performance across accuracy, fairness, and interpretability\. Beyond numerical performance, our method also provides a diverse set of solutions\. As illustrated in Figs\.[11](https://arxiv.org/html/2609.05946#S4.F11)and[12](https://arxiv.org/html/2609.05946#S4.F12), and further demonstrated by the Pareto front visualizations for theMONBM\-CGandMONBM\-CDLGinstantiations in Figs\. S\-16 and S\-17 of theSupplementary Material, our framework exhibits excellent diversity in the objective space\. This enables stakeholders to select models that offer distinct and controllable trade\-offs across all optimized dimensions based on application\-specific priorities—an advantage that single\-point baselines inherently lack\.
### 4\.3Case Study: Threshold\-Based Model Selection
To illustrate how stakeholders may identify a deployable solution from a broad Pareto front with diverse trade\-offs, we present a case study based on a single run of theAdultdataset using theMONBM\-CDinstantiation\. Starting from the initial population of 100 solutions generated byMONBM, we first extracted the non\-dominated Pareto front, resulting in 43 candidate solutions\. We then further selected knee points from this Pareto set\. Knee points correspond to regions with relatively strong local trade\-off characteristics and, in the absence of explicit preferences, are commonly regarded as representative compromise solutions within their neighborhoods\[[59](https://arxiv.org/html/2609.05946#bib.bib46)\]\. Following this procedure, 13 knee\-point solutions were retained as the candidate set for subsequent decision making\.
Next, to simulate a realistic deployment scenario, we assume that a stakeholder specifies two threshold constraints according to operational requirements: \(i\) prediction accuracy must remain within 2% of the highest accuracy achieved among the candidate solutions; and \(ii\) theDPmetric must remain below 0\.05 to satisfy fairness requirements\. Applying these constraints progressively filters the candidate set and ultimately retains two feasible deployment candidates:
- •Model 0:ACC=86\.21%\\text\{ACC\}=86\.21\\%,DP=0\.0433\\text\{DP\}=0\.0433
- •Model 1:ACC=85\.13%\\text\{ACC\}=85\.13\\%,DP=0\.0302\\text\{DP\}=0\.0302
The final deployment decision can then be made according to application\-specific priorities\. For example, a stakeholder emphasizing predictive performance may preferModel 0, whereas scenarios with stricter fairness requirements may considerModel 1more appropriate\. Importantly, this case study is intended only to illustrate a simple a posteriori decision\-making process rather than prescribe a universal selection strategy\.
Overall, this example shows that the solution set generated byMONBMcan be progressively narrowed through the screening of knee points and intuitive threshold constraints, allowing stakeholders to efficiently move from a large Pareto set to a small set of practically deployable candidates while preserving flexibility in final decision making\.
## 5Discussion
This section aims to clarify the scope of the application and the limitations of the conclusions obtained in this paper, especially those exploring the relationships between dimensions\. In terms of interpretability, this paper employs NN\-based GAM as a self\-interpretable model to explore the relationship between its interpretability and other dimensions\. The corresponding conclusions can be generalized to other GAMs, which also need to consider the interpretability of global and local explanations\[[60](https://arxiv.org/html/2609.05946#bib.bib24),[36](https://arxiv.org/html/2609.05946#bib.bib25)\]\. Further, widely recognized self\-interpretable models include traditional linear models, decision trees, and decision rules\. The interpretability of the explanations they provide depends on the number of terms in the linear model, the depth of the decision tree, and the number of decision rules, respectively\[[4](https://arxiv.org/html/2609.05946#bib.bib5),[30](https://arxiv.org/html/2609.05946#bib.bib4)\]\. This is very similar to the local interpretability of NN\-based GAMs considered in this paper\. Therefore, the conclusions considering the relationship between interpretability and other dimensions in these models can be referred to the corresponding conclusions of this paper on local interpretability\.
Regarding fairness, this paper utilizes the widely used group fairness metricDPto assess the fairness of the models\. However, researchers have proposed numerous metrics to assess the fairness of ML models, which portray different aspects of fairness\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]\. Furthermore, existing studies have categorized these metrics, grouping highly correlated metrics into the same category\[[2](https://arxiv.org/html/2609.05946#bib.bib13),[58](https://arxiv.org/html/2609.05946#bib.bib7)\]\. Therefore, we believe that the corresponding conclusions obtained can be extended to fairness metrics within the same category, including two other widely used fairness metrics, equal opportunity \(EOP\) andEOD, as well as metrics like predictive equality \(PE\), discovery difference \(DD\)\. To empirically support this generalizability, we conducted a sensitivity analysis onEODin Section S\-VIII of theSupplementary Material\. The results indicate that the observed trade\-off patterns and decision\-logic adjustments underEODare largely consistent with those obtained usingDP\. It should be acknowledged that the conclusions presented here require further validation when applied to other categories of fairness metrics\. Nonetheless, they apply to three widely used fairness metrics:DP,EOP, andEOD\. In the future, we plan to consider accuracy, interpretability, and multiple fairness metrics simultaneously to analyze their relationships\.
Furthermore, we clarify that the current instantiation of theMONBMframework is designed specifically for tabular datasets, where each feature has a distinct and relatively independent semantic meaning\. In contrast, for high\-dimensional structured data such as images, the interpretation of univariate shape functions becomes less straightforward because individual input units \(e\.g\., pixels\) typically do not carry meaningful semantics in isolation and are strongly influenced by spatial redundancy and translation invariance \(e\.g\., the same visual pattern appearing at different positions within a region of interest\)\. As a result, the currentMONBMinstantiation—and more broadly, standard GAM formulations operating directly on input features—is not directly applicable to raw pixel\-level representations, since interpreting isolated pixel values provides limited explanatory value without surrounding context\. Nevertheless, this limitation does not prevent its extension to structured modalities\. As demonstrated in the original NBM framework\[[42](https://arxiv.org/html/2609.05946#bib.bib11)\], structured inputs can first be transformed into high\-level semantic concepts through a concept extraction stage, allowing additive basis functions to operate on human\-interpretable representations rather than raw inputs\. SinceMONBMfollows a model\-agnostic multi\-objective optimization paradigm, such concept\-level additive architectures may provide a feasible pathway for extending the framework beyond tabular data\. In this setting, the optimization objectives could remain unchanged while the underlying interpretable representation is adapted to the target modality\. Exploring and empirically validating such extensions on large\-scale structured datasets remains an important direction for future work\.
Beyond these scope considerations, we also examined the phenomenon of overfitting\. In our experiments forRQ2, the selected solutions showed consistent results between their performance on the training set \(Tables S\-V\-S\-VIII in theSupplementary Material\) and on the test set \(Tables[2](https://arxiv.org/html/2609.05946#S4.T2)–[S\-VIII](https://arxiv.org/html/2609.05946#S7.T8)\), indicating that overfitting was limited\. Nonetheless, non\-dominance inconsistencies may still occur\. In future work, we plan to use a validation set to mitigate this risk further\.
## 6Conclusion
In this paper, we introduced two explicit quantitative interpretability metrics, smoothness and monotonicity, specifically designed for NN\-based GAMs and experimentally confirmed that they guide the model toward desirable global explanations\. We then formulated the construction of accurate, fair, and interpretable GAMs as a new multi\-objective learning problem and proposed theMONBMframework, the first to bring multi\-objective evolutionary learning to GAMs while jointly optimizing accuracy, fairness, and both local and global interpretability\. Our instantiation retrains only the “feature specific” layer of a pre\-trained NBM, making EMO practical for large\-scale networks\.
UsingMONBM, we reveal the complex relationships among the dimensions and the reasons behind them\. Among them, improvements in fairness, local interpretability, and global interpretability each result in trade\-offs with accuracy, though to varying extents\. Moreover, the relationships between fairness, local interpretability, and global interpretability are complex, sometimes complementary, and sometimes conflicting, depending on the particular dataset and problem characteristics\. These findings also illustrate how EMO can be applied to self\-interpretable models to explore and interpret the complex relationships among multiple objectives, thereby offering a new perspective to analyze the trade\-offs between various trustworthiness objectives\. Finally, the effectiveness of the proposed metrics and the competitiveness ofMONBMare verified by visualizing the explanation results and comparing them with baseline methods\.
In the future, we plan to \(1\) systematically investigate non\-uniform smoothness control by integrating localized expert knowledge and adaptive weighting strategies, \(2\) validate the drawn conclusions on additional self\-interpretable models and different categories of fairness metrics, \(3\) conduct further empirical studies to evaluate the framework’s scalability and performance on high\-dimensional problems as more suitable datasets with sensitive attributes become available, and \(4\) explore potential trade\-off relationships with other trustworthy AI dimensions, such as robustness and security\.
## Data availability
We used publicly available/open\-source datasets for our experiments\. The details and access information for these datasets are provided in the manuscript\.
## CRediT authorship contribution statement
Ziming Wang:Writing – original draft, Methodology, Data curation, Validation;Changwu Huang:Conceptualization, Writing – review & editing, Supervision, Funding acquisition;Ke Tang:Writing – review & editing, Supervision, Funding acquisition;Yew\-Soon Ong:Writing – review & editing, Supervision, Funding acquisition;Xin Yao:Project administration, Resources, Writing – review & editing, Supervision, Funding acquisition;
## Declaration of competing interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper\.
## Acknowledgements
This work was supported by the National Natural Science Foundation of China \(Grant No\. 62250710682\), the Guangdong Major Project of Basic and Applied Basic Research \(Grant No\. 2023B0303000010\), the Guangdong Provincial Key Laboratory \(Grant No\. 2020B121201001\), BNBU Research Grant with No\. of R0700158\-26 at Beijing Normal\-Hong Kong Baptist University, internal grants of Lingnan University, and the IEEE Computational Intelligence Society Graduate Student Research Grant 2024\.
## References
- \[1\]R\. Agarwal, L\. Melnick, N\. Frosst, X\. Zhang, B\. Lengerich, R\. Caruana, and G\. E\. Hinton\(2021\)Neural additive models: interpretable machine learning with neural nets\.Advances in Neural Information Processing Systems34,pp\. 4699–4711\.Cited by:[§1](https://arxiv.org/html/2609.05946#S1.p2.1),[§1](https://arxiv.org/html/2609.05946#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.05946#S2.SS1.p4.1),[§2\.1](https://arxiv.org/html/2609.05946#S2.SS1.p5.1),[§2\.2\.2](https://arxiv.org/html/2609.05946#S2.SS2.SSS2.p2.1),[§2\.4](https://arxiv.org/html/2609.05946#S2.SS4.p2.1),[Table 1](https://arxiv.org/html/2609.05946#S2.T1.4.1.5.1)\.
- \[2\]H\. Anahideh, N\. Nezami, and A\. Asudeh\(2021\)On the choice of fairness: finding representative fairness metrics for a given context\.arXiv preprint arXiv:2109\.05697,pp\. 1–25\.Cited by:[§1](https://arxiv.org/html/2609.05946#S1.p4.1),[§2\.3](https://arxiv.org/html/2609.05946#S2.SS3.p1.1),[item \(2\)](https://arxiv.org/html/2609.05946#S3.I2.ix2.p1.1),[§5](https://arxiv.org/html/2609.05946#S5.p2.1)\.
- \[3\]J\. Angwin, J\. Larson, S\. Mattu, and L\. Kirchner\(2016\)Machine bias\.InEthics of Data and Analytics,pp\. 254–264\.Cited by:[§4\.1\.1](https://arxiv.org/html/2609.05946#S4.SS1.SSS1.p1.1)\.
- \[4\]A\. B\. Arrieta, N\. Díaz\-Rodríguez, J\. Del Ser, A\. Bennetot, S\. Tabik, A\. Barbado, S\. García, S\. Gil\-López, D\. Molina, R\. Benjamins,et al\.\(2020\)Explainable artificial intelligence \(XAI\): concepts, taxonomies, opportunities and challenges toward responsible AI\.Information Fusion58,pp\. 82–115\.Cited by:[§5](https://arxiv.org/html/2609.05946#S5.p1.1)\.
- \[5\]P\. L\. Bartlett and S\. Mendelson\(2002\)Rademacher and gaussian complexities: risk bounds and structural results\.Journal of Machine Learning Research3\(Nov\),pp\. 463–482\.Cited by:[§S\-II\.3](https://arxiv.org/html/2609.05946#S2.SS3a.p2.1)\.
- \[6\]R\. K\. Bellamy, K\. Dey, M\. Hind, S\. C\. Hoffman, S\. Houde, K\. Kannan, P\. Lohia, J\. Martino, S\. Mehta, A\. Mojsilovic,et al\.\(2018\)AI fairness 360: an extensible toolkit for detecting, understanding, and mitigating unwanted algorithmic bias\.arXiv preprint arXiv:1810\.01943\.Cited by:[§S\-III\.2](https://arxiv.org/html/2609.05946#S3.SS2a.p1.1)\.
- \[7\]U\. Bhatt, A\. Weller, and J\. M\. F\. Moura\(2020\)Evaluating and aggregating feature\-based model explanations\.InProceedings of the Twenty\-Ninth International Joint Conference on Artificial Intelligence, IJCAI\-20,pp\. 3016–3022\.Cited by:[§2\.2\.1](https://arxiv.org/html/2609.05946#S2.SS2.SSS1.p1.1),[§2\.2\.1](https://arxiv.org/html/2609.05946#S2.SS2.SSS1.p2.1),[item \(3\)](https://arxiv.org/html/2609.05946#S3.I2.ix3.p1.1)\.
- \[8\]T\. Calders and S\. Verwer\(2010\)Three naive bayes approaches for discrimination\-free classification\.Data Mining and Knowledge Discovery21,pp\. 277–292\.Cited by:[§2\.3](https://arxiv.org/html/2609.05946#S2.SS3.p1.1),[item \(2\)](https://arxiv.org/html/2609.05946#S3.I2.ix2.p1.1)\.
- \[9\]P\. Chalasani, J\. Chen, A\. R\. Chowdhury, X\. Wu, and S\. Jha\(2020\)Concise explanations of neural networks using adversarial training\.InInternational Conference on Machine Learning,pp\. 1383–1391\.Cited by:[§2\.2\.1](https://arxiv.org/html/2609.05946#S2.SS2.SSS1.p1.1),[§2\.2\.1](https://arxiv.org/html/2609.05946#S2.SS2.SSS1.p2.1),[item \(3\)](https://arxiv.org/html/2609.05946#S3.I2.ix3.p1.1)\.
- \[10\]A\. Chandra and X\. Yao\(2006\)Ensemble learning using multi\-objective evolutionary algorithms\.Journal of Mathematical Modelling and Algorithms5\(4\),pp\. 417–445\.Cited by:[§3\.2](https://arxiv.org/html/2609.05946#S3.SS2.p2.1),[§3](https://arxiv.org/html/2609.05946#S3.p1.1)\.
- \[11\]C\. Chang, R\. Caruana, and A\. Goldenberg\(2021\)Node\-gam: neural generalized additive model for interpretable deep learning\.arXiv preprint arXiv:2106\.01613\.Cited by:[§2\.1](https://arxiv.org/html/2609.05946#S2.SS1.p6.1)\.
- \[12\]D\. Chen and W\. Ye\(2022\)Monotonic neural additive models: pursuing regulated machine learning models for credit scoring\.InProceedings of the Third ACM International Conference on AI in Finance,pp\. 70–78\.Cited by:[§2\.2\.2](https://arxiv.org/html/2609.05946#S2.SS2.SSS2.p3.1),[§2\.4](https://arxiv.org/html/2609.05946#S2.SS4.p2.1),[Table 1](https://arxiv.org/html/2609.05946#S2.T1.4.1.10.1)\.
- \[13\]G\. Derringer and R\. Suich\(1980\)Simultaneous optimization of several response variables\.Journal of Quality Technology12\(4\),pp\. 214–219\.Cited by:[item \(3\)](https://arxiv.org/html/2609.05946#S3.I2.ix3.p1.1)\.
- \[14\]D\. Dua and C\. Graff\(2017\)UCI machine learning repository\.University of California, Irvine, School of Information and Computer Sciences\.External Links:[Link](http://archive.ics.uci.edu/ml)Cited by:[§4\.1\.1](https://arxiv.org/html/2609.05946#S4.SS1.SSS1.p1.1)\.
- \[15\]C\. Dwork, M\. Hardt, T\. Pitassi, O\. Reingold, and R\. Zemel\(2012\)Fairness through awareness\.InProceedings of the 3rd Innovations in Theoretical Computer Science Conference,pp\. 214–226\.Cited by:[§2\.3](https://arxiv.org/html/2609.05946#S2.SS3.p1.1),[item \(2\)](https://arxiv.org/html/2609.05946#S3.I2.ix2.p1.1)\.
- \[16\]Z\. Gong, H\. Chen, B\. Yuan, and X\. Yao\(2018\)Multiobjective learning in the model space for time series classification\.IEEE Transactions on Cybernetics49\(3\),pp\. 918–932\.Cited by:[§3\.3\.3](https://arxiv.org/html/2609.05946#S3.SS3.SSS3.p1.1)\.
- \[17\]T\. J\. Hastie\(2017\)Generalized additive models\.InStatistical Models in S,pp\. 249–307\.Cited by:[§1](https://arxiv.org/html/2609.05946#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.05946#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2609.05946#S2.SS1.p4.1),[Table 1](https://arxiv.org/html/2609.05946#S2.T1.4.1.3.1)\.
- \[18\]C\. Huang, Z\. Zhang, B\. Mao, and X\. Yao\(2023\)An overview of artificial intelligence ethics\.IEEE Transactions on Artificial Intelligence4\(4\),pp\. 799–819\.Cited by:[§1](https://arxiv.org/html/2609.05946#S1.p1.1)\.
- \[19\]F\. Kamiran and T\. Calders\(2012\)Data preprocessing techniques for classification without discrimination\.Knowledge and Information Systems33\(1\),pp\. 1–33\.Cited by:[§S\-II\.1](https://arxiv.org/html/2609.05946#S2.SS1a.p1.1),[§2\.3](https://arxiv.org/html/2609.05946#S2.SS3.p3.1),[§S\-III\.2](https://arxiv.org/html/2609.05946#S3.SS2a.p1.1),[item \(1\)](https://arxiv.org/html/2609.05946#S4.I1.ix1.p1.1)\.
- \[20\]F\. Kamiran, A\. Karim, and X\. Zhang\(2012\)Decision theory for discrimination\-aware classification\.In2012 IEEE 12th International Conference on Data Mining,pp\. 924–929\.Cited by:[§S\-II\.1](https://arxiv.org/html/2609.05946#S2.SS1a.p1.1),[§2\.3](https://arxiv.org/html/2609.05946#S2.SS3.p3.1),[§S\-III\.2](https://arxiv.org/html/2609.05946#S3.SS2a.p1.1),[item \(1\)](https://arxiv.org/html/2609.05946#S4.I1.ix1.p1.1)\.
- \[21\]N\. Kemmerzell and A\. Schreiner\(2024\)Quantifying the trade\-offs between dimensions of trustworthy AI\-an empirical study on fairness, explainability, privacy, and robustness\.InGerman Conference on Artificial Intelligence \(Künstliche Intelligenz\),pp\. 128–146\.Cited by:[§1](https://arxiv.org/html/2609.05946#S1.p5.1)\.
- \[22\]M\. Kim, S\. Lee, and J\. Kim\(2026\)Transparent networks for multivariate time series\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 37510–37518\.Cited by:[§2\.1](https://arxiv.org/html/2609.05946#S2.SS1.p7.1)\.
- \[23\]M\. Kraus, D\. Tschernutter, S\. Weinzierl, and P\. Zschech\(2024\)Interpretable generalized additive neural networks\.European Journal of Operational Research317\(2\),pp\. 303–316\.Cited by:[item \(1\)](https://arxiv.org/html/2609.05946#S2.I1.ix1.p1.1),[item \(2\)](https://arxiv.org/html/2609.05946#S2.I1.ix2.p1.1),[§2\.1](https://arxiv.org/html/2609.05946#S2.SS1.p2.1),[§S\-II\.1](https://arxiv.org/html/2609.05946#S2.SS1a.p1.1),[§2\.3](https://arxiv.org/html/2609.05946#S2.SS3.p4.1),[§2\.4](https://arxiv.org/html/2609.05946#S2.SS4.p2.1),[Table 1](https://arxiv.org/html/2609.05946#S2.T1.4.1.11.1),[item \(1\)](https://arxiv.org/html/2609.05946#S4.I1.ix1.p1.1)\.
- \[24\]A\. Krogh and J\. Hertz\(1991\)A simple weight decay can improve generalization\.Advances in Neural Information Processing Systems4\.Cited by:[§2\.2\.2](https://arxiv.org/html/2609.05946#S2.SS2.SSS2.p2.1)\.
- \[25\]T\. Le Quy, A\. Roy, V\. Iosifidis, W\. Zhang, and E\. Ntoutsi\(2022\)A survey on datasets for fairness\-aware machine learning\.Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery12\(3\),pp\. e1452\.Cited by:[§S\-I](https://arxiv.org/html/2609.05946#S1a.p1.1),[§4\.1\.1](https://arxiv.org/html/2609.05946#S4.SS1.SSS1.p1.1)\.
- \[26\]B\. Li, J\. Li, K\. Tang, and X\. Yao\(2015\)Many\-objective evolutionary algorithms: a survey\.ACM Comput\. Surv\.48\(1\),pp\. 1–35\.Cited by:[§3\.2](https://arxiv.org/html/2609.05946#S3.SS2.p4.1)\.
- \[27\]B\. Li, K\. Tang, J\. Li, and X\. Yao\(2016\)Stochastic ranking algorithm for many\-objective optimization based on multiple indicators\.IEEE Transactions on Evolutionary Computation20\(6\),pp\. 924–938\.Cited by:[1st item](https://arxiv.org/html/2609.05946#S2.I1.i1.p1.1),[§3\.3\.3](https://arxiv.org/html/2609.05946#S3.SS3.SSS3.p1.1)\.
- \[28\]J\. Li, B\. Mao, Z\. Liang, Z\. Zhang, Q\. Lin, and X\. Yao\(2021\)Trust and trustworthiness: what they are and how to achieve them\.In2021 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events \(PerCom Workshops\),pp\. 711–717\.Cited by:[§1](https://arxiv.org/html/2609.05946#S1.p1.1)\.
- \[29\]M\. Li and X\. Yao\(2019\)Quality evaluation of solution sets in multiobjective optimisation: a survey\.ACM Computing Surveys \(CSUR\)52\(2\),pp\. 1–38\.Cited by:[§4\.1\.4](https://arxiv.org/html/2609.05946#S4.SS1.SSS4.p1.1),[§4\.1\.4](https://arxiv.org/html/2609.05946#S4.SS1.SSS4.p2.1)\.
- \[30\]P\. Linardatos, V\. Papastefanopoulos, and S\. Kotsiantis\(2020\)Explainable AI: a review of machine learning interpretability methods\.Entropy23\(1\),pp\. 18\.Cited by:[§5](https://arxiv.org/html/2609.05946#S5.p1.1)\.
- \[31\]I\. Loshchilov and F\. Hutter\(2017\)SGDR: stochastic gradient descent with warm restarts\.InInternational Conference on Learning Representations,Cited by:[§S\-III\.1](https://arxiv.org/html/2609.05946#S3.SS1a.p1.1)\.
- \[32\]I\. Loshchilov and F\. Hutter\(2019\)Decoupled weight decay regularization\.InInternational Conference on Learning Representations,Cited by:[§S\-III\.1](https://arxiv.org/html/2609.05946#S3.SS1a.p1.1)\.
- \[33\]Y\. Lou, R\. Caruana, J\. Gehrke, and G\. Hooker\(2013\)Accurate intelligible models with pairwise interactions\.InProceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,pp\. 623–631\.Cited by:[§2\.1](https://arxiv.org/html/2609.05946#S2.SS1.p6.1)\.
- \[34\]Y\. Lou, R\. Caruana, and J\. Gehrke\(2012\)Intelligible models for classification and regression\.InProceedings of the 18th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,pp\. 150–158\.Cited by:[§S\-II\.3](https://arxiv.org/html/2609.05946#S2.SS3a.p1.1)\.
- \[35\]D\. J\. MacKay\(1992\)Bayesian interpolation\.Neural Computation4\(3\),pp\. 415–447\.Cited by:[§2\.2\.2](https://arxiv.org/html/2609.05946#S2.SS2.SSS2.p2.1)\.
- \[36\]J\. Maindonald\(2010\)Smoothing terms in GAM models\.Technical reportAustralian National University\.Note:Technical report/supplementary materialExternal Links:[Link](https://maths-people.anu.edu.au/~johnm/r-book/xtras/autosmooth.pdf)Cited by:[§2\.2\.2](https://arxiv.org/html/2609.05946#S2.SS2.SSS2.p1.1),[§2\.2\.2](https://arxiv.org/html/2609.05946#S2.SS2.SSS2.p2.1),[§5](https://arxiv.org/html/2609.05946#S5.p1.1)\.
- \[37\]N\. Mehrabi, F\. Morstatter, N\. Saxena, K\. Lerman, and A\. Galstyan\(2021\)A survey on bias and fairness in machine learning\.ACM Computing Surveys \(CSUR\)54\(6\),pp\. 1–35\.Cited by:[§1](https://arxiv.org/html/2609.05946#S1.p4.1),[§2\.3](https://arxiv.org/html/2609.05946#S2.SS3.p1.1)\.
- \[38\]L\. L\. Minku and X\. Yao\(2013\)Software effort estimation as a multiobjective learning problem\.ACM Transactions on Software Engineering and Methodology \(TOSEM\)22\(4\),pp\. 1–32\.Cited by:[§3\.3\.3](https://arxiv.org/html/2609.05946#S3.SS3.SSS3.p1.1)\.
- \[39\]W\. R\. Monteiro and G\. Reynoso\-Meza\(2022\)A review of the convergence between explainable artificial intelligence and multi\-objective optimization\.Authorea Preprints\.Cited by:[§1](https://arxiv.org/html/2609.05946#S1.p5.1)\.
- \[40\]H\. Nori, S\. Jenkins, P\. Koch, and R\. Caruana\(2019\)InterpretML: a unified framework for machine learning interpretability\.arXiv preprint arXiv:1909\.09223\.Cited by:[§1](https://arxiv.org/html/2609.05946#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.05946#S2.SS1.p4.1),[§2\.2\.2](https://arxiv.org/html/2609.05946#S2.SS2.SSS2.p2.1),[Table 1](https://arxiv.org/html/2609.05946#S2.T1.4.1.4.1)\.
- \[41\]D\. Pessach and E\. Shmueli\(2022\)A review on fairness in machine learning\.ACM Computing Surveys \(CSUR\)55\(3\),pp\. 1–44\.Cited by:[§1](https://arxiv.org/html/2609.05946#S1.p4.1)\.
- \[42\]F\. Radenovic, A\. Dubey, and D\. Mahajan\(2022\)Neural basis models for interpretability\.Advances in Neural Information Processing Systems35,pp\. 8414–8426\.Cited by:[§1](https://arxiv.org/html/2609.05946#S1.p2.1),[§1](https://arxiv.org/html/2609.05946#S1.p3.1),[Figure 2](https://arxiv.org/html/2609.05946#S2.F2),[§2\.1](https://arxiv.org/html/2609.05946#S2.SS1.p4.1),[§2\.1](https://arxiv.org/html/2609.05946#S2.SS1.p5.1),[§2\.1](https://arxiv.org/html/2609.05946#S2.SS1.p7.1),[§2\.2\.2](https://arxiv.org/html/2609.05946#S2.SS2.SSS2.p2.1),[§2\.4](https://arxiv.org/html/2609.05946#S2.SS4.p2.1),[Table 1](https://arxiv.org/html/2609.05946#S2.T1.4.1.6.1),[§3\.3\.1](https://arxiv.org/html/2609.05946#S3.SS3.SSS1.p1.1),[§4\.1\.3](https://arxiv.org/html/2609.05946#S4.SS1.SSS3.p1.1),[§5](https://arxiv.org/html/2609.05946#S5.p3.1)\.
- \[43\]T\. P\. Runarsson and X\. Yao\(2000\)Stochastic ranking for constrained evolutionary optimization\.IEEE Transactions on Evolutionary Computation4\(3\),pp\. 284–294\.Cited by:[§3\.3\.3](https://arxiv.org/html/2609.05946#S3.SS3.SSS3.p1.1)\.
- \[44\]A\. Solow, S\. Polasky, and J\. Broadus\(1993\)On the measurement of biological diversity\.Journal of Environmental Economics and Management24\(1\),pp\. 60–68\.Cited by:[§4\.1\.4](https://arxiv.org/html/2609.05946#S4.SS1.SSS4.p1.1)\.
- \[45\]A\. Ţifrea, P\. Lahoti, B\. Packer, Y\. Halpern, A\. Beirami, and F\. Prost\(2024\)FRAPPÉ: a group fairness framework for post\-processing everything\.InProceedings of the 41st International Conference on Machine Learning,pp\. 48321–48343\.Cited by:[§S\-III\.2](https://arxiv.org/html/2609.05946#S3.SS2a.p1.1),[item \(1\)](https://arxiv.org/html/2609.05946#S4.I1.ix1.p1.1)\.
- \[46\]P\. Van der Laan\(2001\)The 2001 census in the Netherlands: integration of registers and surveys\.InConference at the Cathie Marsh Centre,pp\. 1–24\.Cited by:[§4\.1\.1](https://arxiv.org/html/2609.05946#S4.SS1.SSS1.p1.1)\.
- \[47\]Z\. Wang, C\. Huang, Y\. Li, and X\. Yao\(2024\)Multi\-objective feature attribution explanation for explainable machine learning\.ACM Transactions on Evolutionary Learning and Optimization4\(1\),pp\. 1–32\.Cited by:[§1](https://arxiv.org/html/2609.05946#S1.p5.1),[§2\.2\.1](https://arxiv.org/html/2609.05946#S2.SS2.SSS1.p1.1),[§2\.2\.1](https://arxiv.org/html/2609.05946#S2.SS2.SSS1.p2.1),[§3](https://arxiv.org/html/2609.05946#S3.p1.1),[§4\.1\.4](https://arxiv.org/html/2609.05946#S4.SS1.SSS4.p3.1),[§4\.2\.3](https://arxiv.org/html/2609.05946#S4.SS2.SSS3.Px1.p3.1),[Table 7](https://arxiv.org/html/2609.05946#S4.T7),[§S\-VI](https://arxiv.org/html/2609.05946#S6a.p1.1),[§S\-VI](https://arxiv.org/html/2609.05946#S6a.p2.1),[§S\-VI](https://arxiv.org/html/2609.05946#S6a.p9.1)\.
- \[48\]Z\. Wang, C\. Huang, and X\. Yao\(2024\)A roadmap of explainable artificial intelligence: explain to whom, when, what and how?\.ACM Trans\. Auton\. Adapt\. Syst\.19\(4\),pp\. 1–40\.Cited by:[§1](https://arxiv.org/html/2609.05946#S1.p2.1)\.
- \[49\]L\. F\. Wightman\(1998\)LSAC national longitudinal bar passage study\. LSAC research report series\.\.Cited by:[§4\.1\.1](https://arxiv.org/html/2609.05946#S4.SS1.SSS1.p1.1)\.
- \[50\]R\. Xian, L\. Yin, and H\. Zhao\(2023\)Fair and optimal classification via post\-processing\.InInternational Conference on Machine Learning,pp\. 37977–38012\.Cited by:[§S\-II\.1](https://arxiv.org/html/2609.05946#S2.SS1a.p1.1),[§2\.3](https://arxiv.org/html/2609.05946#S2.SS3.p3.1),[§S\-III\.2](https://arxiv.org/html/2609.05946#S3.SS2a.p1.1),[item \(1\)](https://arxiv.org/html/2609.05946#S4.I1.ix1.p1.1)\.
- \[51\]R\. Xian and H\. Zhao\(2024\)A unified post\-processing framework for group fairness in classification\.arXiv preprint arXiv:2405\.04025\.Cited by:[§S\-II\.1](https://arxiv.org/html/2609.05946#S2.SS1a.p1.1),[§2\.3](https://arxiv.org/html/2609.05946#S2.SS3.p3.1),[§S\-III\.2](https://arxiv.org/html/2609.05946#S3.SS2a.p1.1),[item \(1\)](https://arxiv.org/html/2609.05946#S4.I1.ix1.p1.1)\.
- \[52\]S\. Xu, Z\. Bu, P\. Chaudhari, and I\. J\. Barnett\(2023\)Sparse neural additive model: interpretable deep learning with feature selection via group sparsity\.InJoint European Conference on Machine Learning and Knowledge Discovery in Databases,pp\. 343–359\.Cited by:[§2\.1](https://arxiv.org/html/2609.05946#S2.SS1.p6.1),[§S\-II\.1](https://arxiv.org/html/2609.05946#S2.SS1a.p2.1),[§2\.2\.1](https://arxiv.org/html/2609.05946#S2.SS2.SSS1.p2.1),[§2\.4](https://arxiv.org/html/2609.05946#S2.SS4.p2.1),[Table 1](https://arxiv.org/html/2609.05946#S2.T1.4.1.7.1),[§S\-III\.2](https://arxiv.org/html/2609.05946#S3.SS2a.p2.1),[item \(2\)](https://arxiv.org/html/2609.05946#S4.I1.ix2.p1.1)\.
- \[53\]Z\. Yang, A\. Zhang, and A\. Sudjianto\(2021\)GAMI\-Net: an explainable neural network based on generalized additive models with structured interactions\.Pattern Recognition120,pp\. 108192\.Cited by:[§2\.1](https://arxiv.org/html/2609.05946#S2.SS1.p6.1),[§S\-II\.1](https://arxiv.org/html/2609.05946#S2.SS1a.p2.1),[§2\.2\.1](https://arxiv.org/html/2609.05946#S2.SS2.SSS1.p2.1),[§2\.4](https://arxiv.org/html/2609.05946#S2.SS4.p2.1),[Table 1](https://arxiv.org/html/2609.05946#S2.T1.4.1.8.1),[§S\-III\.2](https://arxiv.org/html/2609.05946#S3.SS2a.p2.1),[item \(2\)](https://arxiv.org/html/2609.05946#S4.I1.ix2.p1.1)\.
- \[54\]X\. Yao and Y\. Liu\(1997\)A new evolutionary system for evolving artificial neural networks\.IEEE Transactions on Neural Networks8\(3\),pp\. 694–713\.Cited by:[§3\.2](https://arxiv.org/html/2609.05946#S3.SS2.p3.1),[§3\.3\.3](https://arxiv.org/html/2609.05946#S3.SS3.SSS3.p1.1),[7](https://arxiv.org/html/2609.05946#alg1.l7.3.1),[9](https://arxiv.org/html/2609.05946#alg2.l9.3.1)\.
- \[55\]X\. Yao\(1999\)Evolving artificial neural networks\.Proceedings of the IEEE87\(9\),pp\. 1423–1447\.Cited by:[§3\.2](https://arxiv.org/html/2609.05946#S3.SS2.p3.1),[§3\.3\.3](https://arxiv.org/html/2609.05946#S3.SS3.SSS3.p1.1),[7](https://arxiv.org/html/2609.05946#alg1.l7.3.1),[9](https://arxiv.org/html/2609.05946#alg2.l9.3.1)\.
- \[56\]R\. Zemel, Y\. Wu, K\. Swersky, T\. Pitassi, and C\. Dwork\(2013\)Learning fair representations\.InInternational Conference on Machine Learning,pp\. 325–333\.Cited by:[§S\-II\.1](https://arxiv.org/html/2609.05946#S2.SS1a.p1.1),[§2\.3](https://arxiv.org/html/2609.05946#S2.SS3.p3.1),[§S\-III\.2](https://arxiv.org/html/2609.05946#S3.SS2a.p1.1),[item \(1\)](https://arxiv.org/html/2609.05946#S4.I1.ix1.p1.1)\.
- \[57\]Q\. Zhang, J\. Liu, and X\. Yao\(2024\)Fairness\-aware multiobjective evolutionary learning\.IEEE Transactions on Evolutionary Computation\.Note:, doi: 10\.1109/TEVC\.2024\.3430824Cited by:[§1](https://arxiv.org/html/2609.05946#S1.p4.1),[§1](https://arxiv.org/html/2609.05946#S1.p5.1),[§4\.2\.2](https://arxiv.org/html/2609.05946#S4.SS2.SSS2.Px1.p5.1)\.
- \[58\]Q\. Zhang, J\. Liu, Z\. Zhang, J\. Wen, B\. Mao, and X\. Yao\(2022\)Mitigating unfairness via evolutionary multiobjective ensemble learning\.IEEE Transactions on Evolutionary Computation27\(4\),pp\. 848–862\.Cited by:[§1](https://arxiv.org/html/2609.05946#S1.p4.1),[§1](https://arxiv.org/html/2609.05946#S1.p5.1),[Table S\-XIII](https://arxiv.org/html/2609.05946#S13.T13),[Table S\-XIV](https://arxiv.org/html/2609.05946#S13.T14),[§S\-XIII](https://arxiv.org/html/2609.05946#S13.p1.1),[§S\-II\.1](https://arxiv.org/html/2609.05946#S2.SS1a.p1.1),[§2\.3](https://arxiv.org/html/2609.05946#S2.SS3.p3.1),[§S\-III\.2](https://arxiv.org/html/2609.05946#S3.SS2a.p1.1),[§3](https://arxiv.org/html/2609.05946#S3.p1.1),[item \(1\)](https://arxiv.org/html/2609.05946#S4.I1.ix1.p1.1),[§4\.1\.3](https://arxiv.org/html/2609.05946#S4.SS1.SSS3.p2.1),[§4\.1\.4](https://arxiv.org/html/2609.05946#S4.SS1.SSS4.p1.1),[§4\.1\.4](https://arxiv.org/html/2609.05946#S4.SS1.SSS4.p2.1),[§4\.1\.4](https://arxiv.org/html/2609.05946#S4.SS1.SSS4.p3.1),[§4\.2\.2](https://arxiv.org/html/2609.05946#S4.SS2.SSS2.Px1.p5.1),[§4\.2\.3](https://arxiv.org/html/2609.05946#S4.SS2.SSS3.Px1.p1.1),[§4\.2\.3](https://arxiv.org/html/2609.05946#S4.SS2.SSS3.Px1.p2.1),[§4\.2\.3](https://arxiv.org/html/2609.05946#S4.SS2.SSS3.Px1.p3.1),[Table 7](https://arxiv.org/html/2609.05946#S4.T7),[Table 8](https://arxiv.org/html/2609.05946#S4.T8),[§5](https://arxiv.org/html/2609.05946#S5.p2.1),[§S\-VI](https://arxiv.org/html/2609.05946#S6a.p1.1),[§S\-VI](https://arxiv.org/html/2609.05946#S6a.p3.1)\.
- \[59\]X\. Zhang, Y\. Tian, and Y\. Jin\(2014\)A knee point\-driven evolutionary algorithm for many\-objective optimization\.IEEE Transactions on Evolutionary Computation19\(6\),pp\. 761–776\.Cited by:[§4\.2\.1](https://arxiv.org/html/2609.05946#S4.SS2.SSS1.p1.1),[§4\.2\.1](https://arxiv.org/html/2609.05946#S4.SS2.SSS1.p2.1),[§4\.3](https://arxiv.org/html/2609.05946#S4.SS3.p1.1)\.
- \[60\]C\. Zhong, Z\. Chen, J\. Liu, M\. Seltzer, and C\. Rudin\(2024\)Exploring and interacting with the set of good sparse generalized additive models\.Advances in Neural Information Processing Systems36\.Cited by:[§2\.2\.2](https://arxiv.org/html/2609.05946#S2.SS2.SSS2.p1.1),[§2\.2\.2](https://arxiv.org/html/2609.05946#S2.SS2.SSS2.p3.1),[§5](https://arxiv.org/html/2609.05946#S5.p1.1)\.
- \[61\]R\. Zhou, J\. Bacardit, A\. E\. Brownlee, S\. Cagnoni, M\. Fyvie, G\. Iacca, J\. McCall, N\. van Stein, D\. J\. Walker, and T\. Hu\(2024\)Evolutionary computation and explainable AI: a roadmap to understandable intelligent systems\.IEEE Transactions on Evolutionary Computation\.Note:, doi: 10\.1109/TEVC\.2024\.3476443Cited by:[§1](https://arxiv.org/html/2609.05946#S1.p5.1)\.
- \[62\]L\. Zhu, H\. Li, X\. Zhang, L\. Wu, and H\. Chen\(2024\)Neural partially linear additive model\.Frontiers of Computer Science18\(6\),pp\. 186334\.Cited by:[§2\.1](https://arxiv.org/html/2609.05946#S2.SS1.p6.1),[§2\.2\.1](https://arxiv.org/html/2609.05946#S2.SS2.SSS1.p2.1),[§2\.4](https://arxiv.org/html/2609.05946#S2.SS4.p2.1),[Table 1](https://arxiv.org/html/2609.05946#S2.T1.4.1.9.1),[item \(2\)](https://arxiv.org/html/2609.05946#S4.I1.ix2.p1.1)\.
- \[63\]E\. Zitzler, K\. Deb, and L\. Thiele\(2000\)Comparison of multiobjective evolutionary algorithms: empirical results\.Evolutionary Computation8\(2\),pp\. 173–195\.Cited by:[§4\.1\.4](https://arxiv.org/html/2609.05946#S4.SS1.SSS4.p1.1)\.
- \[64\]E\. Zitzler and L\. Thiele\(1998\)Multiobjective optimization using evolutionary algorithms—a comparative case study\.InInternational Conference on Parallel Problem Solving from Nature,pp\. 292–301\.Cited by:[§4\.1\.4](https://arxiv.org/html/2609.05946#S4.SS1.SSS4.p1.1)\.
## S\-IDataset Under Consideration
In this paper, we consider six datasets that are widely used in the field of fairness\[[25](https://arxiv.org/html/2609.05946#bib.bib39)\], with specific information shown in Table[S\-I](https://arxiv.org/html/2609.05946#S1.T1)\.
Table S\-I:Six datasets used in this paper\. “\|𝓓\|\|\\bm\{\\mathcal\{D\}\}\|”, “\|X\|\|\\textit\{\{X\}\}\|”, “\|Xdisc\.\|\|\\textit\{\{X\}\}\_\{disc\.\}\|”, “\|Xnum\.\|\|\\textit\{\{X\}\}\_\{num\.\}\|”, and “SS” denote the number of samples, the number of features, the number of discrete features, the number of numerical features, and the sensitive attribute under consideration, respectively\.
## S\-IIAlgorithmic and Computational Properties ofMONBM
This section provides a comprehensive analysis of the proposedMONBMframework, focusing on its computational budget, algorithmic complexity, generalization behavior, robustness, and practical execution time\. These analyses aim to clarify the computational characteristics and practical scalability of the proposed framework\.
### S\-II\.1Computational Budget Analysis
For the baseline methods compared withMONBM\-CD, all single\-point methods \(e\.g\.,Fair\-NBM\[[23](https://arxiv.org/html/2609.05946#bib.bib20)\],LFR\-NBM\[[56](https://arxiv.org/html/2609.05946#bib.bib8)\],Re\-NBM\[[19](https://arxiv.org/html/2609.05946#bib.bib9)\],ROC\-NBM\[[20](https://arxiv.org/html/2609.05946#bib.bib19)\],LPP\-NBM\[[50](https://arxiv.org/html/2609.05946#bib.bib53)\],Oracle\-NBM\[[51](https://arxiv.org/html/2609.05946#bib.bib54)\], andFRAPPÉ\-NBM\) were trained for 1,200 epochs\.FairEMOL\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\], a population\-based multi\-objective method, was trained for 200 generations on a deep neural network with approximately 68,000 parameters\.
For the baseline methods compared withMONBM\-CL,SNAM\[[52](https://arxiv.org/html/2609.05946#bib.bib16)\]andGAMI\-Net\[[53](https://arxiv.org/html/2609.05946#bib.bib15)\]were trained for 1,200 and 5,000 epochs, respectively\.
In contrast, our method appliesMONBMto retrain only the “feature specific” layer of a pre\-trained NBM model, which contains approximately 1,500 parameters\. We conduct 200 generations of evolutionary optimization using this efficient retraining strategy on the pre\-trained NBM for 1,000 epochs\.
Overall, although each generation inMONBMinvolves updating a population of 100 models, the per\-model computational cost is relatively low, as only∼\\sim2% of the NBM’s total parameters are updated\. Additionally, unlike single\-point methods that yield only a single solution per run,MONBMproduces a diverse set of trade\-off solutions in a single run, supporting downstream analysis and flexible deployment\.
### S\-II\.2Algorithmic Complexity
For a multi\-objective problem withkkobjectives, a population size ofτ\\tau, a data dimension ofDD, and a data size ofnn, the computational complexity analysis of theMONBMinstantiation \(Algorithm 2\) is as follows:
- •Time Complexity:Each generation involves reproduction \(O\(τ⋅D\)O\(\\tau\\cdot D\)\), partial training forNNindividuals \(O\(k⋅N⋅n⋅D\)O\(k\\cdot N\\cdot n\\cdot D\)\), and indicator\-based fitness evaluation\. Specifically, accuracy and fairness evaluations takeO\(n⋅D\)O\(n\\cdot D\),GItakesO\(D\)O\(D\), whileLIrequiresO\(n⋅DlogD\)O\(n\\cdot D\\log D\)due to feature ranking\. The SRA\-based environmental selection takesO\(k⋅τ2\)O\(k\\cdot\\tau^\{2\}\)\[[27](https://arxiv.org/html/2609.05946#bib.bib31)\]\. Combining these components, the total time complexity per generation isO\(kτ2\+τ⋅n⋅DlogD\)O\(k\\tau^\{2\}\+\\tau\\cdot n\\cdot D\\log D\), indicating quadratic dependence on population size and near\-linear scaling with dataset size and dimensionality\.
- •Space Complexity:The framework maintains the population parametersO\(τ⋅D\)O\(\\tau\\cdot D\)and the training datasetO\(n⋅D\)O\(n\\cdot D\)\. Since the shared bases are frozen and shared, the overall space complexity isO\(\(τ\+n\)D\)O\(\(\\tau\+n\)D\), which scales linearly with the feature dimensionDD\.
### S\-II\.3Robustness and Generalization Analysis
The generalization behavior ofMONBMcan be understood from both its structural design and its explicit functional regularization mechanisms\. Fundamentally, as a class of GAMs,MONBMtends to exhibit stable generalization by limiting high\-order feature interactions\. As established in the bias–variance analysis of additive models\[[34](https://arxiv.org/html/2609.05946#bib.bib62)\], reducing complex feature couplings can lower model variance, which helps mitigate overfitting commonly observed in highly interactive models and supports reliable performance across diverse practical scenarios\.
From the perspective of statistical learning theory\[[5](https://arxiv.org/html/2609.05946#bib.bib58)\], generalization performance is closely related to the capacity of the hypothesis space\. By restricting evolutionary optimization to approximately 2% of the total model parameters—specifically the linear combination weights over frozen nonlinear basis functions—the effective hypothesis space ofMONBMis substantially smaller than that of fully trainable deep models\. This controlled parameterization helps regulate model capacity and reduces the risk of overfitting while preserving sufficient expressive power\.
Furthermore, the explicit smoothness objective \(ℳs\\mathcal\{M\}\_\{s\}\) acts as a functional regularization mechanism\. By discouraging excessive oscillations in the learned shape functions, it helps prevent fitting high\-frequency noise and encourages stable global trends\. Empirically, this behavior is reflected in the consistent performance observed across 15 independent trials on both training and test sets \(Tables 2 to 5 and Tables[S\-V](https://arxiv.org/html/2609.05946#S7.T5)to[S\-VIII](https://arxiv.org/html/2609.05946#S7.T8)\)\. The relatively small performance differences between training and testing results suggest stable generalization performance when optimizing multiple objectives simultaneously\.
### S\-II\.4Wall\-Clock Time Performance
To provide a precise assessment of the practical scalability and efficiency of the framework, we measured the average wall\-clock time for theMONBM\-CDinstantiation across all six benchmark datasets\. All experiments were conducted on a Linux server equipped with dual Intel Xeon Platinum 8358 processors \(totaling 64 physical cores\) and 512GB of RAM\. To ensure a fair and consistent comparison, all computational tasks were restricted to a single\-threaded execution mode\. Specifically, we measured the total runtime required for a fullMONBMexecution \(evolving a population ofτ=100\\tau=100models for 200 generations\) and compared it with the time required to train a single standardNBMmodel until convergence\.
Table S\-II:Comparison of average wall\-clock time \(seconds\) between a singleNBMtraining run and theMONBM\-CDinstantiation across six datasets\.DatasetNBMMONBMAdult1,681\.8919,650\.14Bank1,531\.1719,273\.12COMPAS203\.4735,359\.28Default1,515\.1620,792\.25Dutch2,087\.1022,009\.63LSAT863\.61310,367\.60Mean1,313\.7316,242\.00As summarized in Table[S\-II](https://arxiv.org/html/2609.05946#S2.T2), the total execution time ofMONBMis on average approximately 12\.4 times that of a single NBM training run\. However, unlike single\-model training,MONBMproduces a diverse set of trade\-off solutions in one execution\. Obtaining a comparable approximation of the Pareto front using scalarized optimization would typically require many independent runs with different weight settings, leading to substantially higher cumulative cost\. These results indicate that the partial\-retraining strategy effectively reduces the computational burden of population\-based optimization, makingMONBMpractically feasible for exploring complex trade\-offs among accuracy, fairness, and interpretability\.
## S\-IIIHyperparameters of Pre\-trained NBM and Baseline Methods
### S\-III\.1Hyperparameters of Pre\-trained NBM
For training the pre\-trained NBM, we used the Adam with decoupled weight decay \(AdamW\) optimizer\[[32](https://arxiv.org/html/2609.05946#bib.bib40)\]and trained for 1000 epochs with a batch size of 1024\. The learning rate was decayed from its initial value to zero using cosine annealing\[[31](https://arxiv.org/html/2609.05946#bib.bib41)\]\. The shared MLP architecture included three hidden layers with 256, 128, and 128 nodes, and the basis output parameterB=100B=100\. We adjusted the starting learning rate in the continuous interval \[1e\-5, 1\.0\), weight decay within \[1e\-10, 1\.0\), the output penalty coefficients in the range \[1e\-7, 100\), and the dropout and feature dropout coefficients in the discrete set\{0,0\.05,0\.1,0\.2,0\.3,0\.4,0\.5,0\.6,0\.7,0\.8,0\.9\}\\\{0,0\.05,0\.1,0\.2,0\.3,0\.4,0\.5,0\.6,0\.7,0\.8,0\.9\\\}\. The optimal hyper\-parameters were determined using the validation set and random search\. Table[S\-III](https://arxiv.org/html/2609.05946#S3.T3)lists the optimal hyper\-parameters for NBM across all datasets\.
Table S\-III:The parameter settings of the NBM on six datasets in this paper\.
### S\-III\.2Hyperparameters of Baseline Methods
For the pre\-process techniquesLFR\[[56](https://arxiv.org/html/2609.05946#bib.bib8)\]andreweighting\[[19](https://arxiv.org/html/2609.05946#bib.bib9)\]and the post\-process techniquesROC\[[20](https://arxiv.org/html/2609.05946#bib.bib19)\],LPP\[[50](https://arxiv.org/html/2609.05946#bib.bib53)\],Oracle\[[51](https://arxiv.org/html/2609.05946#bib.bib54)\], andFRAPPÉ\[[45](https://arxiv.org/html/2609.05946#bib.bib59)\]for improving fairness, we used the implementation of the AI Fairness 360 toolkit\[[6](https://arxiv.org/html/2609.05946#bib.bib47)\]or the source code provided by the authors with recommended parameter settings\.FairEMOL\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]was also configured with its recommended parameters, with the MLP architecture modified to include 256, 128, and 128 nodes in three hidden layers, instead of the original single hidden layer with 64 nodes for fair comparison\.
For theSNAM\[[52](https://arxiv.org/html/2609.05946#bib.bib16)\]andGAMI\-Net\[[53](https://arxiv.org/html/2609.05946#bib.bib15)\]methods, the parametersλ\\lambdaandη\\etacontrol the strength of feature selection, respectively, thus indirectly controlling the local interpretability\. Referring to the settings in\[[52](https://arxiv.org/html/2609.05946#bib.bib16)\]and\[[53](https://arxiv.org/html/2609.05946#bib.bib15)\], we setλ\\lambdaandη\\etato ten different parameter values, i\.e\.,λ=\\lambda=\{0\.01, 0\.03, 0\.05, 0\.08, 0\.1, 0\.3, 0\.5, 0\.7, 0\.9, 1\} andη=\\eta=\{5e\-4, 1e\-3, 3e\-3, 5e\-3, 8e\-3, 0\.01, 0\.02, 0\.03, 0\.04, 0\.05\}, to obtain models with varying degrees of local interpretability\.
## S\-IVSensitivity Analysis of the Smoothness Scaling Factorη\\eta
To systematically evaluate the impact of the hyperparameterη\\eta\(and the resulting tolerance thresholdkk\) on model interpretability, we conducted a targeted sensitivity analysis\. In this section, to isolate the effects of smoothness control and prevent interference from other interpretability dimensions, we specifically simplified the optimization problem to a two\-objective configuration: accuracy \(CE\) and smoothness \(ℳs\\mathcal\{M\}\_\{s\}\), while excluding the monotonicity metric \(ℳm\\mathcal\{M\}\_\{m\}\)\. We focused this analysis on theAdultdataset, exploring a range ofη\\etavalues including12,23,1\.0,43\\frac\{1\}\{2\},\\frac\{2\}\{3\},1\.0,\\frac\{4\}\{3\}, and32\\frac\{3\}\{2\}\. The resulting non\-dominated solution sets obtained under these different configurations are visualized in Fig\.[S\-1](https://arxiv.org/html/2609.05946#S4.F1)below\.
Figure S\-1:Pareto front distributions of accuracy \(ACC\) and smoothness \(ℳs\\mathcal\{M\}\_\{s\}\) across different scaling factorsη\\etaon theAdultdataset in one arbitrary trial\.As intuitively expected, the visualization demonstrates that asη\\etadecreases, the smoothness constraint becomes progressively more stringent \(corresponding to a smaller tolerance thresholdkk\)\. This is evidenced by the shift of the Pareto front toward lowerℳs\\mathcal\{M\}\_\{s\}values, yielding significantly more regularized and smoother shape functions\. Such results underscore the practical flexibility of theMONBMframework: it empowers users and stakeholders to explicitly calibrateη\\etaaccording to their specific requirements, enabling a controllable transition between high\-accuracy predictive modeling and highly interpretable, smooth global explanations\.
## S\-VRequired Monotonic Features
On each dataset, the monotonically increasing and decreasing features required in this paper are shown in Table[S\-IV](https://arxiv.org/html/2609.05946#S5.T4)\.
Table S\-IV:The required monotonically increasing and decreasing features on each dataset in this paper\.For example, consider the “hours\_per\_week” feature from theAdultdataset, where the task is to predict whether an individual earns more than $50,000 annually\. The “hours\_per\_week” feature describes the number of hours worked weekly\. We considered that the probability of his/her annual income exceeding $50,000 should increase monotonically as the number of hours worked increases, thus setting this feature to need to satisfy monotonically increasing\. It is worth noting that the setting is based on our experience and is only intended to validate the proposed metrics and methodology\. In practice, the settings should be determined according to specific needs\.
## S\-VIEvaluation Criteria
We used the comparative methods in\[[47](https://arxiv.org/html/2609.05946#bib.bib1)\]and\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]to compare the mutual dominance relationships betweenMONBMand population\-based and single solution\-based methods, respectively\.
For population\-based methods, we referred to the comparison method of\[[47](https://arxiv.org/html/2609.05946#bib.bib1)\]to compare the mutual dominance relationships between all pairs of solutions in two populations\. Specifically, assuming thatMONBMand the baseline method each produce 100 solutions, we performed pairwise comparisons, resulting in100×100=10,000100\\times 100=10,000total comparisons\. For example, if in 3000 of these comparisons the corresponding solution of theMONBMdominates the corresponding solution of the baseline method, the dominance probability ofMONBMover the baseline is3000/10,000=0\.3003000/10,000=0\.300, i\.e\., “Do\.” is 0\.300\.
For single solution\-based methods, we referred to the three metrics of\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]to compare the mutual dominance relationships of the two methods\. Given a set of modelsssand a set of model setsPPgenerated by the baseline method andMONBMin allnntrials, respectively, theDominatemetric evaluates whetherMONBMcan find solutions to dominate the baseline method as follows:
Dominate\(P,s\)=1n∑i=1nsign\[∑j=1\|Pi\|sign\[Pji≺si\]\>0\],\\textnormal\{Dominate\}\(P,s\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\textnormal\{sign\}\\left\[\\sum\_\{j=1\}^\{\|P^\{i\}\|\}\\textnormal\{sign\}\[P\_\{j\}^\{i\}\\prec s^\{i\}\]\>0\\right\],\(13\)wherePiP^\{i\}andsis^\{i\}are the model set and model in theii\-th trial,PjiP\_\{j\}^\{i\}is thejj\-th model inPiP^\{i\},Pji≺siP\_\{j\}^\{i\}\\prec s^\{i\}meansPjiP\_\{j\}^\{i\}dominatessis^\{i\}, andsign\[⋅\]\\textnormal\{sign\}\[\\cdot\]is equal to 1 if\[⋅\]\[\\cdot\]is true, otherwise 0\. Larger values of theDominatemetric indicate that in more trials the solutions obtained by the baseline method are dominated by some of the solutions in the model set obtained byMONBM\.
TheIncomparablemetric calculates the average proportion of solutionsPPobtained byMONBMthat are incomparable to solutionsssobtained by the baseline method across allnntrials, as follows:
Incomparable\(P,s\)=1n∑i=1n1\|Pi\|∑j=1\|Pi\|sign\[\(Pji⊀si\)&\(si⊀Pji\)\]\>0,\\textnormal\{Incomparable\}\(P,s\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\frac\{1\}\{\|P^\{i\}\|\}\\sum\_\{j=1\}^\{\|P^\{i\}\|\}\\textnormal\{sign\}\[\(P\_\{j\}^\{i\}\\not\\prec s^\{i\}\)\\&\(s^\{i\}\\not\\prec P\_\{j\}^\{i\}\)\]\>0,\(14\)wherePji⊀siP\_\{j\}^\{i\}\\not\\prec s^\{i\}meansPjiP\_\{j\}^\{i\}does not dominatesis^\{i\}\. A largerIncomparablemetric means that more models obtained byMONBMare incomparable withss\.
TheDominatedmetric calculates the average proportion of solutions obtained byMONBMare dominated by solutionsssobtained by the baseline method over allnntrials, as follows:
Dominated\(P,s\)=1n∑i=1n1\|Pi\|∑j=1\|Pi\|sign\[si≺Pji\]\.\\textnormal\{Dominated\}\(P,s\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\frac\{1\}\{\|P^\{i\}\|\}\\sum\_\{j=1\}^\{\|P^\{i\}\|\}\\textnormal\{sign\}\[s^\{i\}\\prec P\_\{j\}^\{i\}\]\.\(15\)
A largerDominatedvalues means that more models obtained byMONBMare dominated byss\. In fact, theIncomparableandDominatedmetrics are consistent with the population\-based approach to comparing the proportion of mutually non\-dominated and dominated\[[47](https://arxiv.org/html/2609.05946#bib.bib1)\]\.
## S\-VIIResults on the Training Set
This section supplements the performance of the solutions selected in Section 4\.2\.2 of the main text on the training set to discuss the overfitting phenomenon\.
Table S\-V:Comparison of the mean values of theACCandDPon the training set of the models corresponding to the three points chosen from the final solution set obtained by theMONBM\-CDinstantiation over 15 trials\.ACCvalues are expressed as percentages \(%\\%\)\. The best average is highlighted with a gray background\. The difference between it and the best value is shown in parentheses\.Table S\-VI:Comparison of the mean values of theACCandLIon the training set of the models corresponding to the three points chosen from the final solution set obtained by theMONBM\-CLinstantiation over 15 trials\.ACCvalues are expressed as percentages \(%\\%\)\. The best average is highlighted with a gray background\. The difference between it and the best value is shown in parentheses\.Table S\-VII:Comparison of the mean values of theACCandGIon the training set of the models corresponding to the three points chosen from the final solution set obtained by theMONBM\-CGinstantiation over 15 trials\.ACCvalues are expressed as percentages \(%\\%\)\. The best average is highlighted with a gray background\. The difference between it and the best value is shown in parentheses\.Table S\-VIII:Comparison of the mean values of theACC,DP,LI, andGIon the training set of the models corresponding to the four extreme points chosen from the final solution set obtained by theMONBM\-CDLGinstantiation over 15 trials\.ACCvalues are expressed as percentages \(%\\%\)\. The best average is highlighted with a gray background\. The difference between it and the best value is shown in parentheses\. The “Majority Baseline” row represents the accuracy of a constant classifier that consistently predicts the majority class for each dataset\.
## S\-VIIISensitivity Analysis to Fairness Metrics: Results on Equalized Odds \(EOD\)
To address the concern that the observed trade\-off trends might be specific to theDPmetric, we conducted a sensitivity analysis by optimizing the equalized odds \(EOD\) metric—a label\-aware fairness definition that constrains the model to maintain equal true positive and false positive rates across subgroups\. We implemented theMONBM\-CEinstantiation, considering both accuracy andEODobjectives\.
Table[S\-IX](https://arxiv.org/html/2609.05946#S8.T9)summarizes the performance of the models at the extreme and “middleEOD” points\. The empirical results reveal a trade\-off pattern remarkably consistent with that observed forDP\(see Table 2 in the main text\)\. Specifically, moderateEODimprovement \(defined as a 50% reduction in theEODmetric relative to its extremes, represented by the “middleEOD” point\) has a marginal impact on accuracy, with an averageACCdecrease of only 0\.64% across six datasets\. However, pursuing extreme fairness \(EOD≈\\approx0\) leads to substantial accuracy declines on theAdult,COMPAS, andDutchdatasets, while its impact remains moderate on theBank,Default, andLSATdatasets\. This confirms that the inherent difficulty of balancing accuracy and fairness is a robust phenomenon across different fairness definitions\.
Table S\-IX:Comparison of the mean values of theACCandEODon the test set of the models corresponding to the three points chosen from the final solution set obtained by theMONBM\-CEinstantiation over 15 trials\.ACCvalues are expressed as percentages \(%\\%\)\. The best average is highlighted with a gray background\. The difference between it and the best value is shown in parentheses\.Furthermore, Fig\.[S\-2](https://arxiv.org/html/2609.05946#S8.F2)visualizes the global explanations for theMONBM\-CEinstantiation\. A cross\-comparison with theDP\-optimized results \(Fig\. 5 and Fig\. S\-2\) shows that the shape functions exhibit highly similar trends and feature\-importance shifts\. Notably, to satisfyEODconstraints, the framework still assigns distinct weights to different values of the sensitive attribute \(e\.g\., sex or race\)\. This observation suggests that, given the intrinsic bias present in these datasets, even label\-aware metrics likeEODnecessitate significant adjustments to the sensitive attribute’s contribution to compensate for disparate error distributions among groups\. These results reinforce the conclusion that the interpretability profiles and trade\-off insights provided byMONBMare generalizable and not limited to a single fairness metric\.
\(a\)Global explanation results ofMONBM\-CEinstantiation on theAdultdataset
\(b\)Global explanation results ofMONBM\-CEinstantiation on theBankdataset
\(c\)Global explanation results ofMONBM\-CEinstantiation on theCOMPASdataset
\(d\)Global explanation results ofMONBM\-CEinstantiation on theDefaultdataset
\(e\)Global explanation results ofMONBM\-CEinstantiation on theDutchdataset
\(f\)Global explanation results ofMONBM\-CEinstantiation on theLSATdataset
Figure S\-2:Comparison of global explanations to assess the impact of improved fairness on each feature\. For each dataset, the blue, red, and green lines represent the global explanation results of the model corresponding to the bestCE, “middleEOD”, and bestEODpoints obtained by theMONBM\-CEinstantiation in one arbitrary trial, respectively\. Each subfigure corresponds to the global explanation of a feature, with the x\-axis representing feature values and the y\-axis showing the outputs of the shape functions\. The sensitive attribute is highlighted in red\. The Pearson correlation coefficient between each attribute and the sensitive attribute is annotated in the respective subplots\.
## S\-IXStability and Reliability of Explanations Across Independent Trials
A systematic cross\-analysis of the 15 independent trials \(Figs\.[S\-6](https://arxiv.org/html/2609.05946#S10.F6),[S\-8](https://arxiv.org/html/2609.05946#S10.F8),[S\-10](https://arxiv.org/html/2609.05946#S10.F10),[S\-12](https://arxiv.org/html/2609.05946#S10.F12), and[S\-14](https://arxiv.org/html/2609.05946#S10.F14)\) reveals a clear relationship between predictive performance and explanatory stability\. Specifically, the standard deviation of feature importance remains relatively low near the “Best CE” solutions but increases at the “Middle” and “Extreme” operating points along the fairness and local interpretability dimensions\. Two main factors contribute to this behavior:
1. 1\.When strong fairness or interpretability constraints are imposed, predictive accuracy often decreases\. Under such conditions, the optimization landscape becomes less sharply defined, allowing multiple feature configurations to satisfy the constraints\. Consequently, the learned decision logic becomes more sensitive to variations in the sampled training data\.
2. 2\.The “middle points” are selected as knee points from the non\-dominated set of each independent run\. Since the Pareto front shifts across different training samples, the “middle point” represents a varying compromise solution rather than a fixed parameter set, naturally contributing to the fluctuations in aggregate results\.
Overall, while the main trade\-off trends identified byMONBMremain consistent across trials, these analyses indicate that the stability of explanations is closely linked to predictive performance and the position of the model on the trade\-off surface\.
It is important to note that the observed variance in per\-feature importance does not imply that the method itself is unreliable\. Since GAM\-based explanations are 100% faithful to the underlying model, differences in shape functions across runs primarily reflect variations in the training samples rather than shortcomings of the framework\.
## S\-XAdditional Explanation Results
This section supplements the global and local explanation results for datasets not shown on the various instantiations in Section 4\.2 of the main text\.
\(a\)Global explanation results for continuous features ofMONBM\-CGinstantiation on theCOMPASdataset
\(b\)Global explanation for continuous features ofMONBM\-CGinstantiation on theDefaultdataset
\(c\)Global explanation for continuous features ofMONBM\-CGinstantiation on theDutchdataset
\(d\)Global explanation for continuous features ofMONBM\-CGinstantiation on theLSATdataset
Figure S\-3:Comparison of global explanations of continuous features to validate the proposed smoothness and monotonicity metrics\. For each dataset, the blue, red, and green lines represent the global explanation results of the continuous features by the model corresponding to the bestCE, “middleGI,” and bestGIpoints obtained by theMONBM\-CGinstantiation in one arbitrary trial, respectively\. Each subfigure corresponds to the global explanation of a feature, with the x\-axis representing feature values and the y\-axis showing the outputs of the shape functions\. Each continuous feature has a smoothness requirement, and features highlighted in red also have a monotonicity requirement\.\(a\)Distribution of shape functions for continuous features across non\-dominated knee\-point solutions for theMONBM\-CGinstantiation on theAdultdataset\.
\(b\)Distribution of shape functions for continuous features across non\-dominated knee\-point solutions for theMONBM\-CGinstantiation on theBankdataset\.
\(c\)Distribution of shape functions for continuous features across non\-dominated knee\-point solutions for theMONBM\-CGinstantiation on theCOMPASdataset\.
\(d\)Distribution of shape functions for continuous features across non\-dominated knee\-point solutions for theMONBM\-CGinstantiation on theDefaultdataset\.
\(e\)Distribution of shape functions for continuous features across non\-dominated knee\-point solutions for theMONBM\-CGinstantiation on theDutchdataset\.
\(f\)Distribution of shape functions for continuous features across non\-dominated knee\-point solutions for theMONBM\-CGinstantiation on theLSATdataset\.
Figure S\-4:Variation of global explanations for continuous features among non\-dominated knee\-point solutions\. In each subfigure, the collection of curves represents the shape functions of continuous features for the set of non\-dominated knee\-point solutions obtained from the final generation of theMONBM\-CGinstantiation in one arbitrary trial\. The color gradient, ranging from deep purple \(HighACC, prioritizing predictive accuracy\) to bright yellow \(HighGI, emphasizing smoothness and monotonicity\), illustrates the gradual transition in the continuous shape functions across the trade\-off surface\.\(a\)Global explanation results ofMONBM\-CDinstantiation on theAdultdataset
\(b\)Global explanation results ofMONBM\-CDinstantiation on theBankdataset
\(c\)Global explanation results ofMONBM\-CDinstantiation on theDefaultdataset
\(d\)Global explanation results ofMONBM\-CDinstantiation on theDutchdataset
Figure S\-5:Comparison of global explanations to assess the impact of improved fairness on each feature\. For each dataset, The blue, red, and green lines represent the global explanation results of the model corresponding to the bestCE, “middleDP”, and bestDPpoints obtained by theMONBM\-CDinstantiation in one arbitrary trial, respectively\. Each subfigure corresponds to the global explanation of a feature, with the x\-axis representing feature values and the y\-axis showing the outputs of the shape functions\. The sensitive attribute is highlighted in red\. The Pearson correlation coefficient between each attribute and the sensitive attribute is annotated in the respective subplots\.\(a\)Mean global explanation results ofMONBM\-CDinstantiation over 15 independent trials on theAdultdataset
\(b\)Mean global explanation results ofMONBM\-CDinstantiation over 15 independent trials on theBankdataset
\(c\)Mean global explanation results ofMONBM\-CDinstantiation over 15 independent trials on theCOMPASdataset
\(d\)Mean global explanation results ofMONBM\-CDinstantiation over 15 independent trials on theDefaultdataset
\(e\)Mean global explanation results ofMONBM\-CDinstantiation over 15 independent trials on theDutchdataset
\(f\)Mean global explanation results ofMONBM\-CDinstantiation over 15 independent trials on theLSATdataset
Figure S\-6:Comparison of global explanations to assess the impact of improved fairness on each feature\. For each dataset, the blue, red, and green lines and their corresponding shaded areas represent the mean global explanation results and standard deviations of the models corresponding to the bestCE, “middleDP”, and bestDPpoints obtained by theMONBM\-CDinstantiation averaged over 15 independent trials, respectively\. Each subfigure corresponds to the global explanation of a feature, with the x\-axis representing feature values and the y\-axis showing the outputs of the shape functions\. The sensitive attribute is highlighted in red\. The Pearson correlation coefficient between each attribute and the sensitive attribute is annotated in the respective subplots\.\(a\)Average local explanation ofMONBM\-CLinstantiation on theAdultdataset
\(b\)Average local explanation ofMONBM\-CDinstantiation on theAdultdataset
\(c\)Average local explanation ofMONBM\-CLinstantiation on theCOMPASdataset
\(d\)Average local explanation ofMONBM\-CDinstantiation on theCOMPASdataset
\(e\)Average local explanation ofMONBM\-CLinstantiation on theDefaultdataset
\(f\)Average local explanation ofMONBM\-CDinstantiation on theDefaultdataset
\(g\)Average local explanation ofMONBM\-CLinstantiation on theDutchdataset
\(h\)Average local explanation ofMONBM\-CDinstantiation on theDutchdataset
\(i\)Average local explanation ofMONBM\-CLinstantiation on theLSATdataset
\(j\)Average local explanation ofMONBM\-CDinstantiation on theLSATdataset
Figure S\-7:Comparison of the average local explanations to assess the impact of improvedLIon each feature\. For each dataset, the three columns respectively display the average local explanation results of the model corresponding to: \(1\) the bestCEpoint, \(2\) the “middleLI” or “middleDP” point, and \(3\) the bestLIor bestDPpoint, obtained by theMONBM\-CLorMONBM\-CDinstantiation in one arbitrary trial\. For each local explanation, the y\-axis represents the features and the x\-axis represents the average of the absolute values of the shape function outputs corresponding to the features for all data points in the test set\.\(a\)Mean local explanation and standard deviationMONBM\-CLinstantiation over 15 independent trials on theAdultdataset
\(b\)Mean local explanation and standard deviationMONBM\-CDinstantiation over 15 independent trials on theAdultdataset
\(c\)Mean local explanation and standard deviationMONBM\-CLinstantiation over 15 independent trials on theBankdataset
\(d\)Mean local explanation and standard deviationMONBM\-CDinstantiation over 15 independent trials on theBankdataset
\(e\)Mean local explanation and standard deviationMONBM\-CLinstantiation over 15 independent trials on theCOMPASdataset
\(f\)Mean local explanation and standard deviationMONBM\-CDinstantiation over 15 independent trials on theCOMPASdataset
\(g\)Mean local explanation and standard deviationMONBM\-CLinstantiation over 15 independent trials on theDefaultdataset
\(h\)Mean local explanation and standard deviationMONBM\-CDinstantiation over 15 independent trials on theDefaultdataset
\(i\)Mean local explanation and standard deviationMONBM\-CLinstantiation over 15 independent trials on theDutchdataset
\(j\)Mean local explanation and standard deviationMONBM\-CDinstantiation over 15 independent trials on theDutchdataset
\(k\)Mean local explanation and standard deviationMONBM\-CLinstantiation over 15 independent trials on theLSATdataset
\(l\)Mean local explanation and standard deviationMONBM\-CDinstantiation over 15 independent trials on theLSATdataset
Figure S\-8:Comparison of the average local explanations to assess the impact of improvedLIon each feature\. For each dataset, the three columns respectively display the mean and standard deviation of the average local explanation results of the model corresponding to: \(1\) the bestCEpoint, \(2\) the “middleLI” or “middleDP” point, and \(3\) the bestLIor bestDPpoint, obtained by theMONBM\-CLorMONBM\-CGinstantiation averaged over 15 independent trials\. For each local explanation, the y\-axis represents the features and the x\-axis represents the average of the absolute values of the shape function outputs corresponding to the features for all data points in the test set, with error bars representing the standard deviation across the 15 trials\.\(a\)Global explanation results ofMONBM\-CGinstantiation on theAdultdataset
\(b\)Global explanation results ofMONBM\-CGinstantiation on theCOMPASdataset
\(c\)Global explanation results ofMONBM\-CGinstantiation on theDefaultdataset
\(d\)Global explanation results ofMONBM\-CGinstantiation on theDutchdataset
\(e\)Global explanation results ofMONBM\-CGinstantiation on theLSATdataset
Figure S\-9:Comparison of global explanations to assess the impact of improvedGIon each feature\. For each dataset, the blue, red, and green lines represent the global explanation results of the model corresponding to the bestCE, “middleGI,” and bestGIpoints obtained by theMONBM\-CGinstantiation in one arbitrary trial, respectively\. Each subfigure corresponds to the global explanation of a feature, with the x\-axis representing feature values and the y\-axis showing the outputs of the shape functions\. The continuous features are highlighted in red\.\(a\)Mean global explanation results ofMONBM\-CGinstantiation over 15 independent trials on theAdultdataset
\(b\)Mean global explanation results ofMONBM\-CGinstantiation over 15 independent trials on theBankdataset
\(c\)Mean global explanation results ofMONBM\-CGinstantiation over 15 independent trials on theCOMPASdataset
\(d\)Mean global explanation results ofMONBM\-CGinstantiation over 15 independent trials on theDefaultdataset
\(e\)Mean global explanation results ofMONBM\-CGinstantiation over 15 independent trials on theDutchdataset
\(f\)Mean global explanation results ofMONBM\-CGinstantiation over 15 independent trials on theLSATdataset
Figure S\-10:Comparison of global explanations to assess the impact of improved fairness on each feature\. For each dataset, the blue, red, and green lines and their corresponding shaded areas represent the mean global explanation results and standard deviations of the models corresponding to the bestCE, “middleGI”, and bestGIpoints obtained by theMONBM\-CGinstantiation averaged over 15 independent trials, respectively\. Each subfigure corresponds to the global explanation of a feature, with the x\-axis representing feature values and the y\-axis showing the outputs of the shape functions\. The continuous features are highlighted in red\.\(a\)Average local explanation ofMONBM\-CGinstantiation onAdultdataset
\(b\)Average local explanation ofMONBM\-CGinstantiation onCOMPASdataset
\(c\)Average local explanation ofMONBM\-CGinstantiation onDefaultdataset
\(d\)Average local explanation ofMONBM\-CGinstantiation onDutchdataset
\(e\)Average local explanation ofMONBM\-CGinstantiation onLSATdataset
Figure S\-11:Comparison of the average local explanations to assess the impact of improvedGIon each feature\. For each dataset, the three columns represent the average local explanation results of the model corresponding to the bestCE, “middleGI,” and bestGIpoints obtained by theMONBM\-CGinstantiation in one arbitrary trial, respectively\. For each local explanation, the y\-axis represents the features and the x\-axis represents the average of the absolute values of the shape function outputs corresponding to the features for all data points in the test set\.\(a\)Mean local explanation and standard deviation ofMONBM\-CGinstantiation over 15 independent trials on theAdultdataset
\(b\)Mean local explanation and standard deviation ofMONBM\-CGinstantiation over 15 independent trials on theBankdataset
\(c\)Mean local explanation and standard deviation ofMONBM\-CGinstantiation over 15 independent trials on theCOMPASdataset
\(d\)Mean local explanation and standard deviation ofMONBM\-CGinstantiation over 15 independent trials on theDefaultdataset
\(e\)Mean local explanation and standard deviation ofMONBM\-CGinstantiation over 15 independent trials on theDutchdataset
\(f\)Mean local explanation and standard deviation ofMONBM\-CGinstantiation over 15 independent trials on theLSATdataset
Figure S\-12:Comparison of the average local explanations to assess the impact of improvedGIon each feature\. For each dataset, the three columns represent the mean and standard deviation of the average local explanation results of the model corresponding to the bestCE, “middleGI,” and bestGIpoints obtained by theMONBM\-CGinstantiation averaged over 15 independent trials, respectively\. For each local explanation, the y\-axis represents the features and the x\-axis represents the average of the absolute values of the shape function outputs corresponding to the features for all data points in the test set, with error bars representing the standard deviation across the 15 trials\.\(a\)Average local explanation ofMONBM\-CDLGinstantiation on theBankdataset
\(b\)Average local explanation ofMONBM\-CDLGinstantiation on theCOMPASdataset
\(c\)Average local explanation ofMONBM\-CDLGinstantiation on theDefaultdataset
\(d\)Average local explanation ofMONBM\-CDLGinstantiation on theLSATdataset
Figure S\-13:Comparison of the average local explanations to explore the interrelationships among the four objectives\. For each dataset, the four columns represent the average local explanation results of the model corresponding to the bestCE, bestDP, bestLI, and bestGIpoints obtained by theMONBM\-CDLGinstantiation in one arbitrary trial, respectively\. For each local explanation, the y\-axis represents the features and the x\-axis represents the average of the absolute values of the shape function outputs corresponding to the features for all data points in the test set\.\(a\)Mean local explanation and standard deviation ofMONBM\-CDLGinstantiation over 15 independent trials on theAdultdataset
\(b\)Mean local explanation and standard deviation ofMONBM\-CDLGinstantiation over 15 independent trials on theBankdataset
\(c\)Mean local explanation and standard deviation ofMONBM\-CDLGinstantiation over 15 independent trials on theCOMPASdataset
\(d\)Mean local explanation and standard deviation ofMONBM\-CDLGinstantiation over 15 independent trials on theDefaultdataset
\(e\)Mean local explanation and standard deviation ofMONBM\-CDLGinstantiation over 15 independent trials on theDutchdataset
\(f\)Mean local explanation and standard deviation ofMONBM\-CDLGinstantiation over 15 independent trials on theLSATdataset
Figure S\-14:Comparison of the average local explanations to explore the interrelationships among the four objectives\. For each dataset, the four columns represent the mean and standard deviation of the average local explanation results of the model corresponding to the bestCE, bestDP, bestLI, and bestGIpoints obtained by theMONBM\-CDLGinstantiation averaged over 15 independent trials, respectively\. For each local explanation, the y\-axis represents the features and the x\-axis represents the average of the absolute values of the shape function outputs corresponding to the features for all data points in the test set, with error bars representing the standard deviation across the 15 trials\.\(a\)Distribution ofMONBM\-CDLGsolutions on theBankdataset
\(b\)Distribution ofMONBM\-CDLGsolutions on theCOMPASdataset
\(c\)Distribution ofMONBM\-CDLGsolutions on theDefaultdataset
\(d\)Distribution ofMONBM\-CDLGsolutions on theLSATdataset
Figure S\-15:Distribution of allMONBM\-CDLGnon\-dominated solutions across pairs of trustworthy objectives to examine their interrelationships\. Each subplot visualizes the solutions obtained from one arbitrary run of theMONBM\-CDLGinstantiation on a given dataset, showing pairwise relationships betweenDP,LI, andGI\.Figure S\-16:Performance of the non\-dominant knee point solutions obtained byMONBM\-CGin terms of model accuracy andGImetrics on one arbitrary trial\.Figure S\-17:Performance of the non\-dominated knee point solutions obtained byMONBM\-CDLGin terms of model accuracy,DP,LI, andGImetrics on one arbitrary trial\.
## S\-XIAblation Analysis of Global Interpretability: Smoothness versus Monotonicity
To better understand how global interpretability interacts with other trustworthy dimensions, we conduct an ablation study to examine the individual effects of its two components: smoothness \(ℳs\\mathcal\{M\}\_\{s\}\) and monotonicity \(ℳm\\mathcal\{M\}\_\{m\}\)\. Specifically, we investigate whether the relationships observed in the original four\-objective formulation are primarily driven by one of these two constraints\. To this end, we decomposeMONBM\-CDLGinto two variants: Accuracy–Fairness–LI–Smoothness \(MONBM\-CDLS\) and Accuracy–Fairness–LI–Monotonicity \(MONBM\-CDLM\)\. Consistent with the analysis in the main text, experiments are conducted on theAdultandDutchdatasets, representing typical and atypical relationship patterns, respectively\. The extreme solutions extracted from the final non\-dominated sets are summarized in Table[S\-X](https://arxiv.org/html/2609.05946#S11.T10)and Table[S\-XI](https://arxiv.org/html/2609.05946#S11.T11)below\.
Table S\-X:Comparison of the mean values of theACC,DP,LI, andℳs\\mathcal\{M\}\_\{s\}on the test set of the models corresponding to the four extreme points chosen from the final solution set obtained by theMONBM\-CDLSinstantiation on theAdultandDutchdatasets over 15 trials\.ACCvalues are expressed as percentages \(%\\%\)\.Table S\-XI:Comparison of the mean values of theACC,DP,LI, andℳm\\mathcal\{M\}\_\{m\}on the test set of the models corresponding to the four extreme points chosen from the final solution set obtained by theMONBM\-CDLMinstantiation on theAdultandDutchdatasets over 15 trials\.ACCvalues are expressed as percentages \(%\\%\)\.The results of both ablated formulations generally support the observations reported in the main text\. First, compared with fairness and local interpretability, both smoothness and monotonicity impose relatively limited performance degradation on predictive accuracy\. For example, models optimized toward extreme global constraints \(BestGI\) maintain relatively high predictive performance across datasets, whereas pursuing extremeDPoften causes substantially larger accuracy reductions\. This suggests that global interpretability constraints mainly reshape local geometric properties of the learned shape functions \(making them smoother or more monotonic\), rather than fundamentally altering the overall decision boundary\.
Second, a persistent trade\-off between local interpretability and both global interpretability components can still be observed\. ImprovingLIencourages sparse feature usage, whereas smoothness and monotonicity regulate the global structure of feature responses\. These two forms of interpretability therefore operate through different mechanisms and cannot always be improved simultaneously\.
More importantly, the ablation study provides additional insight into the atypical behavior observed on theDutchdataset\. Comparing the two formulations reveals a clear difference between smoothness and monotonicity\. Under monotonicity optimization \(Bestℳm\\mathcal\{M\}\_\{m\}, Table[S\-XI](https://arxiv.org/html/2609.05946#S11.T11)\), fairness remains relatively stable \(DP= 0\.054\), indicating that encouraging monotonic feature trends does not substantially interfere with fair decision making\. In contrast, optimizing smoothness alone \(Bestℳs\\mathcal\{M\}\_\{s\}, Table[S\-X](https://arxiv.org/html/2609.05946#S11.T10)\) leads to a much larger demographic disparity \(DP= 0\.576\)\. This contrast suggests that the fairness–interpretability conflict observed inDutchis primarily associated with smoothness rather than monotonicity\. The underlying reason is that the pursuit of smoothness suppresses local fluctuations in continuous shape functions and promotes a more uniform feature response\. On theDutchdataset, this smoothing effect reduced the contribution of other features, causing the sensitive attribute “sex” to become the dominant feature and exacerbating decision\-making bias\.
Overall, these findings suggest that global interpretability should not be treated as a single homogeneous concept\. Different interpretability constraints may interact with fairness in different ways, and such interactions depend strongly on the characteristics of the underlying data distribution\.
## S\-XIIPreliminary Exploration on Adaptive Selection Pressure
In many\-objective optimization scenarios such as the four\-objectiveMONBM\-CDLGinstantiation, evolutionary search often becomes less effective due to increasingly severe conflicts among objectives and weakened selection pressure\. As the number of objectives grows, distinguishing promising solutions becomes more difficult, which may reduce convergence performance and the overall quality of the final solution set\. To explore whether this limitation can be mitigated, we conduct a preliminary study by introducing an adaptive selection pressure mechanism into the SRA adopted in our framework\.
Specifically, the parameterpcp\_\{c\}in SRA controls the probability that stochastic ranking prioritizes convergence\-oriented information during pairwise comparison and sorting\. A largerpcp\_\{c\}increases the likelihood that solutions with better objective performance are ranked higher, thereby strengthening selection pressure and encouraging convergence\. In contrast, a smallerpcp\_\{c\}allows more emphasis on maintaining diverse candidate solutions and preserving exploration capability\. Therefore,pcp\_\{c\}naturally serves as a controllable mechanism for balancing exploration and convergence throughout the evolutionary process\. Based on this observation, we introduce a dynamic parameter schedule to gradually adjust selection pressure during optimization:
pc\(g\)=pstart\+\(pend−pstart\)×\(gGmax\)2,p\_\{c\}\(g\)=p\_\{start\}\+\(p\_\{end\}\-p\_\{start\}\)\\times\\left\(\\frac\{g\}\{G\_\{max\}\}\\right\)^\{2\},\(16\)whereggrepresents the current generation number,GmaxG\_\{max\}denotes the maximum number of generations,pstartp\_\{start\}represents the initial probability, andpendp\_\{end\}denotes the target probability at the final generation\. This quadratic schedule enables broader exploration in early generations by maintaining relatively weak selection pressure, while progressively increasing convergence pressure during later stages of optimization\.
In our implementation,pstartp\_\{\\textit\{start\}\}is initialized as 0\.45, consistent with the fixed configuration adopted in the originalMONBMframework, andpendp\_\{\\textit\{end\}\}is set to 0\.8\. To evaluate the effect of this adaptive mechanism, we conduct experiments on theAdultdataset using the four\-objectiveMONBM\-CDLGinstantiation\. Table[S\-XII](https://arxiv.org/html/2609.05946#S12.T12)presents the comparison between the originalMONBMwith fixedpcp\_\{c\}\(denoted asMONBMw/o AS\) and the adaptive variant \(denoted asMONBMwith AS\)\. Following the analysis approach in the main experiments, we report the four extreme solutions selected from the final solution set, where each solution corresponds to the model achieving the best value for one optimization objective\.
Table S\-XII:Comparison of the mean values of theACC,DP,LI, andGIon the test set of the models corresponding to the four extreme points chosen from the final solution set obtained by theMONBM\-CDLGinstantiation with and without the adaptive strategy \(MONBMwith AS vs\.MONBMw/o AS\) on theAdultdataset over 15 trials\.ACCvalues are expressed as percentages \(%\\%\)\. For each metric, the better average betweenMONBMwith AS andMONBMw/o AS is highlighted in bold\.The experimental results show that introducing adaptive selection pressure generally improves solution quality, especially for the fairness\-, local interpretability\-, and global interpretability\-oriented solutions, while maintaining competitive predictive performance\. Although the performance gap relative to lower\-dimensional optimization settings is not completely eliminated—which is expected given the inherent trade\-offs in many\-objective learning—the results suggest that adaptive pressure adjustment is a promising direction for improving convergence under severe objective conflicts\. Developing more advanced adaptive mechanisms or dedicated many\-objective operators remains an important topic for future research\.
## S\-XIIIComparingMONBMInstantiations with Baseline Methods on All Four Objectives
This section compares the mutual dominance relationships between theMONBM\-CD,MONBM\-CL, andMONBM\-CDLGinstantiations and their corresponding baseline methods across all four objectives considered, namelyACC,DP,LI, andGI\. The results are shown in Tables[S\-XIII](https://arxiv.org/html/2609.05946#S13.T13),[S\-XIV](https://arxiv.org/html/2609.05946#S13.T14), and[S\-XV](https://arxiv.org/html/2609.05946#S13.T15)\. For the four\-objective instantiation \(MONBM\-CDLG\), we compared it against all ten baseline methods considered in this paper\. It is worth noting thatFairEMOL\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\], being a black\-box method, lacks both local and global interpretability and is therefore excluded from this comparison\.
From Tables[S\-XIII](https://arxiv.org/html/2609.05946#S13.T13)and[S\-XIV](https://arxiv.org/html/2609.05946#S13.T14), we can see that compared to the two\-objective comparison \(i\.e\., the results in Tables 7 and 8 in the main text\), the probability ofMONBM\-CDorMONBM\-CLdominating or being dominated by other methods on all four objectives has significantly decreased, even reaching 0\.0\. This is intuitive, as none of the methods considered additional objectives during optimization, and these objectives often conflict with each other\. Therefore,MONBMis also difficult to outperform baseline methods in terms of additional objectives and naturally difficult to be dominated by other methods\. However,MONBM\-CDinstantiation still has a non\-negligible probability of finding solutions that dominate other methods, which is partially attributed to its prominent advantages in the two dimensions considered, namely accuracy and fairness\.
Table S\-XIII:Dominance relationship betweenMONBM\-CDand other methods in terms ofACC,DP,LI, andGImetrics\. Results are averaged over 15 trials\. “Do\.”, “Non\-Do\.”, and “Be\-Do\.” mean the probability thatMONBM\-CDdominates, does not dominate each other, and is dominated by the corresponding method\. Dominance metrics for all baseline methods follow\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]\.Table S\-XIV:Dominance relationship betweenMONBM\-CLand other methods in terms ofACC,DP,LI, andGImetrics\. Results are averaged over 15 trials\. “Do\.”, “Non\-Do\.”, and “Be\-Do\.” mean the probability thatMONBM\-CLdominates, does not dominate each other, and is dominated by the corresponding method\. Dominance metrics for all baseline methods follow\[[58](https://arxiv.org/html/2609.05946#bib.bib7)\]\.Regarding the comparison results of theMONBM\-CDLGinstantiation in Table[S\-XV](https://arxiv.org/html/2609.05946#S13.T15), several key findings can be observed\. First, the solutions obtained byMONBM\-CDLGare not dominated by any of the ten baseline methods across all datasets \(the “Be\-Do\.” probability is consistently 0\.000\), confirming the framework’s robustness in maintaining a Pareto\-optimal frontier\. Second, on baselines with relatively poor overall performance, such asLFR\-NBMandROC\-NBM,MONBM\-CDLGexhibits a high probability of finding solutions that dominate them across all four objectives simultaneously \(e\.g\., a mean “Do\.” probability of 0\.789 againstLFR\-NBM\)\.
However, for the majority of competitive baselines,MONBM\-CDLGeither rarely or only with low probability achieves complete dominance across all four objectives\. This is intuitive: while our specialized instantiations \(MONBM\-CDandMONBM\-CL\) consistently outperform baselines in their respective target dimensions, this localized advantage is often insufficient to guarantee simultaneous superiority across all four conflicting objectives once two additional dimensions are introduced\. Nevertheless, the primary value ofMONBM\-CDLGlies in its ability to provide a diverse set of high\-quality compromise solutions that remain competitive and non\-dominated across the entire trustworthiness spectrum, offering researchers more flexible trade\-off options than specialized baselines\.
Table S\-XV:Dominance relationship betweenMONBM\-CDLGand 10 baseline methods in terms ofACC,DP,LI, andGImetrics\. Results are averaged over 15 trials\. “Do\.”, “Non\-Do\.”, and “Be\-Do\.” mean the probability that our method dominates, does not dominate each other, and is dominated by the corresponding method\.Similar Articles
A Unified Framework for Gradient Aggregation in Multi-Objective Optimization
This paper presents a unified theoretical framework for gradient aggregation in multi-objective optimization, establishing convergence rates to Pareto stationarity. The authors introduce a sufficient alignment condition and demonstrate its application to existing and new algorithms, such as capped MGDA.
Multi-Objective Multi-Agent Bandits: From Learning Efficiency to Fairness Optimization
This paper introduces Pareto UCB1 Gossip and Simulated NSW UCB Gossip for multi-objective multi-agent multi-armed bandits, addressing both learning efficiency and fairness in stochastic environments.
AHEAD: Advancing Multi-Class Label Aggregation with Interpretable Cross-Annotator Modeling
This paper proposes AHEAD, a cross-annotator learning framework for multi-class label aggregation that uses graph neural networks to model annotator reliability, achieving significant accuracy improvements on 10 real-world datasets spanning NLP, CV, Video, and Audio.
Generalized Multimodal Foundation Model
This paper introduces a generalized multimodal foundation model capable of handling arbitrary modality combinations and prediction tasks, achieving competitive performance through training on large-scale synthetic datasets with diverse causal structures.
Generative Interpretability via Scalable Neuro-Symbolic Models
The paper argues for a paradigm shift from post-hoc to generative interpretability in AI, proposing neuro-symbolic models to enable human-understandable checkpoints and causal intervention for safe LLM deployment in agentic systems.