AquaAugmentor: A Novel Feature Augmentation Algorithm for Water Potability Prediction

arXiv cs.LG Papers

Summary

This paper introduces AquaAugmentor, a novel feature augmentation algorithm to enhance the predictive performance of machine learning and deep learning models for water potability classification using chemical attributes.

arXiv:2607.15775v1 Announce Type: new Abstract: Access to potable water is crucial for health, economic development, and sustainability. However, accurately classifying water quality remains a significant challenge due to the complexity and variability of water source data. This paper addresses the challenge of predicting water potability through machine learning and deep learning algorithms. It introduces a novel feature augmentation algorithm, AquaAugmentor, to enhance the predictive performance of these models for low-dimensional datasets. Utilizing a dataset that includes chemical attributes of water, such as pH, hardness, solids, chloramines, sulfate, and others. This study evaluates the performance of the models with and without AquaAugmentor. Each model applied to classify water as potable or non-potable and its performance is then evaluated and compared based on test accuracy and AUC score. The results highlight the strengths and limitations of our proposed algorithm, providing insights into the most effective techniques for improving the predictive performance of water quality classification. This study contributes to the broader efforts of ensuring safe water access and serves as a framework for employing machine learning in environmental quality assessments. The findings aim to assist researchers, policymakers, and public health officials in making informed decisions based on reliable machine learning predictions.
Original Article
View Cached Full Text

Cached at: 07/20/26, 09:31 AM

# AquaAugmentor: A Novel Feature Augmentation Algorithm for Water Potability Prediction
Source: [https://arxiv.org/html/2607.15775](https://arxiv.org/html/2607.15775)
Muntasir Tabasum1, Al Zadid Sultan Bin Habib2, Tanpia Tasnim3, Md\. Ekramul Islam4, Md Younus Ahamed5, Md Asif Bin Syed6

###### Abstract

Access to potable water is crucial for health, economic development, and sustainability\. However, accurately classifying water quality remains a significant challenge due to the complexity and variability of water source data\. This paper addresses the challenge of predicting water potability through machine learning and deep learning algorithms\. It introduces a novel feature augmentation algorithm, AquaAugmentor, to enhance the predictive performance of these models for low\-dimensional datasets\. Utilizing a dataset that includes chemical attributes of water, such as pH, hardness, solids, chloramines, sulfate, and others\. This study evaluates the performance of the models with and without AquaAugmentor\. Each model applied to classify water as potable or non\-potable and its performance is then evaluated and compared based on test accuracy and AUC score\. The results highlight the strengths and limitations of our proposed algorithm, providing insights into the most effective techniques for improving the predictive performance of water quality classification\. This study contributes to the broader efforts of ensuring safe water access and serves as a framework for employing machine learning in environmental quality assessments\. The findings aim to assist researchers, policymakers, and public health officials in making informed decisions based on reliable machine learning predictions\.

## IIntroduction

Access to clean and safe drinking water is a basic human right, yet it remains a major challenge around the world\. The World Health Organization \(WHO\) reports that more than two billion people live in water\-stressed countries, where access to safe water is either limited or entirely unavailable\[[18](https://arxiv.org/html/2607.15775#bib.bib14)\]\. Efficient testing and verification of water quality is crucial for both public health and environmental sustainability\. However, traditional water quality testing methods can be time\-consuming, often requiring specialized equipment and trained personnel\. As a result, there is increasing interest in using advanced technologies to streamline and improve the accuracy of water quality testing\. Innovative solutions, such as real\-time monitoring systems, biosensors, and machine learning algorithms, are being explored to provide rapid and precise assessments of water potability\. These advancements have the potential to greatly reduce the risks associated with contaminated water, ensuring safer drinking water for more people and supporting global efforts to achieve the United Nations’ Sustainable Development Goals \(SDGs\) related to clean water and sanitation\[[12](https://arxiv.org/html/2607.15775#bib.bib17),[11](https://arxiv.org/html/2607.15775#bib.bib18)\]\. Machine learning offers promising solutions to these challenges by enabling the rapid, cost\-effective, and accurate classification of water potability based on chemical and physical parameters\. Recent studies have demonstrated the potential of various machine learning algorithms in environmental monitoring and public health applications, including water quality assessment\[[5](https://arxiv.org/html/2607.15775#bib.bib15)\],\[[9](https://arxiv.org/html/2607.15775#bib.bib16)\]\. For instance, Support Vector Machines \(SVM\) and neural networks have been particularly noted for their predictive accuracy in complex environmental data scenarios\[[5](https://arxiv.org/html/2607.15775#bib.bib15)\]\. These models can process large datasets to identify patterns and correlations that might be missed by traditional methods, providing more reliable predictions about water quality\. Additionally, machine learning models can be continuously updated and improved with new data, ensuring they remain effective as environmental conditions and pollution sources change\. Integrating machine learning with sensor networks and real\-time data collection can develop more responsive and adaptive water quality monitoring systems\. These advancements hold great potential for enhancing public health protection, reducing the incidence of waterborne diseases, and supporting the sustainable management of water resources\[[20](https://arxiv.org/html/2607.15775#bib.bib19)\]\. This research explores and compares the effectiveness of various machine learning and deep learning models in classifying water as potable or non\-potable\. Each model’s performance was evaluated using test accuracy and AUC score to identify the most suitable approaches for real\-world applications\. Our novel contribution, AquaAugmentor, a feature augmentation algorithm, significantly enhances the predictive performance of these models, especially for low\-dimensional datasets\. This advancement highlights the potential of integrating advanced computational techniques to improve water quality assessment, contributing to broader environmental science efforts to ensure safe water access\. In this paper, section[II](https://arxiv.org/html/2607.15775#S2)highlights the literature review of related existing works\. Section[III](https://arxiv.org/html/2607.15775#S3)elaborates the methodological framework\. Section[IV](https://arxiv.org/html/2607.15775#S4)discusses the results, and section[V](https://arxiv.org/html/2607.15775#S5)concludes the paper with a blueprint of future works\.

## IIRelated Work

Access to potable water is a global concern that impacts health, sustainability, and economic development\. Recent studies have focused on evaluating and classifying water potability using machine learning or deep learning, exploring various models to predict water quality from chemical and physical parameters\. Mukati et al\. explored the effectiveness of ensemble methods like Random Forest, Decision Trees, and XGBoost, emphasizing the improved predictive accuracy these methods provide over traditional statistical techniques\[[13](https://arxiv.org/html/2607.15775#bib.bib1)\]\. This aligns with findings from De Luna et al\., who tested AdaBoost, XGBoost, and ExtraTree Classifier, finding that more complex ensemble models could sometimes significantly enhance prediction accuracy, particularly in handling non\-linear and complex dataset structures\[[6](https://arxiv.org/html/2607.15775#bib.bib2)\]\. Gao et al\. employed binomial Logistic Regression and K\-Nearest Neighbor \(KNN\) algorithms, demonstrating their suitability for smaller datasets and emphasizing the importance of each water quality feature’s independent influence on potability\[[7](https://arxiv.org/html/2607.15775#bib.bib3)\]\. SVM was particularly noted for its high accuracy in datasets with clear margin separations\. Khanna et al\. introduced a novel approach using deep learning for water potability classification, highlighting its capability to capture deeper insights from complex interdependencies among water quality parameters\[[10](https://arxiv.org/html/2607.15775#bib.bib4)\]\. Their findings suggest that deep learning could offer a significant step forward in predictive accuracy and reliability\. Yusuf et al\. demonstrated the effectiveness of Random Forest and Decision Trees in achieving higher classification accuracy, underscoring the significance of feature selection in model performance\[[3](https://arxiv.org/html/2607.15775#bib.bib5)\]\. Deep learning models, especially Artificial Neural Networks \(ANNs\) and Long Short\-Term Memory \(LSTM\) networks, have been applied to predict water quality with high precision, addressing complex non\-linear relationships within the data\. Such models were highlighted by Suleiman et al\. for their exceptional ability to classify groundwater potability, reflecting their growing application in environmental sciences\[[8](https://arxiv.org/html/2607.15775#bib.bib6)\]\. As Alipio discussed, integrating Internet of Things \(IoT\) technologies with machine learning models transforms water quality monitoring\. This approach leverages real\-time data acquisition and machine learning\-driven analytics to enhance the responsiveness and accuracy of water potability assessments\[[16](https://arxiv.org/html/2607.15775#bib.bib7)\]\. Comparative studies of machine learning algorithms reveal varying strengths across different techniques\. Haq et al\. compared several machine learning algorithms and found that models like XGBoost and ExtraTree classifiers provided superior performance due to their robust handling of diverse datasets\[[19](https://arxiv.org/html/2607.15775#bib.bib8)\]\. Argreen et al\.\[[4](https://arxiv.org/html/2607.15775#bib.bib9)\]explore the application of deep learning models for water potability classification in rural areas of the Philippines, highlighting the effectiveness of these models in enhancing water quality monitoring\. Ahmad Musleh\[[14](https://arxiv.org/html/2607.15775#bib.bib10)\]conducted a comprehensive comparative study of six machine learning algorithms for water potability classification, finding that Random Forest and J48 achieved the highest accuracy and precision in predicting water quality\. Alipio\[[2](https://arxiv.org/html/2607.15775#bib.bib11)\]developed a data\-driven IoT\-based system for real\-time water quality monitoring and potability classification for rural areas\. He matched his results with conventional laboratory tests and demonstrated high accuracy and minimal data transmission delays\. Patel et al\.\[[15](https://arxiv.org/html/2607.15775#bib.bib12)\]proposed a machine learning\-based water potability prediction model using the Synthetic Minority Over\-sampling Technique \(SMOTE\) and explainable AI techniques, achieving significant accuracy improvements and providing insights into the feature importance of water quality assessment\. Abuzir and Abuzir\[[1](https://arxiv.org/html/2607.15775#bib.bib13)\]demonstrated that Multi\-Layer Perceptron \(MLP\) outperforms J48 and Naïve Bayes in water quality classification, highlighting the significance of feature selection and dimensionality reduction using PCA for improving prediction accuracy\. These existing works incorporate a broad spectrum of current research, highlighting the dynamic nature of machine learning applications in environmental monitoring\. Integrating advanced computational techniques, such as deep learning and IoT\-enabled machine learning models, is noteworthy and offers promising avenues for enhancing water quality assessment\. However, these approaches are often dataset\-specific and vary significantly across different datasets\. The dimensionality of data can differ, posing challenges for achieving good accuracy with low\-dimensional datasets\. Our proposed AquaAugmentor algorithm addresses this challenge by leveraging the concept of feature augmentation, providing a robust solution for improving predictive performance with low\-dimensional datasets\. Our work advances sustainable water management with AI\-driven predictions, crucial for industry 5\.0\.

## IIIMethodology

### III\-AInput Dataset

The dataset\[[17](https://arxiv.org/html/2607.15775#bib.bib20)\]used in this study comprises 3,276 samples, each with nine chemical attributes of water: pH, Hardness, Solids, Chloramines, Sulfate, Conductivity, Organic Carbon, Trihalomethanes, and Turbidity\. These features are crucial for determining water potability\. The dataset includes potable and non\-potable classes, with 1,998 samples labeled as non\-potable and 1,278 samples labeled as potable\. Before proceeding to preprocessing and analysis, we performed mean imputation to address missing values, ensuring a complete and reliable dataset for accurate modeling\.

### III\-BFeature Augmentation

Feature augmentation is the process of enhancing the original feature set by generating new features through various methods, including polynomial combinations, statistical summarizations, and domain\-specific calculations\. Polynomial Features:It generates new features by taking polynomial combinations of existing features up to a specified degree\. Statistical Features:It adds features that summarize the data, such as mean, standard deviation, minimum, and maximum values for each sample\. Domain Specific Features:It creates new features based on domain knowledge, such as ratios and other relationships between existing features\. By combining these methods, feature augmentation significantly expands the feature space, potentially improving the model’s ability to capture complex patterns in the data\. In this study, we applied feature augmentation to enhance the original dataset\. This involved generating polynomial features, adding statistical summaries, and creating domain\-specific ratios\. This comprehensive approach aimed to provide the model with a richer and more informative feature set, thereby improving its performance\. Given a datasetDDwithnnsamples andmmfeatures, where the last column is the target variableyy:

𝐗=D\[:,:m\]\\mathbf\{X\}=D\[:,:m\]\(1\)𝐲=D​\[:,m\]\\mathbf\{y\}=D\[:,m\]\(2\)Where,𝐗\\mathbf\{X\}is the feature matrix of shape\(n,m\)\(n,m\)and𝐲\\mathbf\{y\}is the target vector of shape\(n,\)\(n,\)\. Given a feature matrix𝐗\\mathbf\{X\}of shape\(n,m\)\(n,m\), the polynomial feature matrix𝐗poly\\mathbf\{X\}\_\{\\text\{poly\}\}includes all polynomial combinations of the features up to a specified degreedd:

𝐗poly=Poly​\(𝐗,d\)\\mathbf\{X\}\_\{\\text\{poly\}\}=\\text\{Poly\}\(\\mathbf\{X\},d\)\(3\)WherePoly​\(𝐗,d\)\\text\{Poly\}\(\\mathbf\{X\},d\)generates a new matrix including all polynomial combinations of features in𝐗\\mathbf\{X\}up to degreeddand the shape of𝐗poly\\mathbf\{X\}\_\{\\text\{poly\}\}depends on the number of polynomial features generated\. For each sampleiiin the feature matrix𝐗\\mathbf\{X\}, calculate the mean \(μi\\mu\_\{i\}\), standard deviation \(σi\\sigma\_\{i\}\), minimum \(mini\\min\_\{i\}\), and maximum \(maxi\\max\_\{i\}\):

μi=1m​∑j=1mXi​j\\mu\_\{i\}=\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}X\_\{ij\}\(4\)σi=1m​∑j=1m\(Xi​j−μi\)2\\sigma\_\{i\}=\\sqrt\{\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\(X\_\{ij\}\-\\mu\_\{i\}\)^\{2\}\}\(5\)mini=minj=1m⁡Xi​j\\min\_\{i\}=\\min\_\{j=1\}^\{m\}X\_\{ij\}\(6\)maxi=maxj=1m⁡Xi​j\\max\_\{i\}=\\max\_\{j=1\}^\{m\}X\_\{ij\}\(7\)WhereXi​jX\_\{ij\}is the value of thejj\-th feature for theii\-th sample, andmmis the number of features\. To create new features based on domain knowledge by calculating specific ratios of the original features, we can assume that𝐗a,b\\mathbf\{X\}\_\{a,b\}denote the element\-wise ratio of columnsaaandbbof𝐗\\mathbf\{X\}:

Featurek=X​\[:,ak\]X​\[:,bk\]\\text\{Feature\}\_\{k\}=\\frac\{X\[:,a\_\{k\}\]\}\{X\[:,b\_\{k\}\]\}\(8\)Table[I](https://arxiv.org/html/2607.15775#S3.T1)shows how we can choose a set of index pairs\(ak,bk\)\(a\_\{k\},b\_\{k\}\)based on domain knowledge\.

TABLE I:Index pairs for domain\-specific features\.We need to combine the original polynomial features, statistical features, and domain\-specific features into a combined feature matrix𝐗comb\\mathbf\{X\}\_\{\\text\{comb\}\}:

𝐗comb=concat​\(𝐗poly,μ,σ,min,max,F1,F2,…,F14\)\\mathbf\{X\}\_\{\\text\{comb\}\}=\\text\{concat\}\(\\mathbf\{X\}\_\{\\text\{poly\}\},\\mu,\\sigma,\\min,\\max,\\text\{F\}\_\{1\},\\text\{F\}\_\{2\},\\ldots,\\text\{F\}\_\{14\}\)\(9\)Whereconcatdenotes the concatenation operation along the feature axis\.𝐗comb\\mathbf\{X\}\_\{\\text\{comb\}\}is the final expanded feature matrix including all derived features\.𝐗poly\\mathbf\{X\}\_\{\\text\{poly\}\}represents the polynomial features\.μ,σ,min,max\\mu,\\sigma,\\min,\\maxare the statistical features \(mean, standard deviation, minimum, and maximum\)\.F1,F2,…,F14\\text\{F\}\_\{1\},\\text\{F\}\_\{2\},\\ldots,\\text\{F\}\_\{14\}represents the domain\-specific features generated as ratios of the original featuresaka\_\{k\}andbkb\_\{k\}as described in Table[I](https://arxiv.org/html/2607.15775#S3.T1)\. Algorithm[1](https://arxiv.org/html/2607.15775#alg1)represents the pseudocode of the AquaAugmentor algorithm, Table[II](https://arxiv.org/html/2607.15775#S3.T2)analyzes the computational complexities, and Fig\.[1](https://arxiv.org/html/2607.15775#S3.F1)illustrates the workflow of the AquaAugmentor algorithm\.

Algorithm 1AquaAugmentor Algorithm0:Dataset

DDwith

nnsamples and

mmfeatures, target variable

yyin the last column

0:Combined feature matrix

XcombX\_\{\\text\{comb\}\}
1:Step 1: Separate Features and Target

2:

X=D\[:,:m\]X=D\[:,:m\],

y=D​\[:,m\]y=D\[:,m\]
3:Step 2: Create Polynomial Features

4:

p​o​l​y=PolynomialFeatures​\(degree=d,interaction\_only=True,include\_bias=False\)poly=\\text\{PolynomialFeatures\}\(\\text\{degree\}=d,\\newline \\text\{interaction\\\_only\}=\\text\{True\},\\text\{include\\\_bias\}=\\text\{False\}\)
5:

Xpoly=p​o​l​y\.f​i​t​\_​t​r​a​n​s​f​o​r​m​\(X\)X\_\{\\text\{poly\}\}=poly\.fit\\\_transform\(X\)
6:Step 3: Add Statistical Features

7:Initialize

s​t​a​t​\_​f​e​a​t​u​r​e​sstat\\\_features
8:foreach sample

iiin

XXdo

9:Calculate

μi\\mu\_\{i\},

σi\\sigma\_\{i\},

mini\\min\_\{i\},

maxi\\max\_\{i\}
10:Append

μi\\mu\_\{i\},

σi\\sigma\_\{i\},

mini\\min\_\{i\},

maxi\\max\_\{i\}to

s​t​a​t​\_​f​e​a​t​u​r​e​sstat\\\_features
11:endfor

12:Step 4: Add Domain\-Specific Features

13:

i​n​d​e​x​\_​p​a​i​r​s=\{\(3,5\),\(2,1\),\(5,2\),\(7,6\),\(3,6\),\(3,4\),\(0,5\),\(8,2\),\(7,4\),\(0,6\),\(4,1\),\(5,1\),\(8,3\),\(6,1\)\}index\\\_pairs=\\\{\(3,5\),\(2,1\),\(5,2\),\(7,6\),\(3,6\),\(3,4\),\\newline \(0,5\),\(8,2\),\(7,4\),\(0,6\),\(4,1\),\(5,1\),\(8,3\),\(6,1\)\\\}
14:Initialize

d​o​m​a​i​n​\_​f​e​a​t​u​r​e​sdomain\\\_features
15:foreach pair

\(ak,bk\)\(a\_\{k\},b\_\{k\}\)in

i​n​d​e​x​\_​p​a​i​r​sindex\\\_pairsdo

16:Calculate

f​e​a​t​u​r​ek=X​\[:,ak\]X​\[:,bk\]feature\_\{k\}=\\frac\{X\[:,a\_\{k\}\]\}\{X\[:,b\_\{k\}\]\}
17:Append

f​e​a​t​u​r​ekfeature\_\{k\}to

d​o​m​a​i​n​\_​f​e​a​t​u​r​e​sdomain\\\_features
18:endfor

19:Step 5: Combine All Features

20:

Xcomb=concat​\(Xpoly,μ,σ,min,max,feature1,feature2,…,feature14\)X\_\{\\text\{comb\}\}=\\text\{concat\}\(X\_\{\\text\{poly\}\},\\mu,\\sigma,\\min,\\max,\\text\{feature\}\_\{1\},\\newline \\text\{feature\}\_\{2\},\\ldots,\\text\{feature\}\_\{14\}\)
21:return

XcombX\_\{\\text\{comb\}\}

TABLE II:Computational complexity of the AquaAugmentor algorithm\.![Refer to caption](https://arxiv.org/html/2607.15775v1/feature_engineering_algorithm.png)Figure 1:Workflow of the AquaAugmentor algorithm\.
### III\-CData Balancing and Normalization

After applying the feature augmentation, the number of features increased from 9 to 62\. We then combined the target attribute with the augmented features and applied the SMOTE to balance the class distribution\. We split the dataset into training, validation, and test sets in a 70:15:15 ratio\. Finally, we normalized the data using StandardScaler to ensure that each feature contributed equally to the model’s performance\.

### III\-DPredictive Modeling

We employed a variety of machine learning and deep learning models to classify water potability\. The machine learning models used include Linear Regression, Random Forest, Linear SVM, KNN, MLP, CatBoost, AdaBoost, XGBoost, GBM, Naive Bayes, and Logistic Regression\. In addition, several deep learning models were implemented, such as 1\-D CNN, LSTM, Autoencoder, Variational Autoencoder, TabNet, TabTransformer, and a hybrid CNN\-LSTM model\. These models were selected to evaluate their effectiveness in accurately classifying water quality and to determine the impact of our novel feature augmentation algorithm, AquaAugmentor, on their predictive performance\.

## IVResults Analysis

The results shown in Table[III](https://arxiv.org/html/2607.15775#S4.T3)and the accompanying graphs highlight the significant influence of the AquaAugmentor algorithm on the effectiveness of several machine learning and deep learning models\. This approach improves accuracy and increases the AUC scores of most models\. The accuracy of the Random Forest model is significantly enhanced from 65\.71% to 72\.62% and its AUC score from 0\.68 to 0\.80\. This significantly enhances the model’s capacity to classify samples of drinkable and non\-drinkable water accurately\. Similarly, the XGBoost model also got a significant increase in accuracy, rising from 54\.27% to 71\.37%\. Additionally, its AUC score improved from 0\.55 to 0\.80, indicating a notable enhancement in prediction capability and reliability due to using AquaAugmentor\. Linear Regression, which typically faces challenges when dealing with complex datasets, also saw substantial advantages\. The model’s accuracy improved from 61\.22% to 63\.83%, while the AUC score significantly increased from 0\.50 to 0\.70, showing a substantial improvement in distinguishing between potable and non\-potable classes\. The accuracy of models such as LSTM and Autoencoder demonstrated significant enhancements\. LSTM’s accuracy increased from 57\.93% to 66\.67%, and its AUC improved from 0\.53 to 0\.71\. The Autoencoder, a model proficient in dealing with intricate patterns, improved its accuracy from 57\.93% to 67\.67% and its AUC from 0\.51 to 0\.80, showcasing AquaAugmentor’s capability to boost deep learning models\.

TABLE III:Results for different machine learning and deep learning models \(\* = with AquaAugmentor algorithm\)\.ModelAccuracyAUCAccuracy\*AUC\*Linear Regression61\.22%0\.5063\.83%0\.70Random Forest65\.71%0\.6872\.62%0\.80Linear SVM61\.02%0\.5166\.67%0\.73KNN64\.90%0\.6464\.90%0\.69MLP61\.02%0\.5769\.50%0\.77CatBoost56\.30%0\.5771\.29%0\.80AdaBoost54\.88%0\.5464\.44%0\.70XGBoost54\.27%0\.5571\.83%0\.80GBM56\.30%0\.5766\.94%0\.73Naive Bayes43\.70%0\.4755\.09%0\.58Logistic Regression50%0\.5464\.50%0\.701\-D CNN55\.66%0\.5570%0\.75LSTM57\.93%0\.5366\.67%0\.73Autoencoder57\.93%0\.5167\.67%0\.76Variational Autoencoder56\.30%0\.5366\.67%0\.70TabNet57\.93%0\.5365\.11%0\.71TabTransformer57\.93%0\.5064\.17%0\.71CNN\-LSTM56\.78%0\.5463\.17%0\.71The bar chart depicted in Fig\.[2](https://arxiv.org/html/2607.15775#S4.F2)clearly illustrates the enhancements in test accuracies among various models, both with and without the implementation of AquaAugmentor\. Improved models with AquaAugmentor routinely achieve better results than models without it, highlighting the algorithm’s ability to enhance model performance\. For instance, the increase in accuracy of the Random Forest model is noticeable, as is the improvement in the XGBoost and LSTM models\. This graphic offers a concise and easily understandable comparison of the efficacy of AquaAugmentor\. Fig\.[3](https://arxiv.org/html/2607.15775#S4.F3)illustrates a radar chart illustrating the AUC scores, providing a comprehensive perspective on the model’s performance across multiple dimensions\. This graphic showcases the extensive enhancements made to several models, with notable increases in the areas shown by the radar plot for models improved with AquaAugmentor\. Models such as CatBoost, XGBoost, and MLP demonstrate significant improvements in their AUC scores, highlighting the efficacy of AquaAugmentor in enhancing the model’s capacity to differentiate between classes\. The clustered heatmap in Fig\.[4](https://arxiv.org/html/2607.15775#S4.F4)illustrates the interconnections among the different characteristics following augmentation\. This heatmap visually represents the correlations between different features, providing insights into the relationships and interdependencies among the features after augmentation\. Specifically, certain characteristics have strong connections, creating separate groups essential for the enhanced performance observed in the models\. Furthermore, the t\-test findings illustrated in Fig\.[5](https://arxiv.org/html/2607.15775#S4.F5)offer a statistical confirmation of the relevance of the feature enhancements implemented by AquaAugmentor\. The bar chart of p\-values illustrates that numerous attributes have p\-values below the 0\.05 significance level, signifying that the disparities in these variables between potable and non\-potable water samples are statistically significant\. The statistical significance confirms the improvements in model performance and verifies the efficiency of AquaAugmentor in enhancing features\.

![Refer to caption](https://arxiv.org/html/2607.15775v1/bar_chart_accuracy_scores.png)Figure 2:Comparison of test accuracies for different models\.![Refer to caption](https://arxiv.org/html/2607.15775v1/radar_chart_auc_scores.png)Figure 3:Comparison of AUC scores for different models\.It is crucial to acknowledge that several research studies of different domains claim very high accuracies of 90% or higher, and these findings are frequently not reproducible in practical situations due to environmental data’s intricate and fluctuating nature\. Real\-world datasets, such as those employed for water potability classification, often contain noise and display substantial fluctuation, rendering the achievement of high accuracies impractical\. On the other hand, the enhancements accomplished by AquaAugmentor are practical and show tangible usefulness\. AquaAugmentor greatly improves models’ accuracy and AUC scores, increasing their robustness and reliability\. This makes them more appropriate for real\-world applications where maintaining safe drinking water is crucial\. When determining which models are appropriate, machine learning models such as Random Forest, XGBoost, and GBM benefit from AquaAugmentor since they help them properly manage enriched and sophisticated feature sets\. Deep learning models, such as LSTM and Autoencoders, provide significant enhancements, suggesting their appropriateness for intricate pattern identification and feature interactions\. Nevertheless, it is important to take into account the restrictions\. The process of feature augmentation using AquaAugmentor can result in higher computing complexity and longer training durations, especially when applied to deep learning models\. Furthermore, the method’s efficacy may differ based on the excellence and characteristics of the initial dataset, requiring meticulous preprocessing and validation to guarantee optimal performance\. Although AquaAugmentor has a few limitations, it is a valuable tool for improving the performance of models used in water quality categorization and other environmental data applications\.

![Refer to caption](https://arxiv.org/html/2607.15775v1/clustered_correlation_heatmap.png)Figure 4:Clustered correlation heatmap matrix after applying the AquaAugmentor algorithm\.![Refer to caption](https://arxiv.org/html/2607.15775v1/t_test_p_values_indexed.png)Figure 5:t\-Test p\-values for each feature after applying the AquaAugmentor algorithm\.
## VConclusions

In this study, AquaAugmentor algorithm represents remarkable progress in environmental data science, namely in the classification of water potability, a major concern in the public health sector\. This originality resides in its capacity to enhance low\-dimensional datasets by adding features obtained from polynomial expansions, statistical measurements, and domain\-specific information\. This greatly improves the predictive performance of machine learning and deep learning models\. The results exhibit consistent performance improvements across multiple models, highlighting AquaAugmentor’s capacity to enhance accuracy and dependability in water quality predictions\. The statistical validation using t\-tests validated the importance of these changes, which is consistent with the observed performance benefits and reinforces the practical usefulness of AquaAugmentor\. AquaAugmentor’s practical advancements showcase its potential for effectively ensuring safe drinking water, in contrast to the frequently unexplained high accuracies mentioned in many research works\. Further research could investigate adaptive feature augmentation methods to dynamically modify feature complexity according to dataset properties, reducing computing costs and training durations of the AquaAugmentor algorithm\. Incorporating AquaAugmentor with real\-time data collecting systems, such as IoT sensor networks, could facilitate ongoing monitoring and prompt analysis of water quality\. This would offer timely insights and interventions\.

## Acknowledgment

We would like to express our gratitude to the Center for Research Innovation and Transformation \(CRIT\) at Green University of Bangladesh for their generous financial support\.

## References

- \[1\]\(2022\)Machine Learning for Water Quality Classification\.Water Quality Research Journal57\(3\),pp\. 152–164\.Cited by:[§II](https://arxiv.org/html/2607.15775#S2.p1.1)\.
- \[2\]M\. I\. Alipio\(2020\)Data\-driven IoT\-based Water Quality Monitoring and Potability Classification System in Rural Areas\.In2020 International Conference on Information and Communication Technology Convergence \(ICTC\),pp\. 634–639\.Cited by:[§II](https://arxiv.org/html/2607.15775#S2.p1.1)\.
- \[3\]M\. I\. Alipio\(2020\)Towards Developing a Classification Model for Water Potability in Philippine Rural Areas\.ASEAN Engineering Journal10\(2\)\.Cited by:[§II](https://arxiv.org/html/2607.15775#S2.p1.1)\.
- \[4\]A\. Blancoet al\.\(2022\)Deep Learning Models for Water Potability Classification in Rural Areas in the Philippines\.In2022 IEEE World AI IoT Congress \(AIIoT\),pp\. 225–231\.Cited by:[§II](https://arxiv.org/html/2607.15775#S2.p1.1)\.
- \[5\]T\. S\. Cheonget al\.\(2022\)Application of Big Data, Deep Learning, Machine Learning, and Other Advanced Analytical Techniques in Environmental Economics and Policy\.Vol\.10,Frontiers Media SA\.Cited by:[§I](https://arxiv.org/html/2607.15775#S1.p1.1)\.
- \[6\]R\. G\. de Lunaet al\.\(2023\)A Comparative Study of Machine Learning Techniques for Water Potability Classification\.InTENCON 2023\-2023 IEEE Region 10 Conference \(TENCON\),pp\. 1345–1350\.Cited by:[§II](https://arxiv.org/html/2607.15775#S2.p1.1)\.
- \[7\]H\. Gaoet al\.\(2022\)Water Potability Analysis and Prediction\.Highlights in Science, Engineering and Technology16,pp\. 70–77\.Cited by:[§II](https://arxiv.org/html/2607.15775#S2.p1.1)\.
- \[8\]M\. I\. K\. Haqet al\.\(2021\)Efficiency of Machine Learning Algorithms in Water Quality Prediction\.Water Quality Research Journal\.Cited by:[§II](https://arxiv.org/html/2607.15775#S2.p1.1)\.
- \[9\]R\. A\. Johnson and D\. W\. Wichern\(2002\)Applied Multivariate Statistical Analysis\.Prentice Hall,Upper Saddle River, NJ\.Cited by:[§I](https://arxiv.org/html/2607.15775#S1.p1.1)\.
- \[10\]A\. Khannaet al\.\(2023\)Deep learning for water potability classification: a novel approach\.Int\. J\. Environ\. Res\. Public Health\.Cited by:[§II](https://arxiv.org/html/2607.15775#S2.p1.1)\.
- \[11\]M\. Khondoker, R\. Gurav, and S\. Hwang\(2024\)Utilization of Water Hyacinth Biomass as Eco\-Friendly Sorbent for Oil Spill Cleanup\.AQUA73\(2\),pp\. 183–199\.Cited by:[§I](https://arxiv.org/html/2607.15775#S1.p1.1)\.
- \[12\]M\. Khondokeret al\.\(2023\)Freshwater Shortage, Salinity Increase, and Global Food Production: A Need for Sustainable Irrigation Water Desalination—A Scoping Review\.Earth4\(2\),pp\. 223–240\.Cited by:[§I](https://arxiv.org/html/2607.15775#S1.p1.1)\.
- \[13\]A\. Mukatiet al\.\(2024\)Understanding the Concept of Water Potability through Machine Learning\.J Curr Trends Comp Sci Res3\(2\),pp\. 01–04\.Cited by:[§II](https://arxiv.org/html/2607.15775#S2.p1.1)\.
- \[14\]F\. Musleh\(2024\)A Comprehensive Comparative Study of Machine Learning Algorithms for Water Potability Classification\.Int\. J\. Comput\. Digit\. Syst\.15\(1\),pp\. 1189–1200\.Cited by:[§II](https://arxiv.org/html/2607.15775#S2.p1.1)\.
- \[15\]J\. Patelet al\.\(2022\)A Machine Learning\-Based Water Potability Prediction Model by Using Synthetic Minority Oversampling Technique and Explainable AI\.Computational Intelligence and Neuroscience2022\(1\),pp\. 9283293\.Cited by:[§II](https://arxiv.org/html/2607.15775#S2.p1.1)\.
- \[16\]A\. A\. Suleimanet al\.\(2023\)Comparative Analysis of Machine Learning and Deep Learning Models for Groundwater Potability Classification\.Engineering Proceedings56\(1\),pp\. 249\.Cited by:[§II](https://arxiv.org/html/2607.15775#S2.p1.1)\.
- \[17\]\(\)Water Quality\.Note:\\urlhttps://www\.kaggle\.com/datasets/adityakadiwal/water\-potability\(Accessed on 06/10/2024\)Cited by:[§III\-A](https://arxiv.org/html/2607.15775#S3.SS1.p1.1)\.
- \[18\]World Health Organization\(2019\)Drinking\-Water\.Note:Accessed: 2024\-06\-10External Links:[Link](https://www.who.int/news-room/fact-sheets/detail/drinking-water)Cited by:[§I](https://arxiv.org/html/2607.15775#S1.p1.1)\.
- \[19\]H\. Yusufet al\.\(2022\)Classification of Water Potability Using Machine Learning Algorithms\.In2022 International Conference on Data Analytics for Business and Industry \(ICDABI\),pp\. 454–458\.Cited by:[§II](https://arxiv.org/html/2607.15775#S2.p1.1)\.
- \[20\]M\. Zhuet al\.\(2022\)A Review of the Application of Machine Learning in Water Quality Evaluation\.Eco\-Environment & Health1\(2\),pp\. 107–116\.Cited by:[§I](https://arxiv.org/html/2607.15775#S1.p1.1)\.

Similar Articles