Beyond Accuracy: Centroid-Guided Contrastive Loss for Structured Fraudulent Job Posting Detection

arXiv cs.AI Papers

Summary

This paper proposes Centroid-Guided Contrastive Loss (CGCL) for structured fraudulent job posting detection, unifying classification and clustering in latent space to achieve state-of-the-art performance.

arXiv:2609.21599v1 Announce Type: new Abstract: Fraudulent job posting detection aims to identify job advertisements that are corrupted either through fake content, misleading information, or negative intent, disrupting the online eco-system of job-seekers and employers. Existing studies in this domain lack effective methods to simultaneously achieve high accuracy and meaningful structure of latent-space representations that capture subtleties among fake posts. To this end, we propose Centroid-Guided Contrastive Loss (CGCL), a loss function which unifies classification with densely formulated clustering to consistently reshape latent-space through a centroid-driven top-$k$ push-and-pull mechanism. The complementary nature of CGCL enables the model to enforce accurate decision boundaries and maintain high clustering compactness, effectively capturing both class separability and latent structure. Extensive experiments demonstrate the state-of-the-art (SOTA) performance of our method on EMSCAD, a public benchmark dataset. The code associated with this work is available at: https://github.com/ali-ahmed925/CGCL_code
Original Article
View Cached Full Text

Cached at: 09/21/26, 09:28 AM

# Beyond Accuracy: Centroid-Guided Contrastive Loss for Structured Fraudulent Job Posting Detection
Source: [https://arxiv.org/html/2609.21599](https://arxiv.org/html/2609.21599)
\{IEEEkeywords\}

Contrastive learning, Centroid\-guided loss, Representation learning, Word embeddings, GloVe, Word2Vec, TF\-IDF, Text classification, Clustering metrics, Fraudulent postings, Latent space structuring\.

\\IEEEspecialpapernotice

This work has been submitted to the IEEE for possible publication\. Copyright may be transferred without notice, after which this version may no longer be accessible\.

Syed Ali Ahmed1\{\}^\{\\textbf\{1\}\}Malaika Raza2\{\}^\{\\textbf\{2\}\}Affiliation:National University of Computer and Emerging Sciences, Karachi, 75020, PakistanMuhammad Shoaib Siddiqui3\{\}^\{\\textbf\{3\}\}\(Senior Member, IEEE\)Affiliation:Faculty of Computer and Information Systems, Islamic University of Madinah, Madinah 42351, Saudi Arabia and Muhammad Rafi4\{\}^\{\\textbf\{4\}\}\(Member, IEEE\)Affiliation:Department of AI & DS, National University of Computer and Emerging Sciences, Karachi, 75020, Pakistan

###### Abstract

Fraudulent job posting detection aims to identify job advertisements that are corrupted either through fake content, misleading information, or negative intent, disrupting the online eco\-system of job\-seekers and employers\. Existing studies in this domain lack effective methods to simultaneously achieve high accuracy and meaningful structure of latent\-space representations that capture subtleties among fake posts\. To this end, we propose Centroid\-Guided Contrastive Loss \(CGCL\), a loss function which unifies classification with densely formulated clustering to consistently reshape latent\-space through a centroid\-driven top\-kkpush\-and\-pull mechanism\. The complementary nature of CGCL enables the model to enforce accurate decision boundaries and maintain high clustering compactness, effectively capturing both class separability and latent structure\. Extensive experiments demonstrate the state\-of\-the\-art \(SOTA\) performance of our method on EMSCAD, a public benchmark dataset\. The code associated with this work is available at:[https://github\.com/ali\-ahmed925/CGCL\_code/tree/main](https://github.com/ali-ahmed925/CGCL_code/tree/main)

††corresponding:CORRESPONDING AUTHORS: SYED ALI AHMED \(email: k224058@nu\.edu\.pk\) and MUHAMMAD SHOAIB SIDDIQUI \(email: shoaib@iu\.edu\.sa\)\.††note:Muhammad Shoaib Siddiqui’s ORCID is 0000\-0002\-5656\-0416\. The authors extend their appreciation to the Deanship of Scientific Research, Islamic University of Madinah, Saudi Arabia, for funding this research work\.\\doi@font

Digital Object Identifier 10\.1109/\\@doiinfo

\\titlefont

\\authorfont

\\afffont\\@affil

\\@IEEEspecialpapernotice

\\corfont\\@corresp

\\afffont\\@authornote

\\abstractbox

\\keybox

## 1INTRODUCTION

Table 1:Comparison of fake job postings detection methods showing feature extraction techniques, imbalance handling approaches, and performance metrics\.\\IEEEPARstart

In the current age of Intelligence, where internet has deeply transformed our modern day life and social media platforms are straightforwardly accessible, companies nowadays, increasingly rely on electronic means to advertise their job postings and recruitment windows\. This digitized approach has not only simplified the application process for job seekers but also accelerated recruitment operations for employers\. However, malicious attempts to corrupt the recruitment ecosystem have emerged due to the prevalence of fake job postings surfacing on popular job\-hunting platforms\. These fraudulent postings are a direct attack on applicant’s personal information, exposing them to a range of cyber threats, such as identity theft, financial scams, and privacy breaches\[[9](https://arxiv.org/html/2609.21599#bib.bib1)\]\. According to a report published by theBetter Business Bureau, employment scams ranked as the second most risky scam type in 2023, with a 5\.2% increase in reported incidents and an average reported loss of $1,995 per victim, up from $1,500 in 2022\[[12](https://arxiv.org/html/2609.21599#bib.bib2)\]\. Therefore, detecting such fraudulent postings is of utmost importance to ensure compliance and integrity, and also to safeguard job\-seekers from financial and identity\-related harms\.

Machine learning in recent years has emerged as a highly promising technique across a wide range of domains, from healthcare\[[13](https://arxiv.org/html/2609.21599#bib.bib3),[14](https://arxiv.org/html/2609.21599#bib.bib4),[15](https://arxiv.org/html/2609.21599#bib.bib5)\]and finance\[[16](https://arxiv.org/html/2609.21599#bib.bib6),[17](https://arxiv.org/html/2609.21599#bib.bib7),[18](https://arxiv.org/html/2609.21599#bib.bib8),[19](https://arxiv.org/html/2609.21599#bib.bib9)\]to video surveillance\[[20](https://arxiv.org/html/2609.21599#bib.bib10),[21](https://arxiv.org/html/2609.21599#bib.bib11),[22](https://arxiv.org/html/2609.21599#bib.bib12)\]and cybersecurity\[[23](https://arxiv.org/html/2609.21599#bib.bib13),[24](https://arxiv.org/html/2609.21599#bib.bib14),[25](https://arxiv.org/html/2609.21599#bib.bib15),[26](https://arxiv.org/html/2609.21599#bib.bib16)\]\. A very impactful application of machine learning is in the field of Natural Language Processing \(NLP\), which is well\-suited for tasks involving unstructured textual data\. This makes it a viable approach for problems like fake job post detection where data is available in natural language format\. Several machine learning models have been employed to detect fraudulent samples ranging from traditional models such as Naive Bayes Classifier \(NBC\), K\-Neighbors Classifier \(KNNs\), Decision Tree Classifier \(DTC\) and more to advanced deep learning architectures like Long Short Term Memory \(LSTMs\), and Gated Recurrent Units \(GRUs\)\. These models capture temporal context and patterns from sequences of textual input, enabling them to classify anomalous samples as ”fraudulent postings”\.

While choosing the right model is undoubtedly a critical aspect of solving the problem, it is only one part of the broader learning scheme\. An often unnoticed and equally essential component of any learning framework is the loss or cost function\. A loss function directly influences the model to learn subtle patterns in the data by computing the difference between original target labels and predicted outputs, thereby optimizing the model’s parameters during training\. Most existing studies rely on standard loss functions like Cross\-entropy Loss or Hinge Loss, which while effective in many cases, may not be able to fully capture nuanced class boundaries and inter\-class separation in such complex tasks where normal and fraudulent samples’ features share subtle similarities\.

To this end, we propose a novelCentroid\-Guided Contrastive Loss \(CGCL\), which meaningfully reshapes the high\-dimensional latent embedding space by incorporating a centroid\-based push\-and\-pull mechanism that enhances intra\-class compactness and inter\-class separability\. We equip CGCL with a topkstrategy to select only thekfarthest in\-class samples during the pull phase and the closest cross\-class samples during the push phase, ensuring a focused and effective feature refinement\. Additionally, CGCL leverages a class\-balanced Cross\-Entropy Loss that guides the classifier towards more robust decision boundaries, maximizing the discrimination between unique classes for improved generalization and classification performance\. We train a Multilayer Perceptron with six hidden layers, each followed by a ReLU activation function— except for the final hidden layer which omits the activation\. We conducted extensive experiments on a publicly available dataset from Kaggle\[[27](https://arxiv.org/html/2609.21599#bib.bib17)\]to validate the effectiveness of our approach\. Our contributions can be summarized as follows:

1. 1\.We propose a novelCentroid\-Guided Contrastive Loss \(CGCL\)that ensures intra\-class compactness and inter\-class separability through topkpush\-pull mechanism\.
2. 2\.We integrate CGCL with class\-balanced Cross\-Entropy Loss, resulting in more robust and discriminative decision boundaries\.
3. 3\.We evaluate our complete framework on a benchmark dataset, demonstrating improved performance in class\-imbalance scenarios\.

Table 2:Job Dataset Schema: Column Specifications and Data Types![Refer to caption](https://arxiv.org/html/2609.21599v1/fraudulent_distribution.png)

![Refer to caption](https://arxiv.org/html/2609.21599v1/Geographic_distribution.png)

![Refer to caption](https://arxiv.org/html/2609.21599v1/Educational_Requirement_Analysis.png)

![Refer to caption](https://arxiv.org/html/2609.21599v1/Distribution_of_Fraudulent_vs_Genuine_by_Employment_Type.png)

Figure 1:Analysis of job postings: \(a\) Distribution of fraudulent vs\. genuine postings, \(b\) Geographic distribution of postings, \(c\) Educational requirement analysis, and \(d\) Distribution by employment type
## 2RELATED WORK

In this section, we review some of the novel works previously done by seasoned researchers that have contributed to advancing fake job post detection\. These studies will serve as the context for our current approach and highlight key trends, challenges, and gaps that our work aims to address\.

Researchers have extensively explored classical machine learning approaches for identifying fraudulent job listings\. Anitaet al\. in\[[2](https://arxiv.org/html/2609.21599#bib.bib18)\]employ several machine learning models to detect fake job posts\. The data cleaning process is specifically emphasized as a critical step in their pipeline\. Among the various models evaluated, Bidirectional LSTMs achieved the best results, reporting an accuracy of 98%\. Similarly, Anbarasuet al\.\[[10](https://arxiv.org/html/2609.21599#bib.bib19)\]trained multiple machine learning algorithms, including Naive Bayes Classifier \(NBC\), Decision Tree Classifier \(DTC\), Multilayer Perceptron \(MLP\), and Stochastic Gradient Descent \(SGD\), with SGD yielding the best results, achieving an impressive 98\.6% overall accuracy\. Another similar work evaluates a broad spectrum of data mining and classification algorithms for predicting fake job posts\. Singhet al\. in\[[7](https://arxiv.org/html/2609.21599#bib.bib28)\]developed a fraud detection model using existing ML models on the EMSCAD dataset, achieving 97\.4% accuracy in identifying fake job postings\. Their approach involved data preprocessing, feature selection, and ensemble classification\.

TF\-IDF is a widely used feature\-extraction technique that transforms text into weighted vectors based on token importance\. Keerthanaet al\.\[[3](https://arxiv.org/html/2609.21599#bib.bib20)\]applied a TF\-IDF vectorizer during preprocessing to encode tabular data, which was then fed into various ML models, with the MLP classifier achieving the highest accuracy of 71%\. Similarly, Duttaet al\.\[[1](https://arxiv.org/html/2609.21599#bib.bib27)\], after applying necessary data preprocessing steps, fed the extracted features to several ML models, with the Random Forest Classifier outperforming others in terms of accuracy\. In a novel work that explores ensemble\-based learning for fake job detection, Shiblyet al\.\[[4](https://arxiv.org/html/2609.21599#bib.bib23)\]leverage boosted decision trees and two\-class decision forest algorithms to detect fake job postings\. In the first algorithm, decision trees are arranged in an ensemble manner, where the errors of previous trees are corrected by subsequent ones before arriving at the final prediction\. Two class decision forests, on the other hand, rely on aggregated outcomes from grouped decision trees\.

Beyond classical methods, several studies have shifted towards deep learning architectures to better capture semantic patterns in job descriptions\. In a detailed approach presented by Pillaiet al\. in\[[8](https://arxiv.org/html/2609.21599#bib.bib21)\], they propose a training framework built on several BiLSTM layers\. First the numeric and textual data is converted into fixed\-size numerical vectors using separate tokenization layers, followed by an embedding layer, multiple BiLSTM layers, and a merging mechanism that fuses textual and numerical features for downstream classification\. Rathudiet al\. in\[[9](https://arxiv.org/html/2609.21599#bib.bib1)\]leverage bidirectional LSTMs with Word2Vec\[[28](https://arxiv.org/html/2609.21599#bib.bib22)\], a vectorizer that converts textual tokens into low dimensional vector representations based on their semantic context\. Coupling these two approaches allows the framework to capture both the semantic meaning of individual words and the sequential dependencies within the text, resulting in an overall accuracy of 97\.1%\.

Some studies have also emphasized the importance of data imbalance mitigation and feature selection\. Amaaret al\. in\[[5](https://arxiv.org/html/2609.21599#bib.bib24)\]employ both TF\-IDF vectorizer and BoW \(Bag of Words\) techniques for feature extraction\. The extracted features are then fed into six machine learning models to evaluate the best performing combination\. Additionally, they achieve over 99% accuracy by incorporating oversampling in their framework\. In a distinct study, Afzalet al\.\[[11](https://arxiv.org/html/2609.21599#bib.bib25)\]argue the reason that existing works often fell short is because feature selection and class imbalance are mostly overlooked\. They leverage Chi\-Square and PCA techniques to select the most relevant features, and apply SMOTE for minority class oversampling to effectively address the issue of class imbalance\. While the above studies rely primarily on traditional feature engineering methods, a recent work explored deep contextual representations, Qayyumet al\.\[[6](https://arxiv.org/html/2609.21599#bib.bib26)\]propose a novel feature extraction technique called Deep Contextualized Word Representation \(DCWR\), that employs a two\-layered BiLSTM to generate context\-aware encoded representation of words by modeling the likelihood of word sequences in both forward and backward directions\. Furthermore, they apply PCA as a feature reduction technique to select the principal components that capture maximum variance in the data\.

## 3METHODOLOGY

### 3\.1DATA COLLECTION

To carry out experiments on our proposed approach, we employed The Employment Scam Aegean Dataset \(EMSCAD\)\[[27](https://arxiv.org/html/2609.21599#bib.bib17)\], a publicly available benchmark dataset for detection of fraudulent job postings\. The dataset is highly imbalanced, comprising 17,014 genuine job advertisements and 866 fraudulent ones, collected between 2012 and 2014\. The tabular dataset contains both numerical and categorical columns, providing sufficient information to train complex neural networks capable of modeling intricate relationships between features\. Table[2](https://arxiv.org/html/2609.21599#S1.T2)presents description of each feature along with their datatypes\.

### 3\.2DATA ANALYSIS

Data Analysis is a crucial step in building a machine learning model as it offers comprehensive insights related to our underlying dataset and aids in determining the selection of an appropriate modeling workflow\. We first examined the distribution of target label \(Figure[1](https://arxiv.org/html/2609.21599#S1.F1)\) \(a\), which reveals a severe class imbalance, with legitimate postings significantly exceeding fraudulent ones\. The handling of such a problem is of utmost importance as it can bias the model towards underfitting and can lead to poor generalization over unseen examples of fraudulent postings\. The solution to this problem shall later be discussed in the paper\.

Additionally, analyzing geographical aspect of the dataset disclosed that majority of the advertisements came from cities like London, New York, and Athens that are considered the employment hubs of their respective countries as shown in Figure[1](https://arxiv.org/html/2609.21599#S1.F1)\(b\)\. This shows us that the distribution of postings is not random, but is sophistically linked to the regions that are employment\-concentrated\. We also investigated the educational qualifications required by the postings \(Figure[1](https://arxiv.org/html/2609.21599#S1.F1)\) \(c\), which ranged from school\-level or equivalent to bachelor’s and master’s degrees, with the majority of jobs demanding at least an undergraduate degree, reflecting the diversity in posted advertisements across employment sectors\.

Furthermore, we inspect what kind of employment types are more prone to get advertised in fraudulent postings, which revealed thatFull\-timepositions exhibit a higher share of fraudulent advertisements compared to other positions as shown in Figure[1](https://arxiv.org/html/2609.21599#S1.F1)\(d\)\. Finally, the data analysis provides a clear picture of all the challenges that need to be resolved before proceeding with effective preprocessing and model development\.

### 3\.3DATA PREPROCESSING

#### 3\.3\.1DATA CLEANING

We first cleaned our data through a preprocessing pipeline\. Data cleaning involves addressing inaccuracies, missing entries, duplicates, and inconsistently formatted data\. We began by handling null values in the data which posed a significant challenge as there were a total of 70183 void entries in the dataset\. However, this number is aggregated across all columns and therefore can be misleading\. To obtain a more accurate assessment, we only selected the columns that contained some amount of null values, and computed the average number of nulls per affected column which resulted in approximately 5848 null entries per column which still is a very concerning number\. To handle this, we removed the integer columns from our dataset, retaining only the categorical features that were sufficient for modeling complex patterns\. Following this, all the null entries were then replaced with empty strings\.

These empty strings individually seem to appear meaningless and insignificant, however, when concatenated with other categorical columns, we obtain a combined column named ”text”\. This methodology preserves the structure of the data, as it eliminates the need to remove rows with missing values, thereby maintaining the overall size and integrity of the dataset\.

Furthermore, we removed all the duplicated rows from our dataset to maintain consistency, and performed additional pre\-processing on the remaining textual data\. In order to do this, we designed a pipeline that carried out several cleaning tasks including converting text to lowercase, removing emails, URLs, and HTML tags, eliminating punctuation and numbers, removing stopwords, and finally lemmatizing the tokens using spaCy, ensuring that our data is suitable for feature engineering and machine learning modeling\.

#### 3\.3\.2FEATURE EXTRACTION

Once the dataset was structurally cleaned, we employed several feature extraction methods such asTF\-IDF Vectorizer, Word2Vec, andGloveto obtain feature embeddings\. Among these, the method that we finally adopted was Glove as it showed improved and consistent results with our clustering approach\. Since Glove is a well\-established embedding technique, that captures global context by estimating the co\-occurance probability between two words, we were able to obtainnthn^\{\\text\{th\}\}\-dimensional embeddings \(wheren=300n=300\) for each word in the vocabulary with minimal parameter tuning\. Therefore, we focus less on its implementation details here and instead present its comparative performance in Section[4](https://arxiv.org/html/2609.21599#S4)\.

Algorithm 1Centroid\-Guided Contrastive Loss \(CGCL\)Input:Embeddings

EE, Logits

ZZ, Labels

YY, hyperparameters:

kk,

α\\alpha,

β\\beta, margin

mm
Output:Loss

ℒ\\mathcal\{L\}
Step 1: Classification Loss \(Weighted CE\)

Compute class counts:

nc=count​\(Y=c\)n\_\{c\}=\\text\{count\}\(Y=c\)for each class

cc
Compute weights:

wc=1nc\+ϵw\_\{c\}=\\frac\{1\}\{n\_\{c\}\+\\epsilon\}
ℒC​E←−1N∑i=1Nwyilogp\(yi∣xi\)\\mathcal\{L\}\_\{CE\}\\leftarrow\-\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}w\_\{y\_\{i\}\}\\log p\(y\_\{i\}\\mid x\_\{i\}\)

Step 2: Normalize Embeddings

E←E‖E‖2E\\leftarrow\\frac\{E\}\{\\\|E\\\|\_\{2\}\}

Step 3: Contrastive Loss \(Centroid\-Guided\)

Initialize

ℒc​o​n​t​r​a​s​t​i​v​e←0\\mathcal\{L\}\_\{contrastive\}\\leftarrow 0,

C←0C\\leftarrow 0
foreach*classc∈unique​\(Y\)c\\in\\text\{unique\}\(Y\)*do

Ec←\{ei∈E∣yi=c\}E\_\{c\}\\leftarrow\\\{e\_\{i\}\\in E\\mid y\_\{i\}=c\\\}
E¬c←\{ei∈E∣yi≠c\}E\_\{\\neg c\}\\leftarrow\\\{e\_\{i\}\\in E\\mid y\_\{i\}\\neq c\\\}
if*\|Ec\|<k\+1\|E\_\{c\}\|<k\+1or\|E¬c\|<k\|E\_\{\\neg c\}\|<k*then

continue

Compute centroid:

μc=1\|Ec\|​∑e∈Ece\\mu\_\{c\}=\\frac\{1\}\{\|E\_\{c\}\|\}\\sum\_\{e\\in E\_\{c\}\}e
//Pull: hardest positives

d\+=\{‖e−μc‖2∣e∈Ec\}d^\{\+\}=\\\{\\\|e\-\\mu\_\{c\}\\\|\_\{2\}\\mid e\\in E\_\{c\}\\\}
Select top\-

kkfarthest:

Hk\+H^\{\+\}\_\{k\}
ℒp​u​l​l=1k​∑e∈Hk\+‖e−μc‖22\\mathcal\{L\}\_\{pull\}=\\frac\{1\}\{k\}\\sum\_\{e\\in H^\{\+\}\_\{k\}\}\\\|e\-\\mu\_\{c\}\\\|\_\{2\}^\{2\}
//Push: hardest negatives

d−=\{‖e−μc‖2∣e∈E¬c\}d^\{\-\}=\\\{\\\|e\-\\mu\_\{c\}\\\|\_\{2\}\\mid e\\in E\_\{\\neg c\}\\\}
Select top\-

kkclosest:

Hk−H^\{\-\}\_\{k\}
ℒp​u​s​h=1k​∑e∈Hk−max⁡\(0,m−‖e−μc‖2\)2\\mathcal\{L\}\_\{push\}=\\frac\{1\}\{k\}\\sum\_\{e\\in H^\{\-\}\_\{k\}\}\\max\(0,m\-\\\|e\-\\mu\_\{c\}\\\|\_\{2\}\)^\{2\}
//Combine push\-pull

ℒc​o​n​t​r​a​s​t​i​v​e←ℒc​o​n​t​r​a​s​t​i​v​e\+\(1−α\)​ℒp​u​l​l\+α​ℒp​u​s​h\\mathcal\{L\}\_\{contrastive\}\\leftarrow\\mathcal\{L\}\_\{contrastive\}\+\(1\-\\alpha\)\\mathcal\{L\}\_\{pull\}\+\\alpha\\mathcal\{L\}\_\{push\}
C←C\+1C\\leftarrow C\+1
if*C\>0C\>0*then

ℒc​o​n​t​r​a​s​t​i​v​e←ℒc​o​n​t​r​a​s​t​i​v​e/C\\mathcal\{L\}\_\{contrastive\}\\leftarrow\\mathcal\{L\}\_\{contrastive\}/C
Step 4: Unified Loss

ℒ←β​ℒC​E\+\(1−β\)​ℒc​o​n​t​r​a​s​t​i​v​e\\mathcal\{L\}\\leftarrow\\beta\\mathcal\{L\}\_\{CE\}\+\(1\-\\beta\)\\mathcal\{L\}\_\{contrastive\}

return*ℒ\\mathcal\{L\}*

![Refer to caption](https://arxiv.org/html/2609.21599v1/CGCL_V6.drawio.png)Figure 2:The pipeline consists of two phases: \(1\) Data Preprocessing & Augmentation, where raw text from the EMSCAD dataset is cleaned, vectorized using Word2Vec, and balanced using the ADASYN oversampling technique\. \(2\) Deep Metric Learning, featuring a hierarchical Feed\-Forward Neural Network \(Input to FC\-16\) optimized by the proposed Centroid\-based Geometric Contrastive Loss \(CGCL\)\. The CGCL formulation \(bottom center\) minimizes intra\-class variance by pulling samples toward their respective centroids \(C1, C2\) while maximizing inter\-class separability\.
#### 3\.3\.3IMBALANCE HANDLING

Imbalance handling is a necessary step when the dataset is largely inclined towards the distribution of a certain class samples\. Class imbalance hinders the generalization capability of machine learning models as they get biased towards majority class samples leading to the model overfitting\.

To handle class imbalance, we leveraged two complementary strategies, first, at the data level, we appliedADASYN \(Adaptive Synthetic Sampling\)to generate synthetic minority samples with a sampling ratio of 0\.4, enabling the model to capture intricate representations of minority class for reliable classification\. The choice of ADASYN overSMOTE \(Synthetic Minority Oversampling Technique\)was because ADASYN oversamples the minority class by focusing on its local context, thereby generating minority class samples that are harder to classify, whereas SMOTE produces uniformly distributed synthetic samples that may not sufficiently capture minority complexity\.

Second, at algorithmic level, we used a class\-weighted cross\-entropy loss, where we assigned weights to the class based on their frequencies\. The class having majority samples was given a lower weight while the one having lesser samples was allocated a higher weight, ensuring that misclassifications of the minority class were penalized more heavily during training\.

### 3\.4LOSS FORMULATION

In this section, we proposeCGCL \(Centroid\-Guided Contrastive Loss\), a custom loss function that unifies both classification and clustering within a single optimization objective\. While many existing works treat fraudulent postings detection as merely a classification task, we argue that relying solely onCross\-Entropyloss is suboptimal, as it does not modulate the structure of latent\-embedding space explicitly\. Additionally, leveraging only the classicContrastiveloss would be ineffective due to its uniform pairwise distance computations, which can dilute the focus on harder examples\. CGCL on the other hand, addresses these limitations by incorporating a classification loss that enforces discriminative decision boundaries and a centroid\-driven contrastive loss to reshape the embedding space, ensuring intra\-class compactness and inter\-class separability\.

#### 3\.4\.1WEIGHTED CROSS\-ENTROPY COMPONENT

To handle class\-imbalance, we adopt a weighted variant of the standard Cross\-Entropy loss\. We do this by assigning weights to the classes in inverse proportion to their frequency in the dataset, unlike the vanilla CE loss which treats all classes equally\. This approach gives more importance to the minority class while significantly reducing the dominance of majority classes\. Formally, the loss is given by:

ℒC​E=−1N∑i=1Nwyilog\(p\(yi∣xi\)\),\\mathcal\{L\}\_\{CE\}=\-\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}w\_\{y\_\{i\}\}\\,\\log\\big\(p\(y\_\{i\}\\mid x\_\{i\}\)\\big\),\(1\)
whereNNis the total number of samples,p⁡\(yi∣xi\)p\(y\_\{i\}\\mid x\_\{i\}\)is the predicted probability of the ground\-truth classyiy\_\{i\}for inputxix\_\{i\}, andwyiw\_\{y\_\{i\}\}denotes the class weight\.

#### 3\.4\.2CENTROID\-GUIDED PUSH AND PULL MECHANISM

This is the core of our proposed algorithm\. Taking inspiration from the well\-known Contrastive loss\[[29](https://arxiv.org/html/2609.21599#bib.bib29)\], we encourage embeddings of samples from the same class to be pulled closer together, while embeddings of samples from different classes are pushed apart in the latent space\. However, unlike traditional contrastive loss, where pairwise distances are computed exhaustively between samples, we design a novel centroid\-driven strategy where centroids for unique classes are computed dynamically to represent their respective clusters\. For every centroid, we identify the farthest top\-kksamples \(hard\-positives\) belonging to the same class as the active centroid, and pull them closer to the centroid \([2](https://arxiv.org/html/2609.21599#S3.E2)\), thereby reinforcing intra\-class compactness\. Similarly, we locate top\-kkfrom other classes \(hard\-negatives\) and push them away from the active centroid by a marginmm\([3](https://arxiv.org/html/2609.21599#S3.E3)\), which can be adjusted during training, ensuring stronger inter\-class separability\. We then combine this push\-pull mechanism in a unified framework, yielding a novel variant of contrastive loss, as formulated in \([4](https://arxiv.org/html/2609.21599#S3.E4)\)\.

ℒpull=1K​∑k=1K‖f⁡\(xkhard\)−μc‖22,\\mathcal\{L\}\_\{\\text\{pull\}\}=\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}\\left\\\|f\(x^\{\\text\{hard\}\}\_\{k\}\)\-\\mu\_\{c\}\\right\\\|^\{2\}\_\{2\},\(2\)
wherexkhardx^\{\\text\{hard\}\}\_\{k\}represents thekk\-th hardest positive sample \(farthest from the centroidμc\\mu\_\{c\}\)\.

ℒpush=1K​∑k=1Kmax⁡\(0,m−‖f⁡\(xkneg\)−μc‖2\)2,\\mathcal\{L\}\_\{\\text\{push\}\}=\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}\\max\\Big\(0,\\,m\-\\\|f\(x^\{\\text\{neg\}\}\_\{k\}\)\-\\mu\_\{c\}\\\|\_\{2\}\\Big\)^\{2\},\(3\)
wherexknegx^\{\\text\{neg\}\}\_\{k\}represents the hard negatives andmmis the margin parameter\.

ℒcontrastive=\(1−α\)⋅ℒpull\+α⋅ℒpush,\\mathcal\{L\}\_\{\\text\{contrastive\}\}=\(1\-\\alpha\)\\cdot\\mathcal\{L\}\_\{\\text\{pull\}\}\+\\alpha\\cdot\\mathcal\{L\}\_\{\\text\{push\}\},\(4\)
whereα∈\[0,1\]\\alpha\\in\[0,1\]balances the contribution of push loss with respect to the pull loss\.

#### 3\.4\.3UNIFIED LOSS FUNCTION

Finally, we integrate the class\-balanced Cross\-Entropy component \([1](https://arxiv.org/html/2609.21599#S3.E1)\) with the centroid\-guided contrastive objective \([4](https://arxiv.org/html/2609.21599#S3.E4)\) into a unified loss formulation\. This combined loss not only enforces robust class separation through discriminative decision boundaries but also explicitly reshapes the latent embedding space via the push\-pull mechanism\. Formally, the complete loss is given by:

ℒCGCL=β⋅ℒCE\+\(1−β\)⋅ℒcontrastive,\\mathcal\{L\}\_\{\\text\{CGCL\}\}=\\beta\\cdot\\mathcal\{L\}\_\{\\text\{CE\}\}\+\(1\-\\beta\)\\cdot\\mathcal\{L\}\_\{\\text\{contrastive\}\},\(5\)
whereβ∈\[0,1\]\\beta\\in\[0,1\]controls the balance between the classification term and the contrastive objective\. Moreover, all embeddings areℓ2\\ell\_\{2\}\-normalized to maintain consistent scale across samples\. For clarity, the step\-by\-step procedure of the proposed CGCL is summarized in Algorithm[1](https://arxiv.org/html/2609.21599#algorithm1)\.

### 3\.5MODEL ARCHITECTURE

For performing classification, we adopt a straightforward Multi\-layer Perceptron \(MLP\) architecture\. First the embeddings are obtained via Glove, which are then passed through the feed\-forward neural network\. The model consists of six fully connected layers, each followed by a ReLU activation function which enables the model to capture non\-linearity in the data\. These deep layers significantly reduce the dimensionality of input features from 300\-dimensional GloVe embeddings to a compact latent representation of size 16\. The resultant embeddings are then sent into a linear layer for classification, formally the model can be expressed as:

z=fembed​\(x\)z=f\_\{\\text\{embed\}\}\(x\)\(6\)
logits=Wcls​z\+bcls\\text\{logits\}=W\_\{\\text\{cls\}\}z\+b\_\{\\text\{cls\}\}\(7\)
zzis the latent representation,WclsW\_\{\\text\{cls\}\}are the classifier weights, andbclsb\_\{\\text\{cls\}\}is the classifier bias\.

Input Layer300\-dim GloVeWord EmbeddingsHidden Layer 1FC: 256 \+ ReLUFeature ExtractionHidden Layer 2FC: 128 \+ ReLUDimensionality ReductionHidden Layer 3FC: 64 \+ ReLUFeature CompressionHidden Layer 4FC: 40 \+ ReLUAbstraction LayerHidden Layer 5FC: 32 \+ ReLUSemantic EncodingEmbedding LayerFC: 16Latent RepresentationClassifier2 ClassesBinary Classificationℝ300\\mathbb\{R\}^\{300\}ℝ256\\mathbb\{R\}^\{256\}ℝ128\\mathbb\{R\}^\{128\}ℝ64\\mathbb\{R\}^\{64\}ℝ40\\mathbb\{R\}^\{40\}ℝ32\\mathbb\{R\}^\{32\}ℝ16\\mathbb\{R\}^\{16\}Figure 3:Enhanced MLP Embedder Architecture\. This diagram illustrates a seven\-layer fully connected neural network used for feature extraction and latent representation learning\.

## 4EXPERIMENTAL RESULTS

We validate our proposed approach on a benchmark dataset in comparison with well\-known studies previously done in the field of fraudulent posts detection\. We perform extensive experiments to justify the effectiveness of our method, including model\-level assessments through hyperparameter tuning and training configurations, and data\-level evaluations to analyze the influence of various feature extractors such as GloVe, Word2Vec, and TF–IDF\. on the performance of our model\.

Table 3:Performance of Word2Vec, GloVe, and TF\-IDF embeddings under identical hyperparameters\(α=0\.6,β=0\.9,k=3,m=1\.0\\textbf\{$\\alpha$\}=0\.6,\\textbf\{$\\beta$\}=0\.9,k=3,m=1\.0\)\. Values are reported asMacro / Micro\. Best results are inbold\.Table 4:Top hyperparameter configurations for GloVe and Word2Vec embeddings\.### 4\.1IMPLEMENTATION DETAILS

Every experiment was carried out in PyTorch on a system with 64GB of RAM and an NVIDIA RTX 4090 GPU\. We assessed three representations for feature extraction: TF\-IDF, Word2Vec, and GloVe \(300\-dimensional pretrained embeddings\)\. The proposed MLP embedder was trained using the Adam optimizer with an initial learning rate of1×10−31\\times 10^\{\-3\}, batch size of 128, and early stopping based on validation loss\. The training was carried out across various epochs such as 500, 1500, 2000 and 3000\. For our CGCL loss, we explored different settings of the hyperparametersα\\alpha,β\\beta,kk, andmm, while keeping all embeddingsℓ2\\ell\_\{2\}\-normalized\. The dataset was split into 80% training and 20% testing, with results reported on the held\-out test set\.

Table 5:Best performing models for each embedding method\.Listed are the optimal hyperparameters and corresponding performance scores\.Table 6:Clustering evaluation metrics for the final selected model \(GloVe,α=0\.5,β=0\.7,k=5\\alpha=0\.5,\\beta=0\.7,k=5\)\.![Refer to caption](https://arxiv.org/html/2609.21599v1/tsne.png)

![Refer to caption](https://arxiv.org/html/2609.21599v1/umap.png)

Figure 4:t\-SNE \(left\) reveals the continuous internal structure of fraud templates, while UMAP \(right\) confirms the global separability and distinct sub\-clustering of legitimate job categories\.
### 4\.2RESULTS

The first set of experiments evaluated different feature extraction techniques under identical hyperparameters across multiple epochs \(1500, 2000, and 3000 for Word2Vec and GloVe\)\. For the TF\-IDF vectorizer, experiments were conducted at 500 and 1000 epochs due to its faster convergence\. Beyond 1000 epochs, no significant performance gains were observed with TF\-IDF\. To ensure a fair comparison, the hyperparameters were kept fixed across all experimental settings\. Details are shown in Table[3](https://arxiv.org/html/2609.21599#S4.T3)\.

![Refer to caption](https://arxiv.org/html/2609.21599v1/evolution_grid_large_vertical_2.png)Figure 5:Evolution of Latent Space\. The progression from epoch 500 to 3000 shows the ’push’ mechanism creating a gradient of confidence for legitimate jobs \(blue\) while compacting fraudulent jobs \(red\) into a dense anomaly cluster\.
Hyperparameters: \(α=0\.5,β=0\.7,k=5\\alpha=0\.5,\\;\\beta=0\.7,\\;k=5\)Table 7:Test\-set metrics \(mean±\\pmstd, five seeds, fraud class\)\. Best per column inbold\.Furthermore, since our designed loss functionCGCLis highly sensitive to hyperparameter settings, in order to obtain the best performing model for each technique, we conducted several experiments with various combinations of parameters\. In total, 27 possible combinations were evaluated per technique\. For clarity and conciseness, we present only the top 5 results for GloVe and Word2Vec, summarized in structured form in Table[4](https://arxiv.org/html/2609.21599#S4.T4)\. Across these experiments, we observed thatα\\alphawithin the range of 0\.5\-1\.0 exhibited consistent results, whileα=1\.5\\alpha=1\.5consistently failed across both techniques\. Additionally, theβ\\betaparameter displayed technique\-specific behavior: The top results were obtained whenβ\\betawas set to 0\.5 on GloVe\-based embeddings, whereas Word2Vec performed better withβ=0\.3\\beta=0\.3, showing that the optimal setting ofβ\\betais embedding\-dependent\.

Table[5](https://arxiv.org/html/2609.21599#S4.T5)presents the best\-performing model for each technique, making it convenient to identify the GloVe\-based embedding model as the final selected model for our study\. It achieves an impressive99\.2%overall accuracy and an F1\-score of98\.7%on our test set\.

Finally, since our loss function is inherently incomplete without thecontrastivecomponent to enforce tight clustering in the latent space and thereby complement the classification objective, we therefore present the evaluations of our final model on established clustering metrics in Table[6](https://arxiv.org/html/2609.21599#S4.T6)\.

In conclusion, our proposed approach shows astounding classification performance with GloVe\-based model achieving an overall accuracy of99\.2%and an F1\-score of98\.7%\. In addition, the clustering performance metrics further validate our model’s robustness in a complex classification task, thereby proving its representational and discriminative efficacy\.

### 4\.3VISUALIZATIONS AND LATENT SPACE ANALYSIS

To prove the effectiveness and authority of CGCL and to understand how latent space evolves during training, we visualized the high\-dimensional embeddings \(dd= 16\) projected into 2D space\. These visualizations validate that the custom dual\-objective loss function successfully enforces both intraclass compactness and inter\-class separability\.

#### 4\.3\.1TEMPORAL EVOLUTION OF CLASS SEPARATION AND INTER\-CLASS COMPACTNESS

Figure[5](https://arxiv.org/html/2609.21599#S4.F5)illustrates the evolution of the latent space over 3000 training epochs\. The blue points represent Non\-Fraudulent \(Class 0\) job postings, while the red points represent Fraudulent \(Class 1\) postings\. At earlier epochs, there is substantial overlap between the two classes, particularly at epoch 500\. As training progresses, the embeddings gradually organize into two distinct regions, and by epoch 3000 the separation becomes considerably clearer\.

The legitimate job postings form a tailed distribution resembling a comet\-like shape\. Samples near the decision boundary share lexical and structural characteristics with fraudulent postings, whereas higher\-confidence legitimate examples are pushed farther away, forming the tail\. At the same time, the non\-fraudulent cluster becomes increasingly compact, indicating that the pull mechanism is successfully reducing intra\-class variation\. A similar trend is observed for fraudulent postings, which evolve from a scattered distribution into a tighter and more coherent cluster as the push mechanism separates them from the legitimate manifold\.

#### 4\.3\.2TOPOLOGICAL STRUCTURE \(T\-SNE AND UMAP\)

We further analyze the learned embeddings using t\-SNE and UMAP projections \(Figure[4](https://arxiv.org/html/2609.21599#S4.F4)\)\. The t\-SNE visualization reveals elongated and continuous structures within the fraudulent class, suggesting that fraudulent postings may vary along a spectrum of scam templates rather than forming a single homogeneous cluster\. Dense regions correspond to frequently occurring patterns, while the continuous trajectories indicate gradual transitions between related fraud types\.

UMAP produces a similar overall structure while providing clearer global separation between classes\[[30](https://arxiv.org/html/2609.21599#bib.bib30)\]\. Both legitimate and fraudulent postings exhibit internal sub\-structures, reflecting fine\-grained semantic variations within each class while maintaining a clear distinction between classes\. These observations complement the quantitative clustering results and provide visual evidence that CGCL promotes both intra\-class compactness and inter\-class separability in the learned latent space\.

Table 8:Per\-seed test\-set results \(F1 / PR\-AUC / MCC\)\. For each seed, best F1, best PR\-AUC, and best MCC across variants are individuallyunderlined\.

### 4\.4ABLATION STUDY

#### 4\.4\.1Introduction

Fraud detection suffers from severe class imbalance, which biases standard cross\-entropy classifiers toward the majority \(legitimate\) class and suppresses fraud recall\. This report evaluates five ablation variants on an identical MLP backbone, isolating the contribution of loss re\-weighting, synthetic oversampling \(ADASYN\), center loss, and the full CGCL model\. All variants share the same architecture, features, and optimizer; only the loss objective and sampling strategy differ\. Results are averaged over five random seeds\{7,13,21,42,99\}\\\{7,13,21,42,99\\\}\.

#### 4\.4\.2EXPERIMENTAL SETUP

##### Model and Training

Each variant uses a three\-hidden\-layer MLP \(256–128–64, ReLU \+ BatchNorm\) trained with Adam \(lr=10−3\\text\{lr\}=10^\{\-3\}, weight decay10−410^\{\-4\}\) for up to 500 epochs\.

##### Ablation Variants

- •CE\_only— Standard cross\-entropy; no imbalance handling\.
- •CE\_weighted— Cross\-entropy with inverse\-frequency class weights\.
- •CE\_ADASYN— Cross\-entropy on ADASYN\-oversampled training data\.
- •CE\_CenterLoss— Cross\-entropy \+ center loss \(λ=0\.5\\lambda=0\.5\) to compact intra\-class embeddings\.
- •CGCL\_full— Full proposed model: contrastive graph objective \+ cross\-entropy\.

#### 4\.4\.3RESULTS

##### Quantitative Summary

Table[7](https://arxiv.org/html/2609.21599#S4.T7)reports mean±\\pmstd of all metrics across five seeds\. CE\_only and CE\_weighted achieve∼\\sim97% accuracy yet score below F1 = 0\.75 on the fraud class—a result of majority\-class bias\. Every variant explicitly addressing imbalance exceeds F1 = 0\.97\.

![Refer to caption](https://arxiv.org/html/2609.21599v1/confusion_matrices.png)Figure 6:Confusion matrices for all five variants \(seed 7\)\. Top row \(L→\\toR\): CE\_only, CE\_weighted, CE\_ADASYN\. Bottom row: CE\_CenterLoss, CGCL\_full\.
##### Confusion Matrices \(Seed 7\)

Figure[6](https://arxiv.org/html/2609.21599#S4.F6)shows confusion matrices for seed 7\. CE\_only and CE\_weighted produce 56 and 51 false negatives respectively\. Enhanced variants improve sharply: CE\_ADASYN yields 13 false negatives, CE\_CenterLoss 5, and CGCL\_full only 2\.

#### 4\.4\.4DISCUSSION

##### Class imbalance as the primary bottleneck\.

The 26\-point F1 gap between CE\_only \(0\.718\) and CE\_ADASYN \(0\.981\) with no architectural change confirms that the MLP backbone is not the bottleneck\. Loss re\-weighting alone yields only marginal improvement \(\+0\.026 F1\), consistent with prior findings that synthetic oversampling is more effective than static weights for highly skewed distributions\.

##### Center loss stability\.

CE\_CenterLoss matches CE\_ADASYN in F1 and MCC while achieving the lowest seed variance across all metrics\. The center\-loss regularizer tightens the fraud\-class embedding cluster, reducing false negatives \(5 vs\. 13 for ADASYN on seed 7\) and stabilising the decision boundary across initializations\.

##### CGCL graph context and variance\.

CGCL\_full achieves the joint\-best Recall \(0\.997\) and highest PR\-AUC \(0\.981\), and virtually eliminates false negatives \(2 on seed 7\)\. However, its F1 variance \(±0\.020\) and MCC variance \(±0\.028\) are notably higher than CE\-based variants\. Seed 99 converged significantly more slowly, suggesting sensitivity to the interaction between graph topology and the contrastive temperature hyper\-parameter—an open problem for future work\.

## 5CONCLUSION

In this paper, we introduced a Centroid\-Guided Contrastive Loss \(CGCL\) loss function that is capable of enforcing accurate decision boundaries while consistently restructuring the latent space into dense and well\-separated clusters\. Our loss function dynamically updates unique clusters during training using centroid\-guided push and mechanism where top\-kkhard\-positives are pulled toward the active cluster while top\-kkhard\-negatives are pushed farther away\. We also integrate Cross\-Entropy loss to complement the contrastive objective, ensuring strong class discrimination alongside compact cluster formation\. Extensive experiments demonstrate that our approach achieves state\-of\-the\-art \(SOTA\) performance in both classification and clustering metrics, highlighting the effectiveness of CGCL as a unified loss function\.

## References

- \[1\]S\. Dutta and S\. K\. Bandyopadhyay\(2020\)Fake job recruitment detection using machine learning approach\.International Journal of Engineering Trends and Technology68\(4\),pp\. 48–53\.Cited by:[Table 1](https://arxiv.org/html/2609.21599#S1.T1.2.2.1),[§2](https://arxiv.org/html/2609.21599#S2.p3.1)\.
- \[2\]C\. Anita, P\. Nagarajan, G\. A\. Sairam, P\. Ganesh, and G\. Deepakkumar\(2021\)Fake job detection and analysis using machine learning and deep learning algorithms\.Revista Geintec\-Gestao Inovacao e Tecnologias11\(2\),pp\. 642–650\.Cited by:[Table 1](https://arxiv.org/html/2609.21599#S1.T1.2.3.1),[§2](https://arxiv.org/html/2609.21599#S2.p2.1)\.
- \[3\]B\. Keerthana, A\. R\. Reddy, and A\. Tiwari\(2021\)Accurate prediction of fake job offers using machine learning\.InMachine Intelligence and Soft Computing: Proceedings of ICMISC 2020,pp\. 101–112\.Cited by:[Table 1](https://arxiv.org/html/2609.21599#S1.T1.2.4.1),[§2](https://arxiv.org/html/2609.21599#S2.p3.1)\.
- \[4\]F\. Shibly, S\. Uzzal, and H\. Naleer\(2021\)Performance comparison of two class boosted decision tree snd two class decision forest algorithms in predicting fake job postings\.Association of Cell Biology Romania\.Cited by:[Table 1](https://arxiv.org/html/2609.21599#S1.T1.2.5.1),[§2](https://arxiv.org/html/2609.21599#S2.p3.1)\.
- \[5\]A\. Amaar, W\. Aljedaani, F\. Rustam, S\. Ullah, V\. Rupapara, and S\. Ludi\(2022\)Detection of fake job postings by utilizing machine learning and natural language processing approaches\.Neural Processing Letters54\(3\),pp\. 2219–2247\.Cited by:[Table 1](https://arxiv.org/html/2609.21599#S1.T1.2.6.1),[§2](https://arxiv.org/html/2609.21599#S2.p5.1)\.
- \[6\]H\. Qayyum, F\. Ali, M\. Nawaz, and T\. Nazir\(2023\)FRD\-lstm: a novel technique for fake reviews detection using dcwr with the bi\-lstm method\.Multimedia Tools and Applications82\(20\),pp\. 31505–31519\.Cited by:[Table 1](https://arxiv.org/html/2609.21599#S1.T1.2.7.1),[§2](https://arxiv.org/html/2609.21599#S2.p5.1)\.
- \[7\]V\. R\. Singh, P\. Sampras, A\. Dhage,et al\.\(2023\)Fake job post prediction using data mining\.Journal of Scientific Research and Technology,pp\. 39–47\.Cited by:[Table 1](https://arxiv.org/html/2609.21599#S1.T1.2.8.1),[§2](https://arxiv.org/html/2609.21599#S2.p2.1)\.
- \[8\]A\. S\. Pillai\(2023\)Detecting fake job postings using bidirectional lstm\.arXiv preprint arXiv:2304\.02019\.Cited by:[Table 1](https://arxiv.org/html/2609.21599#S1.T1.2.9.1),[§2](https://arxiv.org/html/2609.21599#S2.p4.1)\.
- \[9\]A\. D\. Rathudi\(2023\)Fake job post prediction\.Ph\.D\. Thesis,Dublin, National College of Ireland\.Cited by:[Table 1](https://arxiv.org/html/2609.21599#S1.T1.2.10.1),[§1](https://arxiv.org/html/2609.21599#S1.p1.2),[§2](https://arxiv.org/html/2609.21599#S2.p4.1)\.
- \[10\]V\. Anbarasu, S\. Selvakani, and M\. K\. Vasumathi\(2024\)Fake job prediction using machine learning\.ubiquity13\(1\),pp\. 06–14\.Cited by:[Table 1](https://arxiv.org/html/2609.21599#S1.T1.2.11.1),[§2](https://arxiv.org/html/2609.21599#S2.p2.1)\.
- \[11\]H\. Afzal, F\. Rustam, W\. Aljedaani, M\. A\. Siddique, S\. Ullah, and I\. Ashraf\(2024\)Identifying fake job posting using selective features and resampling techniques\.Multimedia Tools and Applications83\(6\),pp\. 15591–15615\.Cited by:[Table 1](https://arxiv.org/html/2609.21599#S1.T1.2.12.1),[§2](https://arxiv.org/html/2609.21599#S2.p5.1)\.
- \[12\]Better Business Bureau\(2022\)Employment scams risk report 2022\.Note:Accessed: 2025\-07\-26External Links:[Link](https://www.bbb.org/article/news-releases/26980-bbb-scam-tracker-risk-rsingeport)Cited by:[§1](https://arxiv.org/html/2609.21599#S1.p1.2)\.
- \[13\]H\. Habehh and S\. Gohel\(2021\)Machine learning in healthcare\.Current genomics22\(4\),pp\. 291–300\.Cited by:[§1](https://arxiv.org/html/2609.21599#S1.p2.1)\.
- \[14\]K\. Shailaja, B\. Seetharamulu, and M\. Jabbar\(2018\)Machine learning in healthcare: a review\.In2018 Second international conference on electronics, communication and aerospace technology \(ICECA\),pp\. 910–914\.Cited by:[§1](https://arxiv.org/html/2609.21599#S1.p2.1)\.
- \[15\]J\. Wiens and E\. S\. Shenoy\(2018\)Machine learning for healthcare: on the verge of a major shift in healthcare epidemiology\.Clinical infectious diseases66\(1\),pp\. 149–153\.Cited by:[§1](https://arxiv.org/html/2609.21599#S1.p2.1)\.
- \[16\]M\. F\. Dixon, I\. Halperin, P\. Bilokon,et al\.\(2020\)Machine learning in finance\.Vol\.1170,Springer\.Cited by:[§1](https://arxiv.org/html/2609.21599#S1.p2.1)\.
- \[17\]F\. Rundo, F\. Trenta, A\. L\. Di Stallo, and S\. Battiato\(2019\)Machine learning for quantitative finance applications: a survey\.Applied Sciences9\(24\),pp\. 5574\.Cited by:[§1](https://arxiv.org/html/2609.21599#S1.p2.1)\.
- \[18\]S\. Ahmed, M\. M\. Alshater, A\. El Ammari, and H\. Hammami\(2022\)Artificial intelligence and machine learning in finance: a bibliometric review\.Research in International Business and Finance61,pp\. 101646\.Cited by:[§1](https://arxiv.org/html/2609.21599#S1.p2.1)\.
- \[19\]B\. Kelly D\. Xiuet al\.\(2023\)Financial machine learning\.Foundations and Trends® in Finance13\(3\-4\),pp\. 205–363\.Cited by:[§1](https://arxiv.org/html/2609.21599#S1.p2.1)\.
- \[20\]W\. Sultani, C\. Chen, and M\. Shah\(2018\)Real\-world anomaly detection in surveillance videos\.In2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition,Vol\.,pp\. 6479–6488\.External Links:[Document](https://dx.doi.org/10.1109/CVPR.2018.00678)Cited by:[§1](https://arxiv.org/html/2609.21599#S1.p2.1)\.
- \[21\]R\. Chalapathy and S\. Chawla\(2019\)Deep learning for anomaly detection: a survey\.arXiv preprint arXiv:1901\.03407\.Cited by:[§1](https://arxiv.org/html/2609.21599#S1.p2.1)\.
- \[22\]G\. Pang, C\. Shen, L\. Cao, and A\. V\. D\. Hengel\(2021\)Deep learning for anomaly detection: a review\.ACM computing surveys \(CSUR\)54\(2\),pp\. 1–38\.Cited by:[§1](https://arxiv.org/html/2609.21599#S1.p2.1)\.
- \[23\]I\. H\. Sarker, A\. Kayes, S\. Badsha, H\. Alqahtani, P\. Watters, and A\. Ng\(2020\)Cybersecurity data science: an overview from machine learning perspective\.Journal of Big data7\(1\),pp\. 41\.Cited by:[§1](https://arxiv.org/html/2609.21599#S1.p2.1)\.
- \[24\]Y\. Xin, L\. Kong, Z\. Liu, Y\. Chen, Y\. Li, H\. Zhu, M\. Gao, H\. Hou, and C\. Wang\(2018\)Machine learning and deep learning methods for cybersecurity\.Ieee access6,pp\. 35365–35381\.Cited by:[§1](https://arxiv.org/html/2609.21599#S1.p2.1)\.
- \[25\]K\. Shaukat, S\. Luo, V\. Varadharajan, I\. A\. Hameed, and M\. Xu\(2020\)A survey on machine learning techniques for cyber security in the last decade\.IEEE access8,pp\. 222310–222354\.Cited by:[§1](https://arxiv.org/html/2609.21599#S1.p2.1)\.
- \[26\]G\. Apruzzese, P\. Laskov, E\. Montes de Oca, W\. Mallouli, L\. Brdalo Rapa, A\. V\. Grammatopoulos, and F\. Di Franco\(2023\)The role of machine learning in cybersecurity\.Digital Threats: Research and Practice4\(1\),pp\. 1–38\.Cited by:[§1](https://arxiv.org/html/2609.21599#S1.p2.1)\.
- \[27\]S\. Bansal\(2018\)Real or fake fake job posting prediction\.Kaggle\.Note:Accessed: 2025\-07\-27External Links:[Link](https://www.kaggle.com/datasets/shivamb/real-or-fake-fake-jobposting-prediction)Cited by:[§1](https://arxiv.org/html/2609.21599#S1.p4.1),[§3\.1](https://arxiv.org/html/2609.21599#S3.SS1.p1.1)\.
- \[28\]T\. Mikolov, K\. Chen, G\. Corrado, and J\. Dean\(2013\)Efficient estimation of word representations in vector space\.arXiv preprint arXiv:1301\.3781\.Cited by:[§2](https://arxiv.org/html/2609.21599#S2.p4.1)\.
- \[29\]P\. Khosla, P\. Teterwak, C\. Wang, A\. Sarna, Y\. Tian, P\. Isola, A\. Maschinot, C\. Liu, and D\. Krishnan\(2020\)Supervised contrastive learning\.Advances in neural information processing systems33,pp\. 18661–18673\.Cited by:[§3\.4\.2](https://arxiv.org/html/2609.21599#S3.SS4.SSS2.p1.1)\.
- \[30\]L\. McInnes, J\. Healy, and J\. Melville\(2018\)UMAP: uniform manifold approximation and projection for dimension reduction\. arxiv\.arXiv preprint arXiv:1802\.0342610\.Cited by:[§4\.3\.2](https://arxiv.org/html/2609.21599#S4.SS3.SSS2.p2.1)\.

Similar Articles

Graph-Based Financial Fraud Detection with Calibrated Risk Scoring and Structural Regularization

arXiv cs.LG

This paper proposes a graph neural network framework for financial fraud detection that integrates transaction records and identity information into node attributes, employs a multi-layer message passing mechanism, and uses weighted supervision and structural consistency regularization to improve risk scoring and probability calibration. Experiments on a public dataset show the method outperforms existing approaches.