TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection
Summary
The paper presents TeleAntiFraud 2.0, an audio-based benchmark for evaluating telecom fraud detection models using a mixed-tree generation pipeline and frozen monthly sets to address evolving fraud scripts and near-domain negatives.
View Cached Full Text
Cached at: 09/21/26, 07:19 AM
Paper page - TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection
Source: https://huggingface.co/papers/2609.18748 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Telecomfraudscriptsevolverapidlyandareoftendesignedtoresembleroutineserviceconversations,creatingtwokeyrequirementsforaudio-basedtelecom-fraudevaluation.First,benchmarksmustincorporatenewlyobservedscampatternswithoutoverwritingpreviouslyestablishedtestsets.Second,theymustdistinguishfraudfromlawful,near-domaincallsratherthanrelyingontopic-separatednegativeexamples.WepresentTeleAntiFraud2.0,constructedwithourMixed-TreeAnti-FraudGenerationPipelineandevaluatedunderamonthlyfrozenevaluationprotocol.Thepipelinetransformsonlinefraud-caseabstractsintoprofile-groundedscenarios,expandsthemthroughmixed-treegeneration,realizesfraudandnon-frauddialoguepathsundersharedcontexts,rendersvalidateddialoguesasrole-matchedspeech,andfreezestheresultingaudio,labels,prompts,manifests,andprovenancerecordsforeachmonthlyevaluationset.Eachfrozensetcontains900Chinesecalls,comprising600fraudand300near-domainnon-fraudcases.Controlledtextexperimentsshowthatthreeclassifiersachieveperfectmacro-averagedF1(Macro-F1)whenevaluatedagainstunrelatedorordinarynegatives,butdropto0.65-0.68withnear-domainsiblingnegatives.Full-setaudioandautomatic-speech-recognitionpluslarge-language-model(ASR+LLM)evaluationsfurtherrevealclass-priorshortcuts,predictioncollapse,andsnapshotsensitivity.Together,thesefindingsestablishnear-domainconstructionandcollapse-awarereportingascorerequirementsforevaluatingaudio-basedtelecom-fraudmodelsunderrealisticconfusableconditions.Theaccompanyingresearchartifactincludestheconstructioncode,evaluationscripts,manifests,anddocumentation.Ourdatasetandcodeareavailableathttps://anonymous.4open.science/r/TeleAntiFraud-2_0-EEB2/.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2609\.18748
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.18748 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.18748 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.18748 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Before the Warning Comes Too Late: Incremental Phone-Scam Detection from Speech
This paper introduces StreamFraudNet, a weakly supervised incremental model for detecting phone scams from raw speech, achieving a ROC-AUC of 0.9953 and operating in real-time to provide early warnings during calls.
Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams
This paper introduces the first public multimodal dataset of 100 Turkish scam and benign phone calls, evaluating seven LLMs under raw audio, ASR transcripts, and human-corrected transcripts. Results show transcript-based inputs outperform direct audio, highlighting the need for inclusive AI safety research in low-resource languages.
Dialogue to Detection: A Multimodal Hybrid NLP Pipeline for Insurance Fraud Detection
This paper presents a synthetic multimodal framework for insurance fraud detection at the first notice of loss (FNOL). It generates dialogue transcripts and two-speaker audio, combining ASR, NER, LLM-RAG, and speaker embeddings into a rule-based risk scoring system.
Application of Artificial Intelligence for Fraudulent Banking Operations Recognition
This paper studies the application of AI and machine learning algorithms for recognizing fraudulent banking transactions, proposing preprocessing techniques and comparing models. An artificial neural network and stacked generalization achieve improved AUC scores, with the best result around 0.954.
Anti-fraud tools can't keep pace with scammers exploiting cheap internet calling
Industry experts at a Broadband Breakfast panel discussed how anti-fraud tools are struggling to keep pace with scammers exploiting cheap internet calling and AI, emphasizing that authentication frameworks like STIR/SHAKEN must be combined with other defenses.