Benchmarking Sensor Robustness in Plasma Diagnostic Models: A Systematic Evaluation on TokaMark
Summary
This paper introduces the first systematic robustness benchmark for plasma diagnostic machine learning models using the TokaMark dataset, evaluating architectures across six sensor failure scenarios and introducing a Robustness Score for cross-architecture comparison.
View Cached Full Text
Cached at: 07/20/26, 09:26 PM
Paper page - Benchmarking Sensor Robustness in Plasma Diagnostic Models: A Systematic Evaluation on TokaMark
Source: https://huggingface.co/papers/2607.11915
Abstract
Plasmadiagnosticmodelsfortokamakfusiondevicesarealmostuniversallyevaluatedonclean,completesensordata.Inpractice,fusiondiagnosticsfailregularly:acquisitionsystemsstartlate,individualsensorsdie,andsignaldropoutsclusterpreciselywhenaplasmadisruptionisapproaching.WepresentthefirstsystematicrobustnessbenchmarkforplasmadiagnosticMLusingtheTokaMarkdatasetof11,573MASTshots,evaluatingXGBoost,LSTM,Transformer,andtheTokaMarkCNNbaselineacrosssixphysically-groundedfailurescenariosandthreeimputationstrategies.WeintroducetheRobustnessScore(RS)forstandardizedcross-architecturecomparison.Ourcentralfindingisthatdisruption-proximatesensorfailure(corruptioninjectedinthefinalwindowtimesteps)collapsessequencemodelperformance(LSTM+212%NRMSE)whileastatisticalfeaturemodelremainscomparativelystable(XGBoost+37%).Forward-fillimputationeliminatesnearlyalldegradationfromrandomdropoutforsequencemodels(LSTM+57%to~0%),butofferslittlehelpwhentheendofthewindowiscorrupted.Shot-levelalarmevaluationusingground-truthdisruptiontimestampsrevealsthatLSTMalarmdetectioncollapsestoTPR=0.00underproximatesensorfailure,whilemean-fillimputationrecoversittoTPR=1.00,areversalofthepatternobservedinNRMSE.Plasmacurrentemergesasthesinglemostcriticaldiagnosticacrossallarchitectures(+73%to+140%uponremoval).Code,data,andtrainedcheckpointsareavailableathttps://github.com/Neerav-Gupta/tokamark-robustness.
View arXiv pageView PDFGitHub3Add to collection
Get this paper in your agent:
hf papers read 2607\.11915
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.11915 in a model README.md to link it from this page.
Datasets citing this paper1
#### Neerav-Gupta/tokamark-robustness-data
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.11915 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
TS-Fault: Benchmarking Time Series Forecasters Against Structural Faults
This paper introduces TS-Fault, a benchmark for evaluating time series forecasting models under structured fault scenarios like broken dependencies and regime changes, finding that clean-data accuracy often anti-correlates with robustness and that foundation models are especially fragile.
Benchmarking Machine Learning Uncertainty Quantification Methodologies for Predicting Turbine Gas Temperature Degradation
This paper benchmarks five uncertainty quantification methods for neural network predictions of turbine gas temperature, evaluating trade-offs in coverage, width, and stability to guide prognostics and health management in engines.
RobustMAD: Evaluating Real-World Robustness of Multimodal Small Language Models for Deployable Anomaly Detection Assistants
RobustMAD introduces a benchmark to evaluate the real-world robustness of multimodal small language models for deployable industrial anomaly detection assistants. It reveals critical failure modes and provides guidance for next-generation systems.
Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment
The paper benchmarks five large language models on multi-sensor physical hazard assessment, revealing that all tested models fail to produce precautionary warnings when multiple sensors are simultaneously elevated below their individual safety limits, while achieving near-perfect accuracy on single-sensor violations.
Unsteady Metrics and Benchmarking Cultures of AI Model Builders
This paper introduces Benchmarking-Cultures-25, a dataset analyzing how AI model builders selectively highlight benchmarks in press releases. It finds a fragmented evaluation landscape with limited cross-model comparability, arguing that benchmarks are used as narrative devices for market positioning rather than standardized scientific measurement.