C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination
Summary
The paper introduces C-Score, a diagnostic framework for evaluating pseudo-label-based semi-supervised learning under open-world unlabeled contamination, demonstrating that clean accuracy is insufficient to detect hidden performance degradation.
View Cached Full Text
Cached at: 08/24/26, 04:32 AM
# C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination Source: [https://arxiv.org/abs/2608.20667](https://arxiv.org/abs/2608.20667) [View PDF](https://arxiv.org/pdf/2608.20667) > Abstract:Pseudo\-label\-based semi\-supervised learning has achieved strong performance due to its simplicity and scalability\. However, it is typically developed under a closed\-world assumption that unlabeled data are drawn from the same distribution as labeled data\. In practical deployment, unlabeled data are often collected from open environments and may contain OOD samples\. Under such contamination, OOD samples may still receive high\-confidence predictions and be incorporated into training as if they were valid target examples\. This creates an important evaluation problem: clean in\-distribution test accuracy may appear stable even when the internal learning dynamics of SSL have already deteriorated\. To address this issue, we study hidden collapse in pseudo\-label\-based SSL under open\-world unlabeled contamination from a diagnostic evaluation perspective\. We present C\-Score, a compact framework that evaluates training behavior in three complementary spaces: prediction, feature representation, and optimization\. C\-Score includes PLE and CCI for unlabeled prediction behavior, Sem\-Drift for deviation from labeled semantic anchors, and Grad\-Align for the compatibility between labeled and unlabeled optimization\. Experiments on CIFAR\-10 and CIFAR\-100 with multiple OOD sources, varying contamination ratios, and four pseudo\-label\-based SSL algorithms show that C\-Score metrics reveal hidden degradation that clean accuracy alone fails to detect: under SVHN contamination, CCI rises over 280% while best\-accuracy remains within 3% of the uncontaminated baseline; near\-OOD sources \(CIFAR\-100, STL\-10\) cause up to 14\.9% accuracy collapse \(FlexMatch, r=0\.5\)\. The results suggest that clean accuracy alone is insufficient for evaluating SSL robustness in open\-world environments, and that internal diagnostic signals are necessary for more reliable robustness assessment under unlabeled contamination\. ## Submission history From: TsaoLun Chen \[[view email](https://arxiv.org/show-email/8c8a041f/2608.20667)\] **\[v1\]**Fri, 21 Aug 2026 01:58:39 UTC \(693 KB\)
Similar Articles
Learning Risk Scores Robust to Unobserved Confounders
This paper proposes a method for learning risk scores from observational data that are robust to unobserved confounding, using sensitivity analysis and Wasserstein distributionally robust optimization. The approach improves calibration over traditional benchmarks and state-of-the-art methods.
Hard Cases, Bad Labels: Testing Error Exposure and Error Location in Uncertainty Sampling Under Bounded Label Noise
This study tests uncertainty sampling in active learning under bounded label noise, comparing error exposure and location effects across datasets to assess robustness and performance.
Accuracy and Order Sensitivity Diverge Under Label-Free Strategies
This paper investigates whether label-free strategies for multiple-choice benchmarks can remove option-order sensitivity in large language models, finding that neither two-stage prompting nor independent hypothesis scoring reliably improves accuracy.
Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards
This paper demonstrates that benchmark contamination inflates absolute scores for large language models but rarely reorders leaderboard rankings, using a paraphrase-controlled measure to show contamination is largely uniform across public models.
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
This paper introduces a framework for validating comparative LLM safety scoring without ground-truth labels, using an 'instrumental-validity chain' to establish deployment evidence. It demonstrates the method using a local-first tool called SimpleAudit on Norwegian safety packs and compares models like Borealis and Gemma 3.