Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

arXiv cs.AI Papers

Summary

This paper evaluates automated safety benchmarks for small language models, finding high ambiguity in judgments that compromises reliability and reveals a capability-safety confound.

arXiv:2608.17183v1 Announce Type: new Abstract: Small Language Models (SLMs) are increasingly deployed in resource-constrained, privacy-sensitive settings, where safety and bias failures can cause security and societal risks. However, existing AI safety\slash security\slash compliance benchmarks are designed for large language models that may not transfer reliably to SLMs. We therefore ask: Can these benchmarks effectively and reliably evaluate SLMs? To answer this question, we conduct a large-scale assessment of the effectiveness and robustness of these automated pipelines by evaluating five widely used benchmark suites across 26 open-source SLMs under a unified judging rubric, which assigns a score of 0, 1, or 0.5 to harmful, safe, or ambiguous/irrelevant responses, respectively. Across the benchmarks, ambiguous judgments dominate and correlate with prompt complexity and model architecture, indicating that {\em LLM-centric safety benchmarks are insufficient as standalone evidence for SLM safety assessment}. In general, the ambiguity rate increases with lexical density, output perplexity, and output length and decreases with lexical sophistication, self-coherence, and reply-prompt similarity. This reveals a capability-safety confound that mixes model capability with apparent safety. Since ambiguity is prevalent, aggregate mean-score leaderboards are mathematically brittle: model rankings change significantly under reasonable ambiguity treatments, even when the underlying outputs remain unchanged.
Original Article
View Cached Full Text

Cached at: 08/19/26, 09:54 AM

# Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models
Source: [https://arxiv.org/html/2608.17183](https://arxiv.org/html/2608.17183)
Fengjun LiAffiliation:E\-mail[\{nyam, fli, bluo\}@ku\.edu](mailto:{nyam,%20fli,%20bluo}@ku.edu)Bo Luo

###### Abstract

Small Language Models \(SLMs\) are increasingly deployed in resource\-constrained, privacy\-sensitive settings, where safety and bias failures can cause security and societal risks\. However, existing AI safety/security/compliance benchmarks are designed for large language models that may not transfer reliably to SLMs\. We therefore ask: Can these benchmarks effectively and reliably evaluate SLMs? To answer this question, we conduct a large\-scale assessment of the effectiveness and robustness of these automated pipelines by evaluating five widely used benchmark suites across 26 open\-source SLMs under a unified judging rubric, which assigns a score of 0, 1, or 0\.5 to harmful, safe, or ambiguous/irrelevant responses, respectively\. Across the benchmarks, ambiguous judgments dominate and correlate with prompt complexity and model architecture, indicating thatLLM\-centric safety benchmarks are insufficient as standalone evidence for SLM safety assessment\. In general, the ambiguity rate increases with lexical density, output perplexity, and output length, and decreases with lexical sophistication, self\-coherence, and reply\-prompt similarity\. This reveals a capability\-safety confound that mixes model capability with apparent safety\. Since ambiguity is prevalent, aggregate mean\-score leaderboards are mathematically brittle: model rankings change significantly under reasonable ambiguity treatments, even when the underlying outputs remain unchanged\.

††footnotetext:This is the author’s accepted version of a paper accepted for publication at ESORICS 2026\. The final authenticated version will be available online at Springer’s Lecture Notes in Computer Science \(LNCS\) series once published\.## 1Introduction

Small language models \(SLMs\), with hundreds of millions to a few billion parameters, have emerged as a distinct choice in edge, IoT, and other resource\-constrained settings with strict latency, cost, and privacy/security compliance limitations\. Many sub\-10B SLMs continue to be released in the open\-source community and increasingly adopted in commercial products\[[50](https://arxiv.org/html/2608.17183#bib.bib1),[39](https://arxiv.org/html/2608.17183#bib.bib5),[53](https://arxiv.org/html/2608.17183#bib.bib3),[14](https://arxiv.org/html/2608.17183#bib.bib4)\]\.

In practice, model safety is evaluated using established benchmark pipelines \(e\.g\.,\[[24](https://arxiv.org/html/2608.17183#bib.bib16)\]\) originally developed for larger, more fluent LLMs\. These automated judges typically compress open\-ended LLM outputs into a single aggregate score derived from ternary labels \(“harmful”, “ambiguous”, or “safe”\)\. However, SLM outputs differ systematically from their larger counterparts, as SLMs often produce shorter, less fluent, and more failure\-prone outputs\[[45](https://arxiv.org/html/2608.17183#bib.bib17)\]\. Prior work also shows that the automated judges in LLM safety benchmarks are highly sensitive to surface features such as length and fluency\[[58](https://arxiv.org/html/2608.17183#bib.bib19)\]\. This raises a critical question: do existing LLM safety benchmarks remain effective, reliable, and decision\-useful for SLMs, or do they merely generate “ambiguous” labels that reflect evaluation difficulty rather than true safety behavior?

Instead of treating benchmark scores as ground truth safety measures, we examine the automated evaluation pipeline itself\. In particular, we study the*capability\-safety confound*and ask*whether this confound is empirically significant for SLMs under standard LLM\-oriented benchmarks\.*We conduct a large\-scale analysis of benchmark prompts, model\-generated responses, judge outputs, and scoring to assess whether LLM\-oriented benchmarks can effectively and reliably assess the safety, security, and compliance properties of SLMs and to identify the root causes of the observed ineffectiveness\. In particular, we aim to answer five research questions: \[RQ1\] What do current safety benchmarks indicate about SLM safety, and how consistent are rankings across model\-benchmark pairs? \[RQ2\] Are benchmark outputs interpretable and decision\-useful? \[RQ3\] How sensitive are aggregate safety rankings, and how much do SLM rankings shift under alternative treatments of ambiguity? \[RQ4\] What factors predict an “ambiguous” judge decision, and does this ambiguity merely reflect model capability \(e\.g\., output quality or prompt difficulty\) rather than safety? And \[RQ5\] what quality\-aware reporting strategies and metric adjustments can mitigate this confound to yield more robust and reliable safety evaluations for SLMs?

To answer these questions, we conduct a large\-scale evaluation of five benchmark suites across 26 SLMs\. We find that ambiguous outcomes concentrate on more complex prompts and lower\-quality generations, which can make mean\-score rankings brittle\. Taken together, our findings support a defensible conclusion without additional human labeling: even when the “true” safety of ambiguous cases is unknown, the automated pipeline is*demonstrably biased*by capability\-related surface features, yielding aggregate rankings that are unstable under reasonable scoring choices\.

The main contributions of this paper is summarized as follows:

- ∙\\bulletWe present a large\-scale measurement study of automated LLM\-safety evaluation pipelines by executing 741,312 model\-prompt evaluations across five benchmark suites over 26 SLMs, with 715,312 judge\-scored safety evaluations and 26,000 BBQ bias evaluations\[[46](https://arxiv.org/html/2608.17183#bib.bib7)\]\.
- ∙\\bulletWe show that ambiguous outcomes are strongly associated with output quality, prompt complexity, and model architecture, which indicates the presence of a capability\-safety confound in automated judging\.
- ∙\\bulletWe identify conditions under which automated pipelines are decision\-useful\. In particular, ambiguity\-heavy suites produce unstable conclusions, while simpler prompt sets may mask these issues\.
- ∙\\bulletWe show that mean\-score leaderboards from LLM benchmarks are unreliable for SLMs, with model orderings shifting substantially under reasonable ambiguity\-handling choices\.
- ∙\\bulletWe provide actionable recommendations to improve the robustness of safety/security evaluation for SLMs\.

The rest of the paper is organized as follows: Section 2 reviews background and related work\. Section 3 describes our methodology and experimental setup\. Section 4 reports our findings and answers RQ1–RQ4, Section 5 discusses implications and recommendations for RQ5, and Section 6 concludes the paper\.

## 2Background and Related Work

### 2\.1Small language models in the LLM era

Small language models \(SLMs\) are not simply “pre\-LLM” model\. Early transformers such as GPT\-2 already enabled general\-purpose generation at a sub\-billion scale\[[41](https://arxiv.org/html/2608.17183#bib.bib36)\]\. In the LLM era, SLMs have evolved into a distinct deployment choice: vendors and open\-source communities continue to release new models with various parameter counts, e\.g\., sub\-1B and near\-1B, which trade peak capability for lower latency, lower cost, and easier on\-device deployment\[[17](https://arxiv.org/html/2608.17183#bib.bib42),[1](https://arxiv.org/html/2608.17183#bib.bib49)\]\. Recent surveys highlight their growing use due to accessibility, ease of fine\-tuning, and permissive licensing, which enable them to be well\-suited for privacy\-sensitive and resource\-constrained applications\[[50](https://arxiv.org/html/2608.17183#bib.bib1),[39](https://arxiv.org/html/2608.17183#bib.bib5),[53](https://arxiv.org/html/2608.17183#bib.bib3)\]\.

The technical trajectory of SLMs also differs from simply “scaling down” transformers\. Advances in architecture, efficiency, distillation, and instruction tuning aim to preserve decision\-useful behavior under tight compute budgets, and recent surveys increasingly position SLMs as components within larger systems \(e\.g\., proxy models, guard models, or collaborators to larger models\) rather than standalone assistants\[[54](https://arxiv.org/html/2608.17183#bib.bib9),[7](https://arxiv.org/html/2608.17183#bib.bib10),[8](https://arxiv.org/html/2608.17183#bib.bib11),[53](https://arxiv.org/html/2608.17183#bib.bib3)\]\. These trends make benchmark\-based evaluation appealing, but they also require ensuring that such benchmarks remain valid when outputs are shorter, less fluent, or more failure\-prone\.

### 2\.2Safety\-security benchmarks and AI governance

Safety and security benchmarks emerged in response to concrete failure modes observed in LLMs, including harmful instruction following, jailbreaks and adversarial prompting, privacy leakage, and biased or stereotyped responses\[[55](https://arxiv.org/html/2608.17183#bib.bib28),[60](https://arxiv.org/html/2608.17183#bib.bib29),[26](https://arxiv.org/html/2608.17183#bib.bib13)\]\. As these harms become operationally relevant, evaluation has evolved from ad hoc red\-team examples to reusable prompt suites, risk taxonomies, and standardized scoring protocols\. Modern benchmarks typically combine targeted prompts \(to elicit safety\-, security\-, privacy\-, or bias\-relevant behaviors\) with scoring methods that map open\-ended outputs into comparable outcomes\[[23](https://arxiv.org/html/2608.17183#bib.bib12),[32](https://arxiv.org/html/2608.17183#bib.bib8),[57](https://arxiv.org/html/2608.17183#bib.bib14),[27](https://arxiv.org/html/2608.17183#bib.bib6)\]\.

Meanwhile, regulation and governance have increased the demand for measurable safety evidence\. TheEU AI Actmandates risk management and requirements for accuracy, robustness, and security for high\-risk AI systems\[[13](https://arxiv.org/html/2608.17183#bib.bib30)\]\. NIST’sAI Risk Management Frameworkemphasizes measurement and evaluation as part of trustworthy AI development\[[38](https://arxiv.org/html/2608.17183#bib.bib31)\]\. China’sInterim Measures for Generative AI Servicesimpose requirements on data use, content safety, privacy, and transparency for gen\-AI services\[[9](https://arxiv.org/html/2608.17183#bib.bib32)\]\. While these frameworks do not mandate a specific benchmark, they raise the stakes for benchmark validity: if safety evaluations are used for compliance or procurement decisions, their scores must accurately reflect the intended safety properties\.

### 2\.3Evaluation frameworks and judge\-based scoring

HELM Safety v1\.0\[[24](https://arxiv.org/html/2608.17183#bib.bib16)\]is a safety evaluation framework that standardizes benchmark suites and automated judging configurations to improve comparability across models and prompts, while SALAD\-Bench\[[27](https://arxiv.org/html/2608.17183#bib.bib6)\]organizes evaluation around a broader safety taxonomy with fine\-grained category coverage\. More broadly, automated evaluation frameworks differ in both benchmark content and scoring design: some use scalar absolute scoring for single responses\[[58](https://arxiv.org/html/2608.17183#bib.bib19),[30](https://arxiv.org/html/2608.17183#bib.bib22)\], some aggregate pairwise preferences into win rates or rankings\[[28](https://arxiv.org/html/2608.17183#bib.bib23),[29](https://arxiv.org/html/2608.17183#bib.bib24)\], and others rely on objective ground\-truth or verifiable checks when possible\[[59](https://arxiv.org/html/2608.17183#bib.bib27),[56](https://arxiv.org/html/2608.17183#bib.bib25),[40](https://arxiv.org/html/2608.17183#bib.bib26)\]\.

HELM\-style safety evaluation occupies a distinct point in this design space\. It applies judge\-based scoring to open\-ended safety prompts and maps outputs into a coarse ternary scale for aggregation\[[24](https://arxiv.org/html/2608.17183#bib.bib16)\]\. While practical and operationally useful, prior work shows that LLM\-as\-a\-judge decisions are sensitive to fluency, verbosity, formatting, and rubric phrasing\[[58](https://arxiv.org/html/2608.17183#bib.bib19),[30](https://arxiv.org/html/2608.17183#bib.bib22)\]\. These sensitivities are especially consequential for SLMs, whose outputs are often shorter, less robust, and hard\-to\-interpret under complex prompts\[[6](https://arxiv.org/html/2608.17183#bib.bib18),[45](https://arxiv.org/html/2608.17183#bib.bib17)\]\. As a result, ternary labels may reflect not only safety\-relevant uncertainty but also evaluation difficulty, hence making the benchmark outputs potentially ambiguous and unreliable\[[25](https://arxiv.org/html/2608.17183#bib.bib21),[12](https://arxiv.org/html/2608.17183#bib.bib20)\]\.

## 3Methodology and Measurement Design

We evaluate the*automated safety evaluation pipeline*at the per\-instance level, pairing each prompt with each target SLM\. We first run 26 SLMs across safety and bias benchmarks, then extract and analyze prompt\-, response\-, and model\-level covariates\. Finally, we examine ambiguity, ranking stability, and metadata effects using these aligned records\.

### 3\.1SLMs, benchmark suites, and settings

We evaluate 26 SLMs spanning 124M to 4B parameters across different model families\. There is no universal parameter\-count cutoff for SLMs, as the term SLM is typically used relative to frontier LLMs and often refers to models designed for lower latency, lower cost, local inference, or resource\-constrained deployment\[[50](https://arxiv.org/html/2608.17183#bib.bib1)\]\. We selected SLMs up to 4B parameters, so that: \(1\) we can include the relatively more powerful variants of the small models, \(2\) we can cover a broader set of model families, and \(3\) we still keep the study focused on resource\-constrained deployment rather than mid\-size LLMs\. The set includes both base and instruction\-tuned models and spans multiple tokenizer and architecture configurations to support model metadata analyses\. Appendix Table[A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2)summarizes the models used in our analyses\.

We evaluate the selected models on five well\-adopted LLM benchmark suites\.

- ∙\\bulletAirBench 2024\[[57](https://arxiv.org/html/2608.17183#bib.bib14)\]is a regulation\- and policy\-aligned safety benchmark derived from government regulations and company policies, with 5,694 prompts spanning 314 granular risk categories\. Its multi\-clause, policy\-framed prompts can increase ambiguity for less fluent or underspecified SLM generations\.
- ∙\\bulletSALAD\-Bench\[[27](https://arxiv.org/html/2608.17183#bib.bib6)\]evaluates safety across a broader taxonomy of unsafe and policy\-sensitive behaviors, emphasizing fine\-grained category coverage, with 21,318 prompts in our evaluation set\. Its diverse, often multi\-constraint prompts are useful for testing whether judge\-based scoring remains stable when SLM outputs are short, partial, or otherwise difficult to interpret\.
- ∙\\bulletHarmBench\[[32](https://arxiv.org/html/2608.17183#bib.bib8)\]is a standardized framework for automated red teaming that measures harmful responses and refusal behaviors under unsafe or adversarial prompts\. In HELM Safety v1\.0, the HarmBench suite corresponds to 400 behavior prompts spanning HarmBench’s four behavior classes; results should be interpreted as applying to this HELM subset and scoring protocol\[[24](https://arxiv.org/html/2608.17183#bib.bib16)\]\.
- ∙\\bulletSimple Safety Tests\[[52](https://arxiv.org/html/2608.17183#bib.bib15)\]\(SST\) provide 100 short, obviously unsafe prompts intended as a quick “lower\-bound” safety check\. We use the same judge\-label setup as HarmBench\[[24](https://arxiv.org/html/2608.17183#bib.bib16)\]and compare ambiguities across model classes\.
- ∙\\bulletBBQ\[[46](https://arxiv.org/html/2608.17183#bib.bib7)\]is a multiple\-choice benchmark designed to measure social bias in question answering by contrasting under\-informative contexts \(ambiguous\) with informative contexts across demographic dimensions\.

These five widely used benchmarks cover most of the risk themes that appear in broader benchmark catalogs: policy and regulation compliance, jailbreak\-oriented safety scenarios, harmful\-response behavior, and social bias\. Because this paper is an*evaluation*of automated judging and aggregation, we treat AirBench and SALAD\-Bench as our primary case studies: their multi\-clause, policy\-relevant prompts provide the high complexity that most strongly exposes ambiguity and pipeline failures\. HarmBench serves as a red\-teaming comparison with high ambiguity under unsafe prompts, Simple Safety Tests provides a shorter lower\-complexity contrast, and BBQ anchors the analysis with an accuracy\-based bias benchmark rather than the ternary safety rubric\.

Experimental setup\.We run AirBench, HarmBench, SST, and BBQ with HELM Safety v1\.0\[[24](https://arxiv.org/html/2608.17183#bib.bib16)\]and SALAD\-Bench with a custom pipeline\[[27](https://arxiv.org/html/2608.17183#bib.bib6)\], executing all 26 SLMs locally with greedy decoding \(temperature00\)\. Open\-ended suites use up to 512 output tokens; BBQ uses single\-token multiple\-choice decoding; SALAD uses max\_new\_tokens=256=256, seed13371337, and chat templates when present\. AirBench, HarmBench, and SST are judged by GPT\-4o \(openai/gpt\-4o\-2024\-05\-13\) \(category rubrics for AirBench\)\. Unless stated otherwise, cross\-benchmark analyses use these GPT labels, with Llama 3\.1 405B Instruct Turbo scores retained for robustness on a 16\-model HarmBench, Simple Safety overlap\. SALAD uses MD\-Judge\-v0\.1 in unsafe\-probability mode\.

### 3\.2Judging and scoring rubric

HELM encodes outcomes as numeric scores\{0,0\.5,1\}\\\{0,0\.5,1\\\}to enable simple aggregation by averaging within and across suites, where 0\.5 denotes an intermediate outcome between a harmful \(0\) and safe response \(1\)\. While the benchmarks produce 0\.5 \(ambiguous response\) when the target model neither returns a harmful result nor explicitly rejects the prompt, the exact interpretation of 0\.5 slightly varies by benchmark: AirBench defines it as ambiguous or partial outputs\[[57](https://arxiv.org/html/2608.17183#bib.bib14)\], HarmBench defines it as noncompliance \(to the prompt\) without an explicit refusal\[[32](https://arxiv.org/html/2608.17183#bib.bib8)\], and Simple Safety Tests use it for potentially \(but not decisively\) unsafe responses\. SALAD\-Bench is different in its original form, as it uses an multi\-dimensional \(MD\) judge evaluator that produces an unsafe probability together with a binary safe/unsafe judgment, instead of a native ternary rubric\[[27](https://arxiv.org/html/2608.17183#bib.bib6)\]\. To make SALAD comparable with the HELM\-style suites in our joint analysis, we inspect the empirical score distributions and use Otsu\-based thresholding to reinterpret SALAD as a ternary outcome, with model\-specific lower and upper bounds mapping low unsafe\-probability cases to safe \(1\), high\-probability cases to harmful \(0\), and the middle region to ambiguous \(0\.5\)\.

In practice, the “Ambiguous” \(0\.5\) label may be considered the semantically correct label for a response\. For example, when the target SLM generates a meaningless gibberish or totally irrelevant response for a prompt, the logically correct label for the judge is “ambiguous”, indicating that the SLM output is neither harmful nor safe \(explicit rejection\)\. However, such semantically correct results are not helpful from a safety/securitybenchmarkingperspective, as they merely indicate that the target SLM is incapable of handling the benchmarking prompts instead of being safe or unsafe, i\.e\., the benchmark is ineffective\. Meanwhile, when the proportion of “ambiguous” \(0\.5\) labels is small, the average score from all the prompts could effectively indicate the ratio of “safe” and “harmful” responses from the target SLM\. However, with a significant portion of “ambiguous” labels, the aggregated score also becomes less meaningful\.

### 3\.3Measurement metrics and analyses

We organize our measurements into three clusters that correspond to the main sources of signal and confounding in automated safety evaluation:*harmfulness and safety outcomes*,*prompt complexity*, and*output quality*\. The first cluster captures the benchmark outcome being reported, while the latter two capture properties that prior NLP and readability work frequently uses to describe text difficulty, fluency, semantic relatedness, and lexical style\[[5](https://arxiv.org/html/2608.17183#bib.bib33),[31](https://arxiv.org/html/2608.17183#bib.bib34),[47](https://arxiv.org/html/2608.17183#bib.bib35)\]\. These established, automatically computable metrics let us interpret all model\-prompt pairs at scale\. Complex prompts may be harder for SLMs to follow, while low\-quality or off\-topic replies may be harder for automated judges to classify reliably\.

∙\\bullet~Harmfulness and safety outcomes\.For each model\-prompt instance in the safety benchmarking suites, the judge assigns a scoresi∈\{0,0\.5,1\}s\_\{i\}\\in\\\{0,0\.5,1\\\}\. We define three metrics, the harmful\-completion rate \(HCR\), the safe\-refusal rate \(SRR\), and the ambiguity rate \(AR\), to capture the fraction of harmful, safe, and ambiguous responses, respectively\. GivenNNevaluated prompts, they are computed as:

HCR=1N∑i=1N𝕀\[si=0\],SRR=1N∑i=1N𝕀\[si=1\],AR=1N∑i=1N𝕀\[si=0\.5\]\.\\mathrm\{HCR\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbb\{I\}\[s\_\{i\}=0\],\\quad\\mathrm\{SRR\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbb\{I\}\[s\_\{i\}=1\],\\quad\\mathrm\{AR\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbb\{I\}\[s\_\{i\}=0\.5\]\.
Our meta\-evaluation focuses onAR\\mathrm\{AR\}and on how model comparisons vary under reasonable alternative treatments of ambiguous outcomes\.

∙\\bullet~Prompt complexity metrics\.We adopt six per\-prompt measures to capture multiple dimensions of input difficulty rather than relying on a single metric\.Token countmeasures prompt length in GPT\-2 tokens\. TheFlesch\-Kincaid gradeandGunning\-Fog indexare common readability metrics to estimate reading level based on sentence and word structure, which are recently used in analyses of LLM\-facing and LLM\-generated text\[[5](https://arxiv.org/html/2608.17183#bib.bib33),[31](https://arxiv.org/html/2608.17183#bib.bib34)\]\.Per\-token perplexity, computed via GPT\-2 cross\-entropy, measures linguistic surprisal, with higher values indicating less predictable wording\[[6](https://arxiv.org/html/2608.17183#bib.bib18)\]\.Instruction\-verb countcaptures the density of imperative or task\-directing verbs, approximating the number of explicit actions requested\. Finally,average dependency distancemeasures the mean token\-to\-head arc length from a dependency parse, reflecting syntactic complexity\.

∙\\bullet~Output quality metrics\.We compute six per\-reply measures to capture whether a generation is sufficiently long, fluent, semantically aligned, and lexically interpretable for reliable judging\.Output token countmeasures reply length, whileoutput perplexity, computed via GPT\-2 per\-token cross\-entropy, captures fluency, with lower values indicating more fluent or typical text\[[6](https://arxiv.org/html/2608.17183#bib.bib18)\]\.Reply–prompt similarity, cosine similarity between prompt and reply sentence embeddings, measures semantic alignment, whileself\-coherence, mean adjacent\-sentence embedding similarity, captures internal consistency\. Both use sentence\-transformer embeddings\[[47](https://arxiv.org/html/2608.17183#bib.bib35)\]\. Finally,lexical sophistication\(rarer vocabulary usage\) andlexical density\(proportion of content words\) characterize lexical style and information packaging rather than safety directly\.

Metric\-ambiguity correlations\.To identify which properties are associated with ambiguity, we aggregate each benchmark prompt across evaluated SLMs and compute a prompt\-levelAR\\mathrm\{AR\}as the fraction of models receiving a score of 0\.5\. We then compute Pearson correlations between this prompt\-levelAR\\mathrm\{AR\}and each prompt and output metric, retaining only associations that remain significant after Benjamini–Hochberg FDR correction \(q<0\.05q<0\.05\)\.

![Refer to caption](https://arxiv.org/html/2608.17183v1/figures/all_benchmarks_model_rankings.png)Figure 1:Safety rankings for 26 SLMs in 14 model families across 5 benchmark suites: higher mean scores indicate the models exhibit safer behavior\.

## 4Evaluations and Findings

### 4\.1Part I: Automated evaluation outcomes and interpretability

#### Observed behavior across benchmark suites and model families

We first summarize safety evaluation results of 26 SLMs across five benchmark suites\. Figure[1](https://arxiv.org/html/2608.17183#S3.F1)reports per\-suite model safety rankings based on aggregate benchmark scores, providing a coarse view of relative safety and compliance behavior under each suite\. For BBQ, we report multiple\-choice accuracy\.

First, we find that model rankings vary across suites\. For example, Gemma 2\-2B IT performs well on HarmBench, BBQ, and AirBench, while Simple Safety is led by Llama 3\.2 Instruct\. Performance also does not follow a consistent trend with model size\. It varies by benchmark and model family rather than increasing monotonically\. In some cases, larger models perform better, e\.g\., BBQ accuracy increases with model size in the Qwen 2\.5 family, while in others, smaller models outperform larger ones, e\.g\., GPT2 Small and Medium exceed larger GPT2 variants on AirBench\. These inconsistencies motivate our subsequent analysis of whether these benchmarks provide decision\-useful signals\.

In our experiments, BBQ shows a more consistent increase in performance with model size\. It uses accuracy\-based scoring rather than judge\-assigned labels, where higher accuracy corresponds to lower bias under its design\[[46](https://arxiv.org/html/2608.17183#bib.bib7)\]\. However, accuracy is also sensitive to format compliance\. For example, DeepSeek\-R1\-Qwen\-1\.5B frequently produces free text \(e\.g\., “Okay”\) instead of an option label \(A/B/C\), leading to unmapped predictions and near\-zero accuracy\.

Answer to RQ1:Across suites, SLM evaluation results depend strongly on the suite and model family and do not follow a consistent monotonic trend with parameter count\. Rankings vary substantially across AirBench, HarmBench, and SALAD\-Bench, but are more stable for Simple Safety Tests and BBQ, partly due to shorter and simpler prompts\. Overall, rankings are not directly comparable across suites: a model that performs well on one benchmark may rank much lower on another\.

#### Decision\-usefulness of safety benchmarks for SLMs

We consider an automated safety benchmarking pipeline as decision\-useful for SLM evaluation if these two basic conditions, effectiveness and consistency, hold:

- ∙\\bulletEffectiveness\.The evaluation pipeline is supposed to generate labels that effectively reflect harmful completions versus safe refusal, while the proportion of “ambiguous” labels should be small\. That is, when a benchmark cannot decisively identify whether a model output is safe or harmful, the benchmark is ineffective and not decision\-useful\.
- ∙\\bulletConsistency\.For the same or similar safety and security aspects, the scores and relative rankings should be consistent across benchmark suites and model classes\. That is, when two benchmarks provide inconsistent or even conflicting scores/rankings across models, such results are not decision\-useful to the users\.

![Refer to caption](https://arxiv.org/html/2608.17183v1/figures/all_benchmarks_score_and_proportions_per_model_row.png)\(a\)SLMs: per\-benchmark score proportions per model across AirBench, HarmBench\.
![Refer to caption](https://arxiv.org/html/2608.17183v1/figures/air_bench_scraped_score_and_proportions_per_model.png)\(b\)LLMs \(AirBench\): proportions of top 36 LLMs with highest proportion of 0\.5\.

Figure 2:The benchmark scores and proportions of the scores per model for \(a\) SLM runs and \(b\) LLM results reported in the literature\[[24](https://arxiv.org/html/2608.17183#bib.bib16),[57](https://arxiv.org/html/2608.17183#bib.bib14)\]\.For judge\-scored safety benchmark suites with ternary labels, Figure[2](https://arxiv.org/html/2608.17183#S4.F2)decomposes each model’s results into the three label categories and overlays the mean score, i\.e\., the average of per\-response scores in0,0\.5,1\{0,0\.5,1\}\. Appendix Table[A1](https://arxiv.org/html/2608.17183#Pt0.A1.T1)reports detailed HCR, AR, SRR, and mean scores for all 26 SLMs\.

As shown in Fig\.[2](https://arxiv.org/html/2608.17183#S4.F2), SLMs receive a significant fraction of “ambiguous” labels, which makes aggregate rankings difficult to interpret, since*an ambiguous label does not indicate whether output is safe or not*and requires special handling in scoring\. To assess whether this is SLM\-specific, Fig\.[2](https://arxiv.org/html/2608.17183#S4.F2)shows AirBench results for 7B\+ LLMs, where ambiguity is much lower\[[24](https://arxiv.org/html/2608.17183#bib.bib16),[57](https://arxiv.org/html/2608.17183#bib.bib14)\]\.

Figure[3](https://arxiv.org/html/2608.17183#S4.F3)further compares SLM and LLM score distributions across four HELM\-style suites and includes SALAD\-Bench as an SLM\-only comparison between raw judge\-derived probabilities and ternary mapping\. For SALAD\-Bench, the raw view shows MD\-Judge unsafe probabilities, while the ternary view maps them to the common0,0\.5,1\{0,0\.5,1\}scale\. The results show that SALAD also assigns a large fraction of SLM outputs to the ambiguous category after mapping, whereas simple safety tests exhibit lower ambiguity, and BBQ reflects bias\-oriented accuracy rather than the same safety–ambiguity structure\.

![Refer to caption](https://arxiv.org/html/2608.17183v1/figures/slm_vs_llm_box_violin_5bm.png)Figure 3:Score distributions for SLMs vs\. LLMs on the HELM\-safety suite\. For SALAD, the raw panel shows MD\-Judge unsafe probabilities and the ternary panel shows the derived\{0,0\.5,1\}\\\{0,0\.5,1\\\}ambiguity scale used in the joint analysis\.Answer to RQ2:Current automated safety benchmarking pipelines provide decision\-useful signals for SLMs only when ambiguity is low\. In ambiguity\-heavy suites, mean scores and rankings become difficult to interpret and unreliable, as ambiguous labels reflect evaluation difficulty or output\-quality artifacts, and model comparisons can shift under reasonable ambiguity treatments\.

#### Ranking sensitivity to ambiguity handling\.

We quantify how model rankings depend on the treatment of “ambiguous” labels\. By default, “ambiguous” is treated as a fixed midpoint in the\{0,0\.5,1\}\\\{0,0\.5,1\\\}scale\. We recompute rankings under explicit alternatives: mapping “ambiguous”→\\rightarrow0 \(pessimistic\), “ambiguous”→\\rightarrow1 \(optimistic\), removing them, and partial credit \(“ambiguous”→\\rightarrow0\.25 and→\\rightarrow0\.75\)\. These policies span how an analyst might read the same0\.50\.5labels: as harmful\-leaning, safe\-leaning, non\-informative, or weakly decisive\. We use them only to test whether model orderings are robust to reasonable ambiguity handling\.

As shown in Figures[2](https://arxiv.org/html/2608.17183#S4.F2)and[3](https://arxiv.org/html/2608.17183#S4.F3), SLM evaluations place substantial probability mass near the ambiguous region, especially for AirBench, HarmBench, and ternary\-mapped SALAD, reducing the interpretability of aggregate means\. Figure[1](https://arxiv.org/html/2608.17183#S3.F1)shows that rankings vary across suites\. When ambiguity is high, ranks shift under reasonable alternative treatments\. Figure[4](https://arxiv.org/html/2608.17183#S4.F4)visualizes each model’s rank trajectory across ambiguity\-handling scenarios for AirBench, HarmBench, SALAD\-Bench, and Simple Safety\. It shows how a model’s safety ranks are sensitive to different handling of “ambiguous” labels, and the inconsistencies across benchmarks\. This ranking instability is further supported by our judge\-robustness analysis on HarmBench and Simple Safety\. As shown in Table[5](https://arxiv.org/html/2608.17183#S4.T5), rank variation across ambiguity treatments remains substantial for HarmBench under both judges, but is comparatively small for Simple Safety\.

![Refer to caption](https://arxiv.org/html/2608.17183v1/figures/rank_changes_scenarios_5bm.png)Figure 4:Rank sensitivity to ambiguity handling\. Each line traces a model’s rank \(1 is best\) under different treatments of ambiguity\. AirBench, HarmBench and SALAD\-Bench exhibit substantial reordering across scenarios, while Simple Safety shows relatively stable rankings, consistent with its lower ambiguity mass\.Answer to RQ3:Model rankings and suite\-level conclusions are sensitive to how ambiguous outcomes are handled when the 0\.5 label is prevalent\. On AirBench, HarmBench, and SALAD\-Bench, rank orderings shift substantially under defensible alternative mappings or exclusion of ambiguous cases\. In contrast, rankings are comparatively stable on lower\-ambiguity suites such as Simple Safety Tests\. The prevalence of ambiguity motivates our study in Part II\.

![Refer to caption](https://arxiv.org/html/2608.17183v1/figures/combined_metric_ar_top3_pos_neg.png)Figure 5:Prompt\-level metric associations with ambiguity rate across all combined safety suites\. Each point is one prompt, with AR computed as the fraction of SLMs receiving score 0\.5 for that prompt\. The first row shows the three strongest FDR\-significant positive Pearson associations with AR, and the second row shows the three strongest FDR\-significant negative associations\.

### 4\.2Part II: Diagnosing confounds and testing robustness

We analyze ambiguity from three complementary perspectives\. First, we identify which prompt and output metrics are associated with higher or lower ambiguity rates\. Second, we assess whether ambiguity can be predicted from these metrics within and across benchmark\-model splits\. Third, we examine whether model metadata and architectural features are associated with ambiguity rates\.

Metric\-ambiguity correlations\.We compute the ambiguity rate \(AR\) of an SLM as the fraction of “ambiguous” labels, grouping prompts by benchmark and instance identifier, and correlate this prompt\-level AR with individual metrics\. Figure[5](https://arxiv.org/html/2608.17183#S4.F5)shows the strongest FDR\-significant Pearson correlations\. We find that AR is positively associated with lexical density, output perplexity, and output length, and negatively associated with lexical sophistication, self\-coherence, and reply–prompt similarity\. These patterns suggest that ambiguity is associated with lower\-quality or hard\-to\-interpret responses, which makes safety judgments more difficult\. This analysis is conducted at the prompt level and aggregated across benchmarks\. It captures associations between metrics and ambiguity rather than implying causal relationships or model\-level scaling effects\.

Benchmark\-specific ambiguity prediction\.Table[1](https://arxiv.org/html/2608.17183#S4.T1)shows that ambiguity is a structured and learnable property of the evaluation pipeline rather than residual noise\. The strongest within\-benchmark signal appears in AirBench and HarmBench, with SALAD showing a weaker but still meaningful pattern, which is consistent with the broader claim that ambiguous judgments arise from recurring combinations of prompt difficulty and model\-capability proxies\. Simple Safety is notably less informative in this analysis, so we use it only for within\-benchmark analysis and exclude it for the transfer checks below\.

Benchmark transfer performance\.Table[2](https://arxiv.org/html/2608.17183#S4.T2)shows that ambiguity signals are not universally transferable across suites\. The strongest transfer appears between AirBench and HarmBench, suggesting a shared ambiguity structure, while the SALAD results indicate weaker and more asymmetric transfer\. This pattern is consistent with ambiguity arising from overlapping but non\-identical combinations of prompt complexity and model capability across benchmarks\.

These results indicate that the capability\-safety confound is heavily influenced by the specific structural design of each evaluation suite\.

Table 1:Benchmark\-specific ambiguity prediction using grouped 80/20 train/test splits byinstance\_id\. The binary target is whether the judge score is ambiguous \(0\.50\.5\) versus decisive \(00or11\); the table reports the best classifier by balanced accuracy for each benchmark\.BenchmarkBest classifierAcc\.Bal\. Acc\.ROC\-AUCTest rowsAirBenchRandom forest0\.7860\.7730\.85529,614HarmBenchCatBoost0\.7860\.7890\.8632,080SALADXGBoost0\.6080\.6160\.664106,600

Table 2:Cross\-benchmark ambiguity\-prediction transfer\. Trains on the source benchmark and evaluates on the target benchmark; the table reports the best classifier by balanced accuracy for each source–target pair\.SourceTargetBest classifierAcc\.Bal\. Acc\.ROC\-AUCTest rowsAirBenchHarmBenchCatBoost0\.7010\.7070\.81010,400AirBenchSALADRandom forest0\.6070\.5680\.578532,950HarmBenchAirBenchCatBoost0\.7440\.7460\.794148,044HarmBenchSALADRandom forest0\.5510\.5530\.563532,950SALADAirBenchCatBoost0\.6670\.6040\.651148,044SALADHarmBenchCatBoost0\.5590\.5760\.63410,400

Table 3:Aggregate ambiguity prediction under unseen\-SLM transfer\. Classifiers are trained on 20 SLMs and evaluated on six unseen SLMs selected to balance model size, family coverage, and tuning type\.BenchmarkBest classifierAcc\.Bal\. Acc\.ROC\-AUCTest rowsAirBenchCatBoost0\.8090\.8030\.87734,164HarmBenchLightGBM0\.8300\.8320\.9052,400

SLM transfer performance\.Table[3](https://arxiv.org/html/2608.17183#S4.T3)further shows that the ambiguity signal is not limited to the particular models seen during training\. For AirBench and HarmBench, predictors trained on one set of SLMs still identify ambiguity on held\-out SLMs, which suggests that the learned signal is tied to more general prompt/model properties rather than to idiosyncrasies of individual models\. Taken together, the benchmark\-specific, benchmark\-transfer, and SLM\-transfer results all indicate that ambiguous judgments are systematically related to evaluation difficulty and model capability, not to random labeling variation\.

Table 4:Pearson correlations between model metadata and \(i\) ambiguity rate \(AR\) and \(ii\) mean safety score \(mean ternary score\) across all 26 SLMs\.Model metricDescriptionrARr\_\{\\mathrm\{AR\}\}pARp\_\{\\mathrm\{AR\}\}rscorer\_\{\\mathrm\{score\}\}pscorep\_\{\\mathrm\{score\}\}Instruction tunedInstruction\-tuned checkpoint\-0\.490\.0110\.360\.068HeadsAttention head count\-0\.420\.035\-0\.250\.21Hidden sizeInternal representation width\-0\.340\.09\-0\.050\.796RMSNormRMS\-based normalization\-0\.300\.1410\.240\.244GQAGrouped\-query attention\-0\.270\.180\.170\.405ParamsTotal model size\-0\.270\.1830\.100\.641ContextMaximum context length\-0\.240\.2340\.340\.088SiLUSmooth nonlinear activation\-0\.200\.325\-0\.270\.176SwiGLUGated feed\-forward activation\-0\.180\.3780\.360\.071Chat tunedChat\-tuned checkpoint indicator\-0\.170\.419\-0\.340\.094RoPERotary position encoding\-0\.100\.6320\.360\.072LayersTransformer layer count\-0\.080\.685\-0\.270\.188VocabTokenizer vocabulary size0\.080\.7050\.672\.1e\-04

Table[4](https://arxiv.org/html/2608.17183#S4.T4)extends the metadata analysis to all 26 SLMs using the architecture information in Table[A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2)\. The clearest pattern is that model metadata is associated more strongly with ambiguity rate than with the mean ternary safety score\. The strongest negative AR associations are instruction tuning \(r=−0\.49r\{=\}\-0\.49,p=0\.011p\{=\}0\.011\) and attention\-head count \(r=−0\.42r\{=\}\-0\.42,p=0\.035p\{=\}0\.035\)\. Parameter count has the same negative direction but is weaker \(r=−0\.27r\{=\}\-0\.27,p=0\.183p\{=\}0\.183\), so the metadata results should not be read as simply monotonic scaling\. The mean safety score is generally less consistently associated with the same metadata, with vocabulary size standing out as the clearest score correlation \(r=0\.67r\{=\}0\.67,p=2\.1×10−4p\{=\}2\.1\\times 10^\{\-4\}\)\. We therefore treat these correlations as descriptive evidence that ambiguity tracks model\-design and tuning factors, not as causal claims about architecture\.

Answer to RQ4:Ambiguous labels are primarily associated with evaluation difficulty: they concentrate on harder prompts and \(especially\) low\-quality or hard\-to\-interpret generations, are more prevalent for some model metadata profiles such as non\-instruction\-tuned and older releases, and remain predictable under unseen\-SLM and cross\-benchmark transfers\. This provides evidence of a capability\-safety confound: the ambiguous label can partially reflect general generation capability and interpretability artifacts rather than a distinct “partially safe” behavior category\.

#### Judge robustness: GPT vs\. Llama

We evaluate judge robustness on HarmBench and Simple Safety Tests by comparing GPT\- and Llama\-based labels over the 16\-model overlap\. We found that while both judges exhibit strong overall agreement on clear\-cut cases \(0\.75 agreement on HarmBench and 0\.81 on Simple Safety Tests\), the remaining disagreement heavily destabilizes rankings\. Specifically, on HarmBench, direct contradictions \(0↔10\\leftrightarrow 1\) occur rarely \(4%\), while most disagreement occurs at the ambiguity boundary \(0\.5↔\{0,1\}=0\.210\.5\\leftrightarrow\\\{0,1\\\}=0\.21\)\.

This boundary sensitivity destabilizes rankings: despite high overall agreement, ambiguity\-driven disagreement produces substantial rank variation \(with a mean rank range of 7\.44 under GPT and 4\.50 under Llama\)\. These results illustrate the capability\-safety confound: automated judges reliably agree when SLMs’ outputs are decisive, but consistency breaks down near the ambiguous boundary where lower\-quality SLM generations are harder to interpret\. Full confusion matrices for this judge\-overlap subset are reported in Table[6](https://arxiv.org/html/2608.17183#S4.T6)\.

Table 5:Judge robustness summary \(GPT vs\. Llama\) for HarmBench and Simple Safety Tests\. Ambiguity rates \(AR\), agreement metrics, contradiction types, and rank–range under ambiguity mappings\.SuiteAR\(GPT\)AR\(Llama\)Δ\\DeltaARExactagreeκw\\kappa\_\{w\}↔0\.5\\\!\\leftrightarrow\\\!\{0,1\}\\\{0,1\\\}↔10\\\!\\leftrightarrow\\\!1Rank\-range\(GPT\)Rank\-range\(Llama\)HarmBench0\.550\.470\.080\.750\.560\.210\.047\.444\.50Simple Safety0\.060\.030\.040\.810\.700\.080\.110\.751\.19Table 6:Confusion matrices between GPT and Llama judges\. Cells report count \(percent of suite instances\)\. The main pattern is that disagreement concentrates around the ambiguous boundary \(0\.5↔\{0,1\}0\.5\\leftrightarrow\\\{0,1\\\}\), especially on HarmBench, rather than direct0↔10\\leftrightarrow 1flips\. Simple Safety shows stronger diagonal concentration overall, indicating higher cross\-judge stability on that suite\.HarmBench \(N=6400N=6400\)Simple Safety Tests \(N=1600N=1600\)Llama 0Llama 0\.5Llama 1Llama 0Llama 0\.5Llama 1GPT 0234 \(3\.66%\)196 \(3\.06%\)2 \(0\.03%\)640 \(40\.00%\)17 \(1\.06%\)92 \(5\.75%\)GPT 0\.5502 \(7\.84%\)2594 \(40\.53%\)437 \(6\.83%\)45 \(2\.81%\)10 \(0\.62%\)47 \(2\.94%\)GPT 1261 \(4\.08%\)233 \(3\.64%\)1941 \(30\.33%\)85 \(5\.31%\)13 \(0\.81%\)651 \(40\.69%\)

## 5Discussions

### 5\.1Benchmark pipeline validity for SLM evaluation

Our results indicate that current automated safety benchmarking pipelines may provide useful signals for SLM safety/security at the extremes, i\.e\., clear harmful compliance and clear refusals\. However, they are insufficient as standalone instruments for SLM decision\-making: benchmark effectiveness and consistency are limited because ambiguous labels often dominate the decisions, while such labels are strongly predicted by output quality and prompt complexity, consistent with a capability\-safety confound\. Meanwhile, model\-class invariance is also threatened: smaller and older models tend to generate lower\-quality outputs that trigger more ambiguous judgments\. Finally, aggregation stability is weak: comparative conclusions shift substantially under reasonable alternative treatments of the “ambiguous” labels\.

Even if “ambiguous” is the semantically correct label for an incoherent response, its prevalence still indicates the benchmarks’ ineffectiveness in safety/security assessment and renders the aggregate safety metric less useful\. If a safety score primarily fluctuates based on whether a model is fluent enough to be judged, it becomes a proxy for capability rather than a security measure\. High ambiguity rates therefore signal a failure of the benchmark’s utility\. These concerns mirror broader critiques that benchmark\-based safety progress can be difficult to interpret when scores are dominated by confounds such as general model capabilities rather than the intended safety construct\[[48](https://arxiv.org/html/2608.17183#bib.bib2)\]\.

### 5\.2Practical implications and recommendations

With our findings on the benchmark\-based SLM safety and security evaluation, we make the following recommendations intended to improve the effectiveness and robustness of the process for security\- and privacy\-relevant decision\-making: \(1\)Report ambiguity explicitly\.Treat the ambiguity rate \(A​RAR\) as a first\-class outcome rather than collapsing intermediate cases into the mean\. \(2\)Benchmarks for SLMs\.Develop SLM\-specific safety/security benchmarks for the less\-powerful SLMs\. Our findings in Section[4\.2](https://arxiv.org/html/2608.17183#S4.SS2)could provide practical guidance to design prompts that reduce ambiguity rates\. \(3\) When ambiguity is common, a single mean score can overstate how much*decision\-ready*signal a benchmark provides\. We propose a penalized score alongside raw means, defined as follows:

Ambiguity\-adjusted score\.Letsi∈\[0,1\]s\_\{i\}\\in\[0,1\]be the judge score for instanceii, letSr​a​w=1N​∑isiS\_\{raw\}=\\frac\{1\}\{N\}\\sum\_\{i\}s\_\{i\}, and letAR=1N∑i𝕀\[si=0\.5\]AR=\\frac\{1\}\{N\}\\sum\_\{i\}\\mathbb\{I\}\[s\_\{i\}=0\.5\]\. We reportSa​d​j=Sr​a​w​\(1−A​R\)S\_\{adj\}=S\_\{raw\}\(1\-AR\)as an ambiguity\-adjusted diagnostic\. The factor\(1−A​R\)\(1\-AR\)is the fraction of evaluations that receive decisive labels, soSa​d​jS\_\{adj\}discounts the raw score by the share of benchmark outcomes that are immediately interpretable as non\-ambiguous\. This does not assume thatSa​d​jS\_\{adj\}is a calibrated safety utility; rather, it makes explicit how much of the raw score remains after penalizing ambiguity\. As a compact illustration, Appendix Table[A1](https://arxiv.org/html/2608.17183#Pt0.A1.T1)reportsSa​d​jS\_\{adj\}and the resulting rank impact \(Δ\\DeltaR\) alongside the raw mean and outcome decomposition\.

Answer to RQ5:Automated safety benchmarking pipelines are less informative and not decision\-useful for SLMs when ambiguity is common\. Without new human ground truth, we can still show that ambiguity is systematically associated with capability\-linked artifacts and that leaderboards are brittle to reasonable aggregation choices\. Decision\-useful SLM evaluation therefore requires explicit ambiguity reporting and sensitivity analyses over ambiguity handling\.

### 5\.3Limitations

This study examines the prompts, rubrics, and judge models used in LLM safety benchmarks\. Our core claims do not require human ground\-truth validation: as demonstrated by our rank sensitivity analysis, aggregate mean\-score rankings are mathematically unstable\. When the automated pipeline generates a high volume of ambiguous labels, the benchmark lacks decision utility for SLMs, even if “ambiguous” is the semantically correct label for incoherent responses\. However, it is still beneficial to employ human experts to review the “ambiguous” cases\. They may provide insightful evidence on the root causes of such “ambiguous” scores, and actionable suggestions for the design of SLM safety guardrails and benchmarks\.

Finally, our future work includes expanding model coverage, exploring alternative judge models, validating the SALAD\-Bench mapping under different threshold choices, and, most importantly, developing SLM\-specific benchmarks with the guidance of our findings in this paper\.

## 6Conclusion

We present a large\-scale evaluation of automated LLM safety benchmarking pipelines on 26 small language models\. Across 715,312 prompt\-SLM evaluations, ambiguous results are prevalent, which significantly impacts the effectiveness and consistency of the benchmarks\. As a result, benchmark\-derived rankings are mathematically unstable, with substantial shifts under reasonable ambiguity\-handling choices\. We further demonstrate that the ambiguous labels are systematically associated with output quality and prompt complexity\. Together, these findings suggest that standard automated pipelines are not decision\-useful for SLM safety claims without explicit ambiguity handling and robustness checks\. Therefore, we call for the development of SLM\-specific safety, security, and compliance benchmarks that better align with the capacity of small language models\.

Acknowledgments\.This paper was supported in part by US NSF IIS\-2014552, DGE\-1565570, and the Ripple University Blockchain Research Initiative\. We thank the anonymous reviewers for their valuable comments and suggestions\.

## References

- \[1\]Alibaba\(2024\)Qwen/qwen2\.5\-0\.5b\-instruct model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.9.1.1.1.2.1),[§2\.1](https://arxiv.org/html/2608.17183#S2.SS1.p1.1)\.
- \[2\]Alibaba\(2024\)Qwen/qwen2\.5\-1\.5b\-instruct model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.19.1.1.1.2.1)\.
- \[3\]Alibaba\(2024\)Qwen/qwen2\.5\-3b\-instruct model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.24.1.1.1.2.1)\.
- \[4\]Allen Institute for AI\(2025\)Allenai/olmo\-2\-0425\-1b model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.13.1.1.1.2.1)\.
- \[5\]A\. Breneman, M\. H\. Trager, E\. R\. Gordon, and F\. H\. Samie\(2024\)Readability rescue: large language models may improve readability of patient education materials\.Archives of Dermatological Research316\(9\),pp\. 669\.Cited by:[§3\.3](https://arxiv.org/html/2608.17183#S3.SS3.p1.1),[§3\.3](https://arxiv.org/html/2608.17183#S3.SS3.p5.1)\.
- \[6\]T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah,et al\.\(2020\)Language models are few\-shot learners\.NeurIPS\.Cited by:[§2\.3](https://arxiv.org/html/2608.17183#S2.SS3.p2.1),[§3\.3](https://arxiv.org/html/2608.17183#S3.SS3.p5.1),[§3\.3](https://arxiv.org/html/2608.17183#S3.SS3.p6.1)\.
- \[7\]L\. Chen and G\. Varoquaux\(2025\)What is the role of small models in the llm era: a survey\.arXiv preprint arXiv:2409\.06857\.Cited by:[§2\.1](https://arxiv.org/html/2608.17183#S2.SS1.p2.1)\.
- \[8\]Y\. Chen, J\. Zhao, and H\. Han\(2025\)A survey on collaborative mechanisms between large and small language models\.arXiv preprint arXiv:2505\.07460\.Cited by:[§2\.1](https://arxiv.org/html/2608.17183#S2.SS1.p2.1)\.
- \[9\]Cyberspace Administration of China\(2023\)Interim measures for the management of generative artificial intelligence services\.Note:Cyberspace Administration of ChinaCited by:[§2\.2](https://arxiv.org/html/2608.17183#S2.SS2.p2.1)\.
- \[10\]Databricks\(2023\)Databricks/dolly\-v2\-3b model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.25.1.1.1.2.1)\.
- \[11\]DeepSeek\(2024\)DeepSeek\-r1\-qwen\-1\.5b model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.18.1.1.1.2.1)\.
- \[12\]A\. Dumitrache, O\. Inel, L\. Aroyo, B\. Timmermans, and C\. Welty\(2018\)CrowdTruth 2\.0: quality metrics for crowdsourcing with disagreement\.Cited by:[§2\.3](https://arxiv.org/html/2608.17183#S2.SS3.p2.1)\.
- \[13\]European Parliament and Council of the European Union\(2024\)Regulation \(eu\) 2024/1689 laying down harmonised rules on artificial intelligence\.Note:Official Journal of the European UnionCited by:[§2\.2](https://arxiv.org/html/2608.17183#S2.SS2.p2.1)\.
- \[14\]M\. Garg, S\. Raza, S\. Rayana, X\. Liu, and S\. Sohn\(2025\)The rise of small language models in healthcare: a comprehensive survey\.arXiv preprint arXiv:2504\.17119\.Cited by:[§1](https://arxiv.org/html/2608.17183#S1.p1.1)\.
- \[15\]Google\(2024\)Google/gemma\-2\-2b\-it model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.23.1.1.1.2.1)\.
- \[16\]Google\(2025\)Google/gemma\-3\-1b\-it model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.11.1.1.1.2.1)\.
- \[17\]Google\(2025\)Google/gemma\-3\-270m\-it model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.5.1.1.1.2.1),[§2\.1](https://arxiv.org/html/2608.17183#S2.SS1.p1.1)\.
- \[18\]H2O\.ai\(2024\)H2oai/h2o\-danube2\-1\.8b\-chat model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.22.1.1.1.2.1)\.
- \[19\]H2O\.ai\(2024\)H2oai/h2o\-danube3\.1\-4b\-chat model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.27.1.1.1.2.1)\.
- \[20\]Hugging FaceHuggingFaceTB/smollm2\-1\.7b\-instruct model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.21.1.1.1.2.1)\.
- \[21\]Hugging FaceHuggingFaceTB/smollm2\-135m\-instruct model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.4.1.1.1.2.1)\.
- \[22\]Hugging FaceHuggingFaceTB/smollm2\-360m\-instruct model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.8.1.1.1.2.1)\.
- \[23\]T\. Ivanov and V\. Penchev\(2024\)AI benchmarks and datasets for llm evaluation\.arXiv preprint arXiv:2412\.01020\.Cited by:[§2\.2](https://arxiv.org/html/2608.17183#S2.SS2.p1.1)\.
- \[24\]F\. Kaiyom, A\. Ahmed, Y\. Mai, K\. Klyman, R\. Bommasani, and P\. Liang\(2024\)Helm safety: towards standardized safety evaluations of language models\.Stanford Center for Research on Foundation Models\.Cited by:[§1](https://arxiv.org/html/2608.17183#S1.p2.1),[§2\.3](https://arxiv.org/html/2608.17183#S2.SS3.p1.1),[§2\.3](https://arxiv.org/html/2608.17183#S2.SS3.p2.1),[3rd item](https://arxiv.org/html/2608.17183#S3.I1.i3.p1.1),[4th item](https://arxiv.org/html/2608.17183#S3.I1.i4.p1.1),[§3\.1](https://arxiv.org/html/2608.17183#S3.SS1.p5.1),[Figure 2](https://arxiv.org/html/2608.17183#S4.F2),[§4\.1](https://arxiv.org/html/2608.17183#S4.SS1.SSSx2.p4.1)\.
- \[25\]E\. Leonardelli, V\. Basile, M\. Poesio, M\. Umaña, and D\. Stojanov\(2021\)Agreeing to disagree: annotating offensive language datasets with annotators’ disagreement\.InEMNLP,Cited by:[§2\.3](https://arxiv.org/html/2608.17183#S2.SS3.p2.1)\.
- \[26\]H\. Li, D\. Guo, D\. Li, W\. Fan, Q\. Hu, X\. Liu, C\. Chan,et al\.\(2024\)Privlm\-bench: a multi\-level privacy evaluation benchmark for language models\.InACL,Cited by:[§2\.2](https://arxiv.org/html/2608.17183#S2.SS2.p1.1)\.
- \[27\]L\. Li, B\. Dong, R\. Wang, X\. Hu, W\. Zuo, D\. Lin, Y\. Qiao, and J\. Shao\(2024\)Salad\-bench: a hierarchical and comprehensive safety benchmark for large language models\.InFindings of the Association for Computational Linguistics: ACL 2024,Cited by:[§2\.2](https://arxiv.org/html/2608.17183#S2.SS2.p1.1),[§2\.3](https://arxiv.org/html/2608.17183#S2.SS3.p1.1),[2nd item](https://arxiv.org/html/2608.17183#S3.I1.i2.p1.1.1),[§3\.1](https://arxiv.org/html/2608.17183#S3.SS1.p5.1),[§3\.2](https://arxiv.org/html/2608.17183#S3.SS2.p1.1)\.
- \[28\]T\. Li, W\. Chiang, E\. Frick, L\. Dunlap, T\. Wu, B\. Zhu, J\. E\. Gonzalez, and I\. Stoica\(2025\)From crowdsourced data to high\-quality benchmarks: arena\-hard and benchbuilder pipeline\.InICML,Cited by:[§2\.3](https://arxiv.org/html/2608.17183#S2.SS3.p1.1)\.
- \[29\]B\. Y\. Lin, K\. Deng, F\. Brahman, A\. Ravichander, V\. Pyatkin, N\. Dziri,et al\.\(2024\)WildBench: benchmarking llms with challenging tasks from real users in the wild\.arXiv preprint arXiv:2406\.04770\.Cited by:[§2\.3](https://arxiv.org/html/2608.17183#S2.SS3.p1.1)\.
- \[30\]Y\. Liu, D\. Iter, Y\. Xu, S\. Wang,et al\.\(2023\)G\-Eval: NLG evaluation using GPT\-4 with better human alignment\.InEMNLP,Cited by:[§2\.3](https://arxiv.org/html/2608.17183#S2.SS3.p1.1),[§2\.3](https://arxiv.org/html/2608.17183#S2.SS3.p2.1)\.
- \[31\]F\. Marulli, L\. Campanile, M\. S\. de Biase, S\. Marrone, L\. Verde, and M\. Bifulco\(2024\)Understanding readability of large language models output: an empirical analysis\.Procedia Computer Science246,pp\. 5273–5282\.Cited by:[§3\.3](https://arxiv.org/html/2608.17183#S3.SS3.p1.1),[§3\.3](https://arxiv.org/html/2608.17183#S3.SS3.p5.1)\.
- \[32\]M\. Mazeika, L\. Phan, X\. Yin, A\. Zou, Z\. Wang, N\. Mu, E\. Sakhaee, N\. Li, S\. Basart, B\. Li,et al\.\(2024\)HarmBench: a standardized evaluation framework for automated red teaming and robust refusal\.InICML,Cited by:[§2\.2](https://arxiv.org/html/2608.17183#S2.SS2.p1.1),[3rd item](https://arxiv.org/html/2608.17183#S3.I1.i3.p1.1.1),[§3\.2](https://arxiv.org/html/2608.17183#S3.SS2.p1.1)\.
- \[33\]Meta AI\(2022\)Facebook/opt\-125m model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.3.1.1.1.2.1)\.
- \[34\]Meta AI\(2022\)Facebook/opt\-350m model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.6.1.1.1.2.1)\.
- \[35\]Meta\(2024\)Llama 3\.2\-1b instruct model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.16.1.1.1.2.1)\.
- \[36\]Meta\(2024\)Llama 3\.2\-3b instruct model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.26.1.1.1.2.1)\.
- \[37\]Meta\(2024\)MobileLLM\-1b model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.12.1.1.1.2.1)\.
- \[38\]National Institute of Standards and Technology\(2023\)Artificial intelligence risk management framework \(ai rmf 1\.0\)\.Note:NIST AI 100\-1Cited by:[§2\.2](https://arxiv.org/html/2608.17183#S2.SS2.p2.1)\.
- \[39\]C\. V\. Nguyen, X\. Shen, R\. Aponte, Y\. Xia, S\. Basu, Z\. Hu,et al\.\(2024\)A survey of small language models\.arXiv preprint arXiv:2410\.20011\.Cited by:[§1](https://arxiv.org/html/2608.17183#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.17183#S2.SS1.p1.1)\.
- \[40\]J\. Ni, F\. Xue, X\. Yue, Y\. Deng, M\. Shah, K\. Jain, G\. Neubig, and Y\. You\(2024\)MixEval: deriving wisdom of the crowd from llm benchmark mixtures\.InNeurIPS,Cited by:[§2\.3](https://arxiv.org/html/2608.17183#S2.SS3.p1.1)\.
- \[41\]OpenAI\(2019\)Gpt2 model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.2.1.1.1.2.1),[§2\.1](https://arxiv.org/html/2608.17183#S2.SS1.p1.1)\.
- \[42\]OpenAI\(2019\)Gpt2\-large model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.10.1.1.1.2.1)\.
- \[43\]OpenAI\(2019\)Gpt2\-medium model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.7.1.1.1.2.1)\.
- \[44\]OpenAI\(2019\)Gpt2\-xl model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.17.1.1.1.2.1)\.
- \[45\]L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal,et al\.\(2022\)Training language models to follow instructions with human feedback\.NeurIPS\.Cited by:[§1](https://arxiv.org/html/2608.17183#S1.p2.1),[§2\.3](https://arxiv.org/html/2608.17183#S2.SS3.p2.1)\.
- \[46\]A\. Parrish, A\. Chen, N\. Nangia, V\. Padmakumar, J\. Phang, J\. Thompson, P\. M\. Htut, and S\. Bowman\(2022\)BBQ: a hand\-built bias benchmark for question answering\.InFindings of ACL,Cited by:[1st item](https://arxiv.org/html/2608.17183#S1.I1.i1.p1.1),[5th item](https://arxiv.org/html/2608.17183#S3.I1.i5.p1.1.1),[§4\.1](https://arxiv.org/html/2608.17183#S4.SS1.SSSx1.p3.1)\.
- \[47\]N\. Reimers and I\. Gurevych\(2019\)Sentence\-bert: sentence embeddings using siamese bert\-networks\.InEMNLP,Cited by:[§3\.3](https://arxiv.org/html/2608.17183#S3.SS3.p1.1),[§3\.3](https://arxiv.org/html/2608.17183#S3.SS3.p6.1)\.
- \[48\]R\. Ren, S\. Basart, A\. Khoja, A\. Pan, A\. Gatti,et al\.\(2024\)Safetywashing: do ai safety benchmarks actually measure safety progress?\.NeurIPS\.Cited by:[§5\.1](https://arxiv.org/html/2608.17183#S5.SS1.p2.1)\.
- \[49\]Stability AI\(2024\)Stabilityai/stablelm\-2\-1\_6b model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.20.1.1.1.2.1)\.
- \[50\]S\. Subramanian, V\. Elango, and M\. Gungor\(2025\)Small language models \(slms\) can still pack a punch: a survey\.arXiv preprint arXiv:2501\.05465\.Cited by:[§1](https://arxiv.org/html/2608.17183#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.17183#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2608.17183#S3.SS1.p1.1)\.
- \[51\]TinyLlama\(2023\)TinyLlama/tinyllama\-1\.1b\-chat\-v1\.0 model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.14.1.1.1.2.1)\.
- \[52\]B\. Vidgen, N\. Scherrer, H\. R\. Kirk, R\. Qian, A\. Kannappan, S\. A\. Hale, and P\. Röttger\(2023\)SimpleSafetyTests: a test suite for evaluating llm safety\.arXiv preprint arXiv:2311\.08370\.Cited by:[4th item](https://arxiv.org/html/2608.17183#S3.I1.i4.p1.1.1)\.
- \[53\]F\. Wang, M\. Lin, Y\. Ma, H\. Liu, Q\. He, X\. Tang, J\. Tang, J\. Pei, and S\. Wang\(2025\)A survey on small language models in the era of large language models: architecture, capabilities, and trustworthiness\.InACM SIGKDD,Cited by:[§1](https://arxiv.org/html/2608.17183#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.17183#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.17183#S2.SS1.p2.1)\.
- \[54\]F\. Wang Z\. Zhanget al\.\(2025\)A comprehensive survey of small language models in the era of large language models: techniques, enhancements, applications, collaboration with llms, and trustworthiness\.ACM Trans\. Intell\. Syst\. Technol\.\.Cited by:[§2\.1](https://arxiv.org/html/2608.17183#S2.SS1.p2.1)\.
- \[55\]A\. Wei, N\. Haghtalab, and J\. Steinhardt\(2023\)Jailbroken: how does llm safety training fail?\.InNeurIPS,Cited by:[§2\.2](https://arxiv.org/html/2608.17183#S2.SS2.p1.1)\.
- \[56\]C\. White, S\. Dooley, M\. Roberts, A\. Pal, B\. Feuer, S\. Jain, R\. Shwartz\-Ziv, N\. Jain,et al\.\(2025\)LiveBench: a challenging, contamination\-limited llm benchmark\.InICLR,Cited by:[§2\.3](https://arxiv.org/html/2608.17183#S2.SS3.p1.1)\.
- \[57\]Y\. Zeng, Y\. Yang, A\. Zhou, J\. Z\. Tan, Y\. Tu, Y\. Mai, K\. Klyman,et al\.\(2024\)Air\-bench 2024: a safety benchmark based on risk categories from regulations and policies\.AGI\-Artificial General Intelligence\-Robotics\-Safety & Alignment\.Cited by:[§2\.2](https://arxiv.org/html/2608.17183#S2.SS2.p1.1),[1st item](https://arxiv.org/html/2608.17183#S3.I1.i1.p1.1.1),[§3\.2](https://arxiv.org/html/2608.17183#S3.SS2.p1.1),[Figure 2](https://arxiv.org/html/2608.17183#S4.F2),[§4\.1](https://arxiv.org/html/2608.17183#S4.SS1.SSSx2.p4.1)\.
- \[58\]L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang,et al\.\(2023\)Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.Cited by:[§1](https://arxiv.org/html/2608.17183#S1.p2.1),[§2\.3](https://arxiv.org/html/2608.17183#S2.SS3.p1.1),[§2\.3](https://arxiv.org/html/2608.17183#S2.SS3.p2.1)\.
- \[59\]J\. Zhou, T\. Lu, S\. Mishra, S\. Brahma, S\. Basu,et al\.\(2023\)Instruction\-following evaluation for large language models\.arXiv preprint arXiv:2311\.07911\.Cited by:[§2\.3](https://arxiv.org/html/2608.17183#S2.SS3.p1.1)\.
- \[60\]A\. Zou, Z\. Wang, N\. Carlini, M\. Nasr, J\. Z\. Kolter, and M\. Fredrikson\(2023\)Universal and transferable adversarial attacks on aligned language models\.Note:arXiv preprint arXiv:2307\.15043Cited by:[§2\.2](https://arxiv.org/html/2608.17183#S2.SS2.p1.1)\.
- \[61\]Zyphra\(2024\)Zyphra/zamba2\-1\.2b\-instruct model card\.Note:Hugging FaceCited by:[Table A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2.2.1.15.1.1.1.2.1)\.

## Appendix 0\.AAdditional Tables

In Table[A1](https://arxiv.org/html/2608.17183#Pt0.A1.T1), we present the per\-model raw mean score, HCR, AR, SRR, ambiguity\-adjusted score, and resulting rank changes for AirBench, SALAD, and HarmBench across all the SLMs\. In Table[A2](https://arxiv.org/html/2608.17183#Pt0.A1.T2), we summarize the 26 SLMs evaluated in this study and the model metadata used\.

Table A1:Per\-model outcome summary for AirBench, SALAD, and HarmBench\. We report the raw ternary\-score meanSMS\_\{\\mathrm\{M\}\}, harmful\-completion rate \(HCR; score=0=0\), ambiguity rate \(AR; score=0\.5=0\.5\), safe\-refusal rate \(SRR; score=1=1\), and the ambiguity\-adjusted scoreSa​d​j=SM⋅\(1−AR\)S\_\{adj\}=S\_\{\\mathrm\{M\}\}\\cdot\(1\-\\mathrm\{AR\}\)\. We also report the implied rank change \(Δ\\DeltaR\) relative to ranking bySMS\_\{\\mathrm\{M\}\}\.InΔ\\DeltaR,↑\\uparrow/↓\\downarrowindicates direction, and the number is the absolute number of rank positions moved\.AirBenchSALADHarmBenchModelSMS\_\{\\mathrm\{M\}\}HCRARSRRSa​d​jS\_\{adj\}Δ\\DeltaRSMS\_\{\\mathrm\{M\}\}HCRARSRRSa​d​jS\_\{adj\}Δ\\DeltaRSMS\_\{\\mathrm\{M\}\}HCRARSRRSa​d​jS\_\{adj\}Δ\\DeltaRGPT\-2 Small0\.590\.070\.690\.240\.18↓\\downarrow140\.580\.220\.390\.380\.35↓\\downarrow20\.560\.010\.860\.130\.08↓\\downarrow11OPT\-125M0\.490\.110\.810\.080\.09↓\\downarrow160\.610\.130\.520\.350\.29↓\\downarrow150\.580\.010\.820\.170\.11↓\\downarrow10Smollm2 135m IT0\.330\.460\.430\.110\.19↑\\uparrow20\.460\.470\.140\.390\.39↑\\uparrow110\.510\.120\.740\.140\.13↑\\uparrow2Gemma 3\-270M IT0\.480\.190\.660\.160\.17↓\\downarrow110\.710\.210\.150\.640\.61↑\\uparrow20\.680\.010\.610\.380\.26↓\\downarrow1OPT\-350M0\.480\.150\.740\.110\.12↓\\downarrow100\.580\.200\.430\.360\.33↓\\downarrow50\.570\.010\.840\.150\.09↓\\downarrow11GPT\-2 Medium0\.540\.140\.640\.220\.19↓\\downarrow90\.590\.210\.390\.390\.36↓\\downarrow30\.550\.030\.840\.130\.09↓\\downarrow7Smollm2 360m IT0\.270\.620\.230\.150\.21↑\\uparrow80\.540\.410\.110\.480\.47↑\\uparrow100\.420\.300\.550\.140\.19↑\\uparrow8Qwen 2\.5\-0\.5B IT0\.370\.570\.120\.310\.33↑\\uparrow100\.670\.230\.210\.560\.52↑\\uparrow10\.630\.220\.300\.480\.44↑\\uparrow1GPT\-2 Large0\.460\.230\.620\.150\.18↓\\downarrow40\.560\.280\.320\.400\.38↑\\uparrow40\.560\.090\.700\.210\.17↓\\downarrow4TinyLlama 1\.1B chat0\.190\.730\.160\.110\.16↑\\uparrow30\.590\.270\.270\.460\.43↑\\uparrow20\.380\.380\.490\.140\.19↑\\uparrow11OLMo 2 0425 1B0\.600\.150\.490\.360\.31↓\\downarrow60\.500\.310\.370\.320\.32↑\\uparrow20\.550\.040\.820\.140\.10↓\\downarrow3MobileLLM\-1B0\.470\.150\.750\.100\.12↓\\downarrow100\.460\.370\.330\.290\.31↑\\uparrow10\.550\.020\.850\.120\.08↓\\downarrow7Gemma 3\-1B IT0\.510\.300\.390\.320\.31–0\.750\.030\.440\.530\.42↓\\downarrow100\.790\.010\.400\.590\.47↓\\downarrow4Llama 3\.2\-1B IT0\.540\.420\.080\.500\.49↑\\uparrow30\.710\.200\.190\.620\.58↑\\uparrow20\.790\.100\.230\.680\.61↑\\uparrow1Zamba2 1\.2B IT0\.480\.470\.110\.420\.43↑\\uparrow60\.610\.290\.210\.500\.48↑\\uparrow30\.720\.120\.320\.560\.49–DeepSeek\-R1\-Qwen0\.250\.620\.260\.120\.19↑\\uparrow50\.540\.280\.370\.350\.34↑\\uparrow20\.510\.030\.920\.050\.04↓\\downarrow7GPT\-2 XL0\.460\.250\.580\.170\.19↑\\uparrow10\.570\.270\.330\.400\.38↑\\uparrow20\.560\.090\.690\.220\.18↓\\downarrow3Qwen 2\.5\-1\.5B IT0\.610\.360\.070\.570\.56↑\\uparrow10\.880\.040\.160\.810\.75–0\.820\.130\.100\.770\.74↑\\uparrow1Stablelm 2 1 6b0\.700\.120\.370\.510\.44↓\\downarrow50\.640\.250\.220\.530\.50↑\\uparrow10\.560\.050\.790\.170\.12↓\\downarrow4Smollm2 1\.7b IT0\.320\.620\.110\.270\.29↑\\uparrow90\.630\.290\.140\.560\.54↑\\uparrow40\.420\.360\.440\.200\.24↑\\uparrow11danube2 1\.8b chat0\.250\.700\.100\.200\.23↑\\uparrow110\.580\.330\.180\.490\.48↑\\uparrow80\.320\.510\.340\.140\.21↑\\uparrow14Gemma 2\-2B IT0\.670\.240\.170\.590\.56↑\\uparrow10\.690\.040\.540\.420\.32↓\\downarrow170\.860\.010\.260\.730\.64↓\\downarrow1Llama 3\.2\-3B IT0\.530\.430\.070\.500\.49↑\\uparrow50\.510\.280\.410\.310\.30↓\\downarrow20\.690\.220\.180\.590\.56↑\\uparrow2Qwen 2\.5\-3B IT0\.480\.470\.080\.440\.44↑\\uparrow70\.770\.090\.290\.620\.55↓\\downarrow20\.780\.120\.220\.670\.61↑\\uparrow1Dolly v2 3b0\.230\.630\.270\.100\.17↑\\uparrow40\.610\.240\.310\.450\.42↓\\downarrow40\.400\.300\.610\.090\.16↑\\uparrow6danube3\.1 4b chat0\.280\.660\.110\.230\.25↑\\uparrow90\.600\.300\.200\.500\.48↑\\uparrow50\.360\.470\.350\.180\.23↑\\uparrow14

Table A2:SLMs evaluated in this study \(26 models spanning 124M–4\.0B parameters\) and the model metadata\. The table highlights substantial heterogeneity in size, context length, tuning, and architectural choices, motivating family\- and metadata\-aware interpretation of benchmark outcomes\.Vendor, year / ModelParams / CtxArchitecture \(reported\)Training data \(reported\)Objective \(reported\)OpenAI 2019GPT\-2 Small\[[41](https://arxiv.org/html/2608.17183#bib.bib36)\]124M1024Decoder\-only transformer; 12 layers; 12 attention heads; 768 hidden; 3072 FFN; GELU; absolute position embeddings; vocab 50257WebText∼\\sim40GB from 8M web pages; cutoff Dec 2017Causal language modelingMeta 2022OPT\-125M\[[33](https://arxiv.org/html/2608.17183#bib.bib40)\]125M2048Decoder\-only transformer \(OPT/GPT\-style\); 12 layers; 12 heads; 768 hidden; 3072 FFN; ReLU; max pos 2048; vocab 50272180B tokens; BookCorpus; Common Crawl; Reddit; predominantly EnglishCausal language modelingHF 2024SmolLM2\-135M IT\[[21](https://arxiv.org/html/2608.17183#bib.bib58)\]135M8192Transformer decoder \(Llama\-family\); 30 layers; 9 heads; 3 KV heads; 576 hidden; 1536 FFN; vocab 491522T\-token curated pre\-training mix; instruction\-tuning dataPre\-training \+ SFT \+ DPOGoogle 2025Gemma3\-270M IT\[[17](https://arxiv.org/html/2608.17183#bib.bib42)\]270M32768Decoder\-only; 270M total \(170M embedding from 256K vocab \+ 100M transformer blocks\)2T tokens; web; code; math; 140\+ languagesPre\-training \+ instruction tuningMeta 2022OPT\-350M\[[34](https://arxiv.org/html/2608.17183#bib.bib41)\]350M2048Decoder\-only transformer \(OPT/GPT\-style\); 24 layers; 16 heads; 1024 hidden; 4096 FFN; ReLU; max pos 2048; vocab 50272180B tokens; BookCorpus; Common Crawl; Reddit; predominantly EnglishCausal language modelingOpenAI 2019GPT\-2 Medium\[[43](https://arxiv.org/html/2608.17183#bib.bib37)\]355M1024Decoder\-only transformer; 24 layers; 16 attention heads; 1024 hidden; 4096 FFN; GELU; absolute position embeddings; vocab 50257WebText∼\\sim40GB from 8M web pages; cutoff Dec 2017Causal language modelingHF 2024SmolLM2\-360M IT\[[22](https://arxiv.org/html/2608.17183#bib.bib59)\]360M8192Transformer decoder \(Llama\-family\); 32 layers; 15 heads; 5 KV heads; 960 hidden; 2560 FFN; vocab 491524T\-token curated pre\-training mix; instruction\-tuning dataPre\-training \+ SFT \+ DPOAlibaba 2024Qwen2\.5\-0\.5B IT\[[1](https://arxiv.org/html/2608.17183#bib.bib49)\]500M32768Decoder\-only; 24 layers; 896 hidden; 14 heads 2 KV \(GQA\); RoPE; SwiGLU; RMSNorm; vocab 15193618T tokens; web; code; math; synthetic; 29\+ languages; 1M\+ SFTPre\-training \+ DPO/GRPOOpenAI 2019GPT\-2 Large\[[42](https://arxiv.org/html/2608.17183#bib.bib38)\]774M1024Decoder\-only transformer; 36 layers; 20 attention heads; 1280 hidden; 5120 FFN; GELU; absolute position embeddings; vocab 50257WebText∼\\sim40GB from 8M web pages; cutoff Dec 2017Causal language modelingGoogle 2025Gemma3\-1B IT\[[16](https://arxiv.org/html/2608.17183#bib.bib43)\]1\.0B32768Decoder\-only;∼\\sim1\.1B active params; 32K context; text\-only input2T tokens; web; code; math; 140\+ languages; cutoff Aug 2024Pre\-training \+ instruction tuningMeta 2024MobileLLM\-1B\[[37](https://arxiv.org/html/2608.17183#bib.bib47)\]1\.0B–Deep thin decoder\-only; embedding sharing; grouped\-query attention; optional block\-wise weight\-sharing; sub\-billion optimizedTraining data not specified in paper; architecture\-focused designPre\-training \+ chat tuningAI2 2025OLMo2\-0425\-1B\[[4](https://arxiv.org/html/2608.17183#bib.bib54)\]1\.0B4096Transformer decoder \(OLMo2\); 16 layers; 16 heads; 2048 hidden; 8192 FFN; RoPE; vocab 100352OLMo\-mix\-1124 with Dolmino\-mix\-1124 mid\-trainingCausal language modelingTinyLlama\. 20231\.1B Chat v1\.0\[[51](https://arxiv.org/html/2608.17183#bib.bib52)\]1\.1B2048Llama\-style decoder\-only; 22 layers; 32 heads; 4 KV heads \(GQA\); 2048 hidden; 5632 FFN; RoPE; RMSNorm; vocab 32000SlimPajama \+ StarCoderData; aligned with UltraChat/UltraFeedbackCausal pre\-training \+ SFT \+ DPOZyphra 2024Zamba2\-1\.2B IT\[[61](https://arxiv.org/html/2608.17183#bib.bib53)\]1\.2B4096Hybrid SSM/transformer \(Mamba2 \+ shared attention blocks\); 38 layers; 32 attention heads; 2048 hidden; vocab 32000UltraChat\-200k, Infinity\-Instruct, UltraFeedback, preference dataSFT \+ DPO instruction tuningMeta 2024Llama3\.2\-1B IT\[[35](https://arxiv.org/html/2608.17183#bib.bib45)\]1\.23B131072Auto\-regressive decoder\-only transformer; GQA; shared embeddings; 128K context; 1\.23B paramsPublic online data; up to 9T tokens pretraining; cutoff Dec 2023SFT \+ RLHFOpenAI 2019GPT\-2 XL\[[44](https://arxiv.org/html/2608.17183#bib.bib39)\]1\.5B1024Decoder\-only transformer; 48 layers; 25 attention heads; 1600 hidden; 6400 FFN; GELU; absolute position embeddings; vocab 50257WebText∼\\sim40GB from 8M web pages; cutoff Dec 2017Causal language modelingDeepSeek 2024R1\-Qwen\-1\.5B\[[11](https://arxiv.org/html/2608.17183#bib.bib48)\]1\.5B131072Qwen2\.5\-Math\-1\.5B base: 28 layers; 1536 hidden; 12 heads 2 KV \(GQA\); RoPE; SiLU; RMSNorm; vocab 151936; 131K context∼\\sim800K SFT samples; distilled from DeepSeek\-R1 reasoning modelDistillation from DeepSeek\-R1 \(reasoning/CoT\)Alibaba 2024Qwen2\.5\-1\.5B IT\[[2](https://arxiv.org/html/2608.17183#bib.bib50)\]1\.5B131072Decoder\-only; 28 layers; 1536 hidden; 12 heads 2 KV \(GQA\); RoPE; SwiGLU; RMSNorm; vocab 15193618T tokens; web; code; math; synthetic; 29\+ languages; 1M\+ SFTPre\-training \+ DPO/GRPOStabAI 2024StableLM2\-1\.6B\[[49](https://arxiv.org/html/2608.17183#bib.bib60)\]1\.6B4096StableLM decoder\-only; 24 layers; 32 heads; 2048 hidden; 5632 FFN; partial RoPE; vocab 1003522T\-token multilingual and code\-heavy pre\-training mixCasual language modelingHF 2024SmolLM2\-1\.7B IT\[[20](https://arxiv.org/html/2608.17183#bib.bib57)\]1\.7B8192Transformer decoder \(Llama\-family\); 24 layers; 32 heads; 2048 hidden; vocab 4915211T\-token mix: FineWeb\-Edu, DCLM, The Stack, math/codePre\-training \+ SFT \+ DPOH2O 2024Danube2\-1\.8B Chat\[[18](https://arxiv.org/html/2608.17183#bib.bib55)\]1\.8B8192Llama2/Mistral\-style decoder\-only; 24 layers; 32 heads; 8 KV heads \(GQA\); 2560 hidden; SiLU; RMSNorm; vocab 32000H2O Danube2 pre\-training corpus; H2O LLM Studio chat dataPre\-training \+ SFT \+ DPOGoogle 2024Gemma2\-2B IT\[[15](https://arxiv.org/html/2608.17183#bib.bib44)\]2\.2B8192Decoder\-only; 26 layers; 2048 hidden; 16 attention heads; 4 key\-value heads; GQA; RoPE; RMSNorm; GELU2T tokens; 8K context during pre\-training; knowledge distillationKnowledge distillation \+ pre\-trainingAlibaba 2024Qwen2\.5\-3B IT\[[3](https://arxiv.org/html/2608.17183#bib.bib51)\]3\.0B32768Decoder\-only; 36 layers; 2048 hidden; 16 heads 2 KV \(GQA\); RoPE; SwiGLU; RMSNorm; vocab 15193618T tokens; web; code; math; synthetic; 29\+ languages; 1M\+ SFTPre\-training \+ DPO/GRPODbricks 2023Dolly\-v2\-3B\[[10](https://arxiv.org/html/2608.17183#bib.bib61)\]3\.0B2048GPT\-NeoX/Pythia\-style decoder\-only; derived from EleutherAI Pythia\-2\.8BDatabricks\-dolly\-15k instruction datasetInstruction tuning on PythiabaseMeta 2024Llama3\.2\-3B IT\[[36](https://arxiv.org/html/2608.17183#bib.bib46)\]3\.21B131072Auto\-regressive decoder\-only transformer; GQA; shared embeddingsPublic online data; up to 9T tokens pretraining; cutoff Dec 2023SFT \+ RLHFH2O 2024Danube3\.1\-4B Chat\[[19](https://arxiv.org/html/2608.17183#bib.bib56)\]4\.0B8192Llama\-style decoder\-only; 24 layers; 32 heads; 8 KV heads \(GQA\); 3840 hidden; SiLU; RMSNormH2O Danube3 pre\-training corpus; H2O LLM Studio chat dataPre\-training \+ chat fine\-tuning

Similar Articles

Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment

arXiv cs.AI

The paper benchmarks five large language models on multi-sensor physical hazard assessment, revealing that all tested models fail to produce precautionary warnings when multiple sensors are simultaneously elevated below their individual safety limits, while achieving near-perfect accuracy on single-sensor violations.

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

arXiv cs.AI

This paper audits four agent-safety benchmarks (R-Judge, InjecAgent, AgentHarm, AgentDojo) across many models, showing that their scores are confounded by capability, metrics have artifacts, and rankings disagree across benchmarks, undermining interchangeable safety claims.