Rank Portability 并不意味着 Feasibility Portability:针对特定目标的联合硬件约束评估
摘要
本文挑战了在跨设备硬件评估中,rank portability 意味着 feasibility portability 的假设,通过基准测试表明,高秩相关性在联合硬件约束下并不能保证安全部署决策。
arXiv:2609.22122v1 Announce Type: new
Abstract: Cross-device hardware evaluation often assumes that if architecture rankings transfer across devices, a proxy device can support target-side model selection. We stress-test this assumption for joint latency-energy feasibility across two public architecture families. On NAS-Bench-201, cross-device rank correlations are moderate, while target-comparable feasible-set overlap remains incomplete. A faithful AdaProxy diagnostic substantially improves latency ranking, showing that the observed boundary failures are not simply due to weak adaptation. Exact finite-sample split-conformal analysis also exposes an evidence bottleneck: a finite one-sided 90% threshold requires at least nine calibration observations. We then replicate the phenomenon on 10,000 GPT architectures across 13 HW-GPT-Bench devices. Relative to an RTX3080 proxy, target latency SRCC ranges from 0.951 to 0.996, yet proxy-reuse violation risk ranges from 33.3% to 100% under matched joint constraints. These results show that rank portability, feasibility portability, and target-specific decision support are distinct evaluation objects. Cross-device evaluations should therefore report which target environments actually support the operating point being claimed.
查看缓存全文
缓存时间: 2026/09/22 09:12
# Rank Portability Does Not Imply Feasibility Portability:Target-Specific Evaluation of Joint Hardware Constraints
Source: [https://arxiv.org/html/2609.22122](https://arxiv.org/html/2609.22122)
###### Abstract
Cross\-device hardware evaluation often relies on a premise that is useful for prediction but unsafe for decision making: if architecture rankings transfer across devices, then a proxy device should support target\-side model selection\. We stress\-test that premise for*joint*latency–energy feasibility\. The evaluation spans two public architecture families: 15,625 NAS\-Bench\-201 architectures with EdgeGPU, Eyeriss, and FPGA resource measurements, and 10,000 common GPT architectures in HW\-GPT\-Bench over 13 hardware devices\. First, on NAS\-Bench\-201, cross\-device latency and energy rank correlations are moderate, while target\-comparable feasible\-set overlap remains incomplete\. A faithful AdaProxy diagnostic substantially improves latency rank correlation \(median SRCC→0\.8590\.623\\\!\\to\\\!0\.859,→0\.9290\.721\\\!\\to\\\!0\.929, and→0\.9730\.873\\\!\\to\\\!0\.973\), demonstrating that the evaluation is not merely exposing a weak adaptation method\. Second, exact finite\-sample split\-conformal arithmetic reveals an evidence bottleneck hidden by empirical quantile heuristics: a finite one\-sided 90% threshold requires at least nine calibration observations, leaving only one fitting observation under a ten\-probe budget\. Third, HW\-GPT\-Bench provides a strong counterexample to equating rank portability with deployment support\. Relative to an RTX3080 proxy, target latency SRCC ranges from 0\.951 to 0\.996 across twelve targets, yet proxy\-reuse violation risk ranges from 33\.3% to 100% under target\-comparable joint constraints\. With valid 90% calibration, pooled target\-evidence risk remains 15\.8% at 20 probes and 14\.8% at 40 probes, and no common evidence budget through 40 meets the pre\-specified target\-wise criterion\. Together with a previously observed pooled\-versus\-target authorization reversal on NAS\-Bench\-201, these results show that rank correlation, pooled frontiers, and target\-specific decision support are distinct evaluation objects\. We argue that cross\-device evaluations should report*target\-specific support*: which hardware environments actually support the operating point being claimed\.
## 1Introduction
Hardware\-aware model selection increasingly depends on transferring measurements across devices\. A proxy device can be used to rank candidate architectures, a few target measurements can adapt a predictor, and uncertainty estimates can filter candidates before deployment\. These tools are valuable because exhaustive physical measurement is expensive\. OneProxy, for example, explicitly exploits cross\-device latency monotonicity and adapts a proxy predictor when rank correlation is weak\([Lu et al\. 2021](https://arxiv.org/html/2609.22122#bib.bib6)\)\. HELP and related few\-shot approaches instead measure the actual unseen target and adapt from those observations\([Lee et al\. 2021](https://arxiv.org/html/2609.22122#bib.bib7)\)\. Hardware\-aware benchmarks make these questions reproducible across thousands of architectures\([Li et al\. 2021](https://arxiv.org/html/2609.22122#bib.bib5);[Sukthanker et al\. 2024](https://arxiv.org/html/2609.22122#bib.bib11)\)\.
The evaluation problem changes when the output is not a latency estimate but a hard deployment decision\. Suppose a system must select the most capable architecture satisfying both latency and energy budgets\. The relevant feasible set on target hardwarehhis
𝒞h\(B\)=\{a:ℓh\(a\)≤Bℓ,eh\(a\)≤Be\}\.\\mathcal\{C\}\_\{h\}\(B\)=\\\{a:\\ell\_\{h\}\(a\)\\leq B\_\{\\ell\},\\;e\_\{h\}\(a\)\\leq B\_\{e\}\\\}\.\(1\)A candidate outside either coordinate is infeasible\. A high rank correlation can still be compatible with substantial disagreement near this two\-dimensional boundary, especially when selection chooses an extreme candidate from a large search space\. Prior work already shows that hardware\-predictor errors can propagate into architecture selection\([Laube et al\. 2022](https://arxiv.org/html/2609.22122#bib.bib10)\)\. Our question is more specific:*what evidence actually supports a target\-specific feasibility decision?*
This paper is an evaluation\-methodology and stress\-testing study\. It does not propose a new NAS algorithm and it does not claim that cross\-device adaptation is ineffective\. Instead, it separates three evaluation objects that are often conflated:
1. 1\.Rank portability:do proxy and target preserve architecture ordering?
2. 2\.Feasibility portability:do they agree on membership in a hard joint resource set?
3. 3\.Decision support:does the evidence justify the selected architecture on each target, rather than only on average after pooling targets?
We make four contributions\. First, we add explicit rank and feasible\-set diagnostics to the NAS\-Bench\-201 hardware study and show that feasible\-set overlap is substantially weaker than rank portability\. Second, we reproduce the public OneProxy/AdaProxy adaptation structure in its legitimate latency\-ranking domain\. Adaptation works: it markedly raises SRCC, which makes the subsequent boundary failures more informative, not less\. Third, we replace “conformal\-style” empirical residual quantiles with the exact finite\-sample order\-statistic rule and expose the resulting evidence\-budget arithmetic\. Fourth, we independently replicate the evaluation phenomenon in HW\-GPT\-Bench, a different architecture family with 10,000 common GPT architectures across 13 devices\. There, latency rankings are extraordinarily portable, yet source\-only feasibility decisions remain unreliable\.
The strongest result is therefore not “hardware differs\.” It is an evaluation reversal:*evidence can be strong for ranking while weak for authorization*\. In the GPT replication, every one of twelve targets has latency SRCC above 0\.95 relative to the RTX3080 proxy, yet proxy\-reuse target violation risk is at least 33\.3% and reaches 100%\. In the NAS\-Bench\-201 study, a pooled 40\-probe risk of 9\.06% appears to satisfy a 10% threshold while EdgeGPU and FPGA individually fail it at 12\.03% and 15\.00%\. These are different failure modes with the same methodological implication: report target\-specific support, not only aggregate performance\.
## 2Related work and claim boundary
#### Hardware\-aware architecture evaluation\.
Hardware\-aware NAS incorporates deployment cost directly into model selection\([Tan et al\. 2019](https://arxiv.org/html/2609.22122#bib.bib1);[Cai et al\. 2019](https://arxiv.org/html/2609.22122#bib.bib2);[Cai et al\. 2020](https://arxiv.org/html/2609.22122#bib.bib3)\)\. NAS\-Bench\-201 and HW\-NAS\-Bench make architecture and hardware evaluation reproducible in finite search spaces\([Dong and Yang 2020](https://arxiv.org/html/2609.22122#bib.bib4);[Li et al\. 2021](https://arxiv.org/html/2609.22122#bib.bib5)\)\. HW\-GPT\-Bench extends hardware\-aware benchmarking to GPT\-family architectures and exposes latency and energy across 13 devices\([Sukthanker et al\. 2024](https://arxiv.org/html/2609.22122#bib.bib11)\)\. Our contribution is not another benchmark; it is a stress test of how cross\-device benchmark evidence is aggregated into deployment claims\.
#### Proxy transfer and target adaptation\.
OneProxy studies cross\-device latency rank monotonicity\. When monotonicity is high, the searched result on a proxy may transfer; when it is low, AdaProxy adapts the proxy predictor using target observations\([Lu et al\. 2021](https://arxiv.org/html/2609.22122#bib.bib6)\)\. We therefore use OneProxy as the correct motivation for the rank\-portability question, not as a straw\-man safety certificate\. HELP, Multi\-Predict, and recent small\-budget search methods use actual target\-device evidence and belong on the target\-evidence side of the comparison\([Lee et al\. 2021](https://arxiv.org/html/2609.22122#bib.bib7);[Akhauri and Abdelfattah 2023](https://arxiv.org/html/2609.22122#bib.bib8);[Capuano et al\. 2025](https://arxiv.org/html/2609.22122#bib.bib9)\)\.
#### Uncertainty and selection\.
Conformal prediction provides finite\-sample marginal coverage under exchangeability when its order\-statistic rule is implemented correctly\([Vovk et al\. 2005](https://arxiv.org/html/2609.22122#bib.bib12);[Angelopoulos and Bates 2023](https://arxiv.org/html/2609.22122#bib.bib13)\)\. Selection complicates interpretation because the candidate of interest is chosen after screening many alternatives; selective prediction and post\-selection inference make this distinction explicit\([Geifman and El\-Yaniv 2019](https://arxiv.org/html/2609.22122#bib.bib14);[Jin and Ren 2024](https://arxiv.org/html/2609.22122#bib.bib15)\)\. We use conformal calibration only as an evaluation baseline and do not claim a new conformal theorem\.
## 3Evaluation object and protocol
### 3\.1Metrics
For a decision cell, letA∈\{0,1\}A\\in\\\{0,1\\\}indicate whether the method admits a selected architecture andU∈\{0,1\}U\\in\\\{0,1\\\}indicate whether that admitted architecture violates at least one target resource bound\. We report
Cov=𝔼\[A\],Risk=𝔼\[U∣A=1\],\\operatorname\{Cov\}=\\mathbb\{E\}\[A\],\\qquad\\operatorname\{Risk\}=\\mathbb\{E\}\[U\\mid A=1\],\(2\)and, where capability is available, normalized capability regret relative to the best target\-feasible architecture\. Direct verification has zero lookup\-table violation by construction; its empirical cost is therefore coverage and capability opportunity cost, not “discovered safety\.”
### 3\.2Target\-comparable feasibility regimes
Absolute latency and energy scales differ dramatically across hardware\. To avoid making one target artificially easy, we construct target\-specific budgets at oracle feasible\-set densities 0\.2, 0\.4, and 0\.6 under balanced, latency\-tight, and energy\-tight profiles\. Each profile rescales the target’s median latency and energy and then selects one common scalar threshold so the joint feasible fraction matches the requested density\. This preserves the non\-compensatory decision in Eq\. \([1](https://arxiv.org/html/2609.22122#S1.E1)\) while making difficulty comparable\.
### 3\.3NAS\-Bench\-201 / HW\-NAS\-Bench arm
The first arm uses all 15,625 NAS\-Bench\-201 architectures with paired latency and energy for EdgeGPU, Eyeriss, and FPGA in HW\-NAS\-Bench\. We evaluate all six ordered source–target pairs\. For each pair we report latency SRCC, energy SRCC, a normalized joint\-load SRCC, and target\-comparable feasible\-set Jaccard overlap\. We also run a faithful AdaProxy latency diagnostic: Pixel3 is the proxy; the public architecture encoding and Eq\. \(3\)\-style scaling\-plus\-sparse\-residual optimization are used; target train/validation counts follow the public NAS\-Bench\-201 setup; and the regularization parameter is selected on target\-validation SRCC\. This diagnostic asks whether the known proxy\-adaptation mechanism works before we discuss joint feasibility\.
### 3\.4Finite\-sample target evidence
For one\-sided split conformal withnncalibration scores and target miscoverageα\\alpha, we use the⌈\(n\+1\)\(1−α\)⌉\\lceil\(n\+1\)\(1\-\\alpha\)\\rceil\-th order statistic of the calibration scores augmented with\+∞\+\\infty\. We never clip an unattainable rank to the largest finite score\. Therefore a finite 90% upper threshold requires at least nine calibration observations\. At a total budget of ten measurements, exact 90% split calibration leaves one fitting observation; at five measurements the finite 90% threshold is unavailable\.
### 3\.5Independent HW\-GPT\-Bench replication
The second arm uses the official HW\-GPT\-Bench GPT\-small stored ground\-truth statistics: sampled latency, sampled energy, and perplexity\. The strict common\-architecture intersection contains 10,000 architectures\. RTX3080 is the proxy and the remaining twelve devices are targets: P100, A100, A40, A6000, H100, RTX2080, V100, three AMD EPYC CPUs, and two Xeon CPUs\. We repeat the same 0\.2/0\.4/0\.6 feasible\-density regimes and three joint\-constraint profiles\. Source\-only proxy reuse selects the best\-perplexity proxy\-feasible architecture\. Target\-evidence mapping uses total budgets 10, 20, and 40, reserves nine observations for exact 90% calibration, and uses the remaining observations to fit a log resource map from proxy to target\. Twenty pre\-specified seeds vary the target probes\.
## 4Results
### 4\.1Feasible\-set portability is weaker than rank portability
Table[1](https://arxiv.org/html/2609.22122#S4.T1)reports all ordered NAS\-Bench\-201 hardware pairs\. Latency SRCC ranges from 0\.506 to 0\.754 and energy SRCC from 0\.539 to 0\.828, yet median target\-comparable feasible\-set Jaccard ranges only from 0\.483 to 0\.626, with minima as low as 0\.244\. The evaluation object therefore changes before any predictor is fit: preserving global order is easier than preserving membership near a hard joint boundary\.
Table 1:Cross\-device portability on NAS\-Bench\-201\. Jaccard is computed over target\-comparable joint latency–energy feasible sets across the pre\-specified density/profile regimes\.
### 4\.2A faithful proxy adaptation improves ranking
A negative evaluation of feasibility transfer would be weak if it merely used a poor target\-adaptation implementation\. Figure[1](https://arxiv.org/html/2609.22122#S4.F1)addresses that concern\. Relative to the Pixel3 proxy, faithful AdaProxy adaptation raises median target latency SRCC from 0\.623 to 0\.859 on EdgeGPU, 0\.721 to 0\.929 on Eyeriss, and 0\.873 to 0\.973 on FPGA\. All three improve; two exceed 0\.90\. Thus our claim is not that proxy adaptation fails\. Rather, strong rank adaptation is compatible with a separate limitation at the joint feasibility boundary\.
Figure 1:Faithful AdaProxy diagnostic\. The target\-adaptation mechanism substantially improves latency ranking, establishing that the later feasibility failures are not explained by refusing to adapt the proxy\.
### 4\.3Finite\-sample calibration creates a real evidence bottleneck
Table[2](https://arxiv.org/html/2609.22122#S4.T2)shows the arithmetic hidden by a small empirical residual quantile\. At 90% nominal coverage, nine calibration points are the minimum for a finite one\-sided split\-conformal threshold\. Consequently,k=10k=10permits only one fitting observation, whilek=5k=5cannot produce a finite 90% threshold at all\. At 95%, the corresponding minimum is 19 calibration observations\. This is not an implementation detail: the guarantee itself consumes the measurement budget\.
Table 2:Exact finite\-sample split\-conformal budget arithmetic\. “Fit” is the maximum remaining target observations after reserving the minimum calibration set\.Nominal coverageTotalkkMin\. calibrationFitFinite threshold?90%590No90%1091Yes90%20911Yes90%40931Yes95%10190No95%20191Yes
### 4\.4Independent replication: nearly perfect rankings, unreliable feasibility
HW\-GPT\-Bench creates the sharper test\. Across twelve targets, latency SRCC with the RTX3080 proxy is between 0\.951 and 0\.996\. Nevertheless, the proxy\-feasible best\-perplexity selection violates the target’s matched joint constraint in 33\.3–100% of evaluation regimes \(Figure[2](https://arxiv.org/html/2609.22122#S4.F2)\)\. The contrast is strongest on A6000: latency SRCC is 0\.996, yet target violation risk is 66\.7%\. A40 has SRCC 0\.992 and 100% violation; P100 has SRCC 0\.990 and 100% violation\. High rank portability is therefore not sufficient evidence for hard\-boundary portability\.
Figure 2:Independent HW\-GPT\-Bench replication\. Every target has latency SRCC above 0\.95 with the RTX3080 proxy, yet source\-only proxy reuse has substantial target violation risk under matched joint latency–energy constraints\. Dashed lines show SRCC 0\.90 and risk 0\.10 only as visual references\.The target\-evidence result is also unfavorable to a universal small\-kkclaim\. Under exact 90% calibration, pooled coverage/risk are 0\.37%/100% atk=10k=10, 88\.6%/15\.8% atk=20k=20, and 90\.5%/14\.8% atk=40k=40\(Figure[3](https://arxiv.org/html/2609.22122#S4.F3)\)\. Thek=10k=10failure is expected from the finite\-sample arithmetic: only one point remains for fitting\. More importantly, increasing to 20 or 40 restores coverage but does not reduce pooled conditional violation below 10%\. The pre\-specified audit finds nok≤40k\\leq 40that simultaneously supports the target\-wise criterion on all hardware environments, and it flags a material target\-specific support gap of at least 0\.20 coverage units\.
Figure 3:HW\-GPT\-Bench pooled target\-evidence frontier using an exact 90% split\-conformal threshold\. Coverage recovers with more fitting data, but unsafe\-among\-admitted risk remains above 10% through 40 probes\. Target\-specific support is reported separately in the supplement\.
### 4\.5Pooling can change the decision
The independent GPT replication does*not*produce a pooled\-below\-threshold reversal; its pooled risk remains above 10%\. The earlier NAS\-Bench\-201 closure does, which is useful precisely because the two benchmark families expose different failure modes\. Atk=40k=40, a target\-ridge uncertainty rule has pooled conditional risk 9\.06%, apparently passing a 10% criterion, while EdgeGPU and FPGA have target\-specific risks 12\.03% and 15\.00% \(Eyeriss: 0%\)\. Pooling would authorize two targets that their own evidence rejects\. The new GPT study independently confirms the broader support problem: the pre\-specified audit finds a material difference between aggregate and common target support, even though the particular pooled\-risk reversal is absent\.
Table 3:Cross\-benchmark stress\-test summary\. The two benchmark families expose different manifestations of the same evaluation problem rather than duplicating one artifact\.
## 5What should hardware evaluations report?
The experiments support four reporting rules\.
#### Separate ranking from boundary validity\.
Rank metrics such as SRCC answer whether architectures are ordered similarly\. They do not answer whether a selected architecture belongs to the target’s feasible set\. The HW\-GPT result makes this distinction unusually clear because correlations above 0\.95 coexist with large hard\-boundary failure\.
#### Report finite\-sample evidence arithmetic\.
A target\-probe budget is not fully available for fitting if part of it is reserved for a valid uncertainty statement\. For split conformal, the calibration sample required by the desired nominal coverage should be shown explicitly\. A heuristic empirical quantile with four calibration points should not be labeled as if it supplied a conventional 90% finite\-sample threshold\.
#### Report target\-specific support next to pooled frontiers\.
A pooled risk–coverage curve is descriptive\. It becomes decision\-relevant only when the target of deployment is sampled from the same mixture and the mixture\-level claim is actually the intended object\. When the claim is “this hardware target satisfies risk≤ϵ\\leq\\epsilon,” target\-specific support is required\. We recommend reporting the range of realizable coverage, the conditional risk at nearby operating points, and whether a common support interval exists across targets\.
#### Treat direct verification as an opportunity\-cost baseline\.
If a candidate is measured on the target before admission, zero lookup\-table violation is constructional\. The question is how much search opportunity was spent to obtain that evidence\. In the NAS\-Bench\-201 closure, active direct verification increased pooled coverage from 73\.5% at five probes to 99\.6% at forty while reducing normalized capability regret from 10\.69% to 2\.54%\. This is a useful evidence\-sufficiency frontier, not a claim that verification “discovers” zero risk\.
## 6Limitations
The two benchmark families differ substantially, which is a strength for replication but prevents a perfectly symmetric experiment\. HW\-NAS\-Bench provides a finite NAS\-Bench\-201 space with EdgeGPU, Eyeriss, and FPGA latency/energy values; HW\-GPT\-Bench provides GPT\-family ground\-truth sampled statistics across many GPU and CPU targets\. Both are lookup\-table evaluations rather than contemporaneous physical deployment measurements, so they condition on the benchmark’s measurement process and do not model thermal drift, firmware changes, repeated\-measurement noise, or runtime interference\.
The HW\-GPT target\-evidence baseline is deliberately simple: a log resource map from proxy to target followed by exact split\-conformal multiplicative calibration\. It is not a claim to outperform HELP, Multi\-Predict, or specialized active hardware search\. Its purpose is to expose the information accounting under a fixed total probe budget\. The faithful AdaProxy experiment is a latency\-rank diagnostic, matching the method’s legitimate objective; we do not extend its original guarantee to joint latency–energy admission\.
The NAS\-Bench\-201 capability\-regret results use aligned CIFAR\-100 capability metadata derived through a public development pipeline\. The final archival package should retain exact provenance and should not redistribute third\-party artifacts whose license is unclear\. The new HW\-GPT replication uses the official repository’s stored ground\-truth statistics and its perplexity field\. Statistical variability in the closure is over pre\-specified probe seeds conditional on the fixed benchmark; it is not population confidence over all possible hardware devices\.
## 7Conclusion
Cross\-device hardware evaluation needs a sharper vocabulary\. Rank portability, feasibility portability, and target\-specific decision support are not interchangeable\. A faithful proxy adaptation can substantially improve ranking; exact finite\-sample calibration can still consume most of a small evidence budget; and, in a second architecture family, latency rankings above 0\.95 can coexist with 33–100% target violation under hard joint constraints\. Pooled summaries add another failure mode: an aggregate point can satisfy a risk threshold while individual targets do not\. The practical recommendation is simple: when an evaluation is used to justify deployment, report the*target\-specific support*of the operating point being claimed\. Aggregate frontiers and proxy correlations remain useful diagnostics, but they should not be treated as authorization evidence by themselves\.
## References
- Tan et al\. \(2019\)Mingxing Tan et al\. MnasNet: Platform\-Aware Neural Architecture Search for Mobile\.*CVPR*, 2019\.
- Cai et al\. \(2019\)Han Cai, Ligeng Zhu, and Song Han\. ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware\.*ICLR*, 2019\.
- Cai et al\. \(2020\)Han Cai et al\. Once\-for\-All: Train One Network and Specialize it for Efficient Deployment\.*ICLR*, 2020\.
- Dong and Yang \(2020\)Xuanyi Dong and Yi Yang\. NAS\-Bench\-201: Extending the Scope of Reproducible Neural Architecture Search\.*ICLR*, 2020\.
- Li et al\. \(2021\)Chaojian Li et al\. HW\-NAS\-Bench: Hardware\-Aware Neural Architecture Search Benchmark\.*ICLR*, 2021\.
- Lu et al\. \(2021\)Bingqian Lu, Jianyi Yang, Weiwen Jiang, Yiyu Shi, and Shaolei Ren\. One Proxy Device Is Enough for Hardware\-Aware Neural Architecture Search\.*Proceedings of the ACM on Measurement and Analysis of Computing Systems*, 5\(3\), 2021\.
- Lee et al\. \(2021\)Hayeon Lee et al\. HELP: Hardware\-Adaptive Efficient Latency Prediction for NAS via Meta\-Learning\.*NeurIPS*, 2021\.
- Akhauri and Abdelfattah \(2023\)Yash Akhauri and Mohamed S\. Abdelfattah\. Multi\-Predict: Few Shot Predictors For Efficient Neural Architecture Search\. 2023\.
- Capuano et al\. \(2025\)Francesco Capuano et al\. Searching on a Budget: Hardware\-Aware Neural Architecture Search with 10 Latency Probes\. 2025\.
- Laube et al\. \(2022\)Kevin Alexander Laube et al\. What to Expect of Hardware Metric Predictors in NAS\.*AutoML Conference*, 2022\.
- Sukthanker et al\. \(2024\)Rhea Sanjay Sukthanker et al\. HW\-GPT\-Bench: Hardware\-Aware Architecture Benchmark for Language Models\.*NeurIPS Datasets and Benchmarks Track*, 2024\.
- Vovk et al\. \(2005\)Vladimir Vovk, Alex Gammerman, and Glenn Shafer\.*Algorithmic Learning in a Random World*\. Springer, 2005\.
- Angelopoulos and Bates \(2023\)Anastasios N\. Angelopoulos and Stephen Bates\. Conformal Prediction: A Gentle Introduction\.*Foundations and Trends in Machine Learning*, 2023\.
- Geifman and El\-Yaniv \(2019\)Yonatan Geifman and Ran El\-Yaniv\. SelectiveNet: A Deep Neural Network with an Integrated Reject Option\.*ICML*, 2019\.
- Jin and Ren \(2024\)Ying Jin and Zhimei Ren\. Confidence on the Focal: Conformal Prediction with Selection\-Conditional Coverage\. 2024\.
## Appendix AClaim\-scope proposition
###### Proposition 1\(No source\-only distribution\-free target certificate without a linking assumption\)\.
Fix a source resource map\(ℓs,es\)\(\\ell\_\{s\},e\_\{s\}\)and an algorithm that observes only source\-side information before admitting an architecture\. Without an assumption restricting the relationship between source and target resource maps, there is no nontrivial distribution\-free guarantee that every admitted architecture belongs to𝒞t\(B\)\\mathcal\{C\}\_\{t\}\(B\)\.
#### Proof\.
Take any source\-side observation for which the algorithm admits some architectureaa\. Construct two target worlds that are identical on all source observations\. In world 1, set\(ℓt\(a\),et\(a\)\)\(\\ell\_\{t\}\(a\),e\_\{t\}\(a\)\)insideBB; in world 2, change one target coordinate ofaato exceed its budget while leaving all source information unchanged\. The source\-only algorithm makes the same decision in both worlds, so it cannot certify target feasibility in both\. A nontrivial target guarantee therefore requires either target evidence or a structural assumption linking source and target resource maps\.□\\square
## Appendix BExecuted scientific closure summary
The pre\-specified closure was executed after one solver\-only repair: HW\-NAS latency was converted from milliseconds to seconds as in the public OneProxy notebook, the unregularized endpoint was solved by least squares, and positive regularization retained the same convex objective with a deterministic solver fallback\. No scientific thresholds, seeds, hardware targets, or decision criteria were changed after observing results\.
The final verifier confirmed that the faithful AdaProxy diagnostic, finite\-sample calibration audit, and 13\-device independent replication all executed successfully\. Three pre\-specified evaluation\-failure diagnostics were positive: high HW\-GPT rank portability with hard\-boundary failure; a material HW\-GPT target\-support gap; and material NAS\-Bench\-201 joint feasible\-set membership mismatch\. The separate HW\-GPT pooled\-authorization\-reversal signal was false; the pooled\-versus\-target reversal reported in the main text comes from the earlier NAS\-Bench\-201 closure\.
## Appendix CHW\-GPT proxy\-reuse table
Table 4:RTX3080 proxy portability on the 12 HW\-GPT targets\. Risk is the fraction of the nine matched joint\-constraint regimes in which the selected proxy\-feasible architecture violates the target constraint\.
## Appendix DReproducibility and assets
The supplementary runner pins the public HW\-NAS\-Bench blob, downloads the official HW\-GPT\-Bench GPT\-small statistics blob, verifies its Git blob SHA, and freezes the target/proxy roles, target\-comparable density regimes, profiles, probe budgets, seeds, and the pre\-specified decision rule\. HW\-NAS\-Bench is distributed under the MIT License; HW\-GPT\-Bench is distributed under Apache\-2\.0\. The supplement should include the executed result CSVs together with the runner before archival submission\. The generated row count is not treated as an independent statistical sample size\.相似文章
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations
This paper applies Generalizability Theory to agent benchmarks, showing leaderboards rank specialization rather than capability, and proposes a framework (DDR) for sizing reliable deployment evaluations.
安全,还是仅仅是能力?对智能体安全基准的效度审计
本文对四个智能体安全基准(R-Judge、InjecAgent、AgentHarm、AgentDojo)在多个模型上进行了审计,表明它们的分数受能力混淆影响,指标存在伪影,且不同基准之间的排名相互矛盾,从而削弱了可互换的安全性声明。
分区分数不是系统分数:分解算法选择中的部署保真度差距
本文介绍了分解算法选择中的部署保真度差距,证明了分区级评估可能与端到端系统性能不同,并对报告和基准测试有影响。
不是能力问题:LLM智能体层级间的控制敏感度是非单调的
本文通过实证测试了“更结构化的控制(harness)能普遍提高LLM智能体可靠性”这一常见假设,发现不同模型层级间存在非单调关系。它引入了HEAT-24基准,并揭示了严格的控制可能会损害前沿聊天模型,但有利于推理模型。
@pvncher: 虽然你说得完全正确,这些路由器确实没有意义,但我已经对运行那些仅仅……的基准测试完全失去了兴趣。
作者讨论了OpenRouter推出的缓存感知模型路由器,该路由器针对质量、速度和成本优化模型选择,同时批评了那些孤立评估AI工具而非真实场景的基准测试。