UO-FIE: 结合精确标签监督与分级效用的事实性推理

arXiv cs.LG 论文

摘要

UO-FIE是一个参数高效的事实性推理系统,结合精确标签监督与分级效用,在FIE2026的微调赛道中以0.8316的宏效用获得第一名。

arXiv:2609.28605v1 Announce Type: new Abstract: The Factivity Inference Evaluation 2026 (FIE2026) classifies Chinese context-hypothesis pairs into nine ordered factivity intervals. Its evaluation metric rewards both exact predictions and proximity to the correct interval, while 64.1% of the 566 training examples belong to a single class. In preliminary experiments, several mDeBERTa classification models predominantly predict the dominant class, whereas a Huber-regression baseline produces more predictions near the correct interval but fewer exact matches. We introduce Utility-Oriented Factivity Inference (UO-FIE), a parameter-efficient system that combines exact-label supervision with graded utility. UO-FIE predicts a distribution over the nine classes and combines hard-label supervision, utility-based soft targets, scheduled class weights, and an ordinal loss. We evaluate expected-utility decoding in controlled comparisons and use ordinal calibration selected on out-of-fold predictions for the submitted system. Based on Qwen3.5-9B with LoRA, UO-FIE ranks first in the fine-tuning track with a macro utility of 0.8316. A separate prompt-based ensemble ranks third in the non-fine-tuning track with a macro utility of 0.8450.
查看原文
查看缓存全文

缓存时间: 2026/09/25 09:31

# Combining Exact-Label Supervision withGraded Utility for Factivity Inference
Source: [https://arxiv.org/html/2609.28605](https://arxiv.org/html/2609.28605)
###### Abstract

The Factivity Inference Evaluation 2026 \(FIE2026\) classifies Chinese context–hypothesis pairs into nine ordered factivity intervals\. Its evaluation metric rewards both exact predictions and proximity to the correct interval, while 64\.1% of the 566 training examples belong to a single class\. In preliminary experiments, several mDeBERTa classification models predominantly predict the dominant class, whereas a Huber\-regression baseline produces more predictions near the correct interval but fewer exact matches\.

We introduce Utility\-Oriented Factivity Inference \(UO\-FIE\), a parameter\-efficient system that combines exact\-label supervision with graded utility\. UO\-FIE predicts a distribution over the nine classes and combines hard\-label supervision, utility\-based soft targets, scheduled class weights, and an ordinal loss\. We evaluate expected\-utility decoding in controlled comparisons and use ordinal calibration selected on out\-of\-fold predictions for the submitted system\.

Based on Qwen3\.5\-9B with LoRA, UO\-FIE ranks first in the fine\-tuning track with a macro utility of 0\.8316\. A separate prompt\-based ensemble ranks third in the non\-fine\-tuning track with a macro utility of 0\.8450\.

Keywords:factivity inference, ordinal classification, graded utility, utility\-aligned learning, parameter\-efficient fine\-tuning

## 1Introduction

Factivity inference asks whether a linguistic context commits to the truth of an event, often through factive or counter\-factive predicates, negation, modality, and reported speech\. Prior work includes FactBank\[[13](https://arxiv.org/html/2609.28605#bib.bib5)\]and neural models of lexicosyntactic inference\[[15](https://arxiv.org/html/2609.28605#bib.bib6)\]\. FIE extends the problem to Chinese and asks systems to express both truth status and judgment strength\[[2](https://arxiv.org/html/2609.28605#bib.bib1)\]\.

FIE2026 adds a distance\-sensitive graded utility to this linguistic problem\[[5](https://arxiv.org/html/2609.28605#bib.bib2)\]\. Predictions occupy one of nine intervals, from strong counter\-factive to strong factive\. Exact intervals matter, but errors at distance one or two retain partial credit; more distant errors receive zero\. Ordinary nine\-class accuracy ignores this graded structure, whereas a smooth scalar objective can reduce exact\-label accuracy\. The 566 training labels are also sharply skewed: 363 occupy the strongest factive interval, while several intermediate intervals have fewer than ten examples\.

Preliminary experiments reveal two recurring empirical patterns \(Table[1](https://arxiv.org/html/2609.28605#S5.T1)\)\. Because the dominant label and its neighborhood cover much of the data, a majority predictor already obtains 0\.7652 micro utility, and several preliminary categorical mDeBERTa runs converge to the same pattern\. Huber regression raises the within\-two rate, the fraction of predictions at distance at most two, from 0\.7827 to 0\.8145, but lowers exact accuracy from 0\.6413 to 0\.1325\. The first pattern predicts mainly the majority interval; the second produces more nearby predictions but fewer exact matches\.

These observations favor a categorical model that preserves all nine labels while accounting for the distances between them\. UO\-FIE uses the official score matrix for this purpose and adapts Qwen3\.5\-9B with LoRA\[[10](https://arxiv.org/html/2609.28605#bib.bib10)\]\. The backbone estimates a nine\-label distribution; an OOF\-calibrated ordinal layer sets decision thresholds to account for class imbalance and task utility\.

We report systems for both official resource settings\. UO\-FIE is the primary fine\-tuning method, while prompt fusion provides an independent secondary submission\.

We make three contributions:

- •We characterize two empirical patterns: majority\-label collapse and regression toward nearby but inexact intervals\.
- •We introduce score\-matrix\-guided categorical learning that retains exact\-label supervision while representing neighborhood utility and ordinal structure\.
- •We provide controlled OOF evidence for utility\-aware training and decoder selection; the complete system ranks first in the official fine\-tuning track\.

## 2Related Work

Event factuality models infer commitment from lexical, syntactic, and discourse evidence\[[13](https://arxiv.org/html/2609.28605#bib.bib5),[15](https://arxiv.org/html/2609.28605#bib.bib6),[16](https://arxiv.org/html/2609.28605#bib.bib15)\]\. FIE2025 established the Chinese shared task\[[2](https://arxiv.org/html/2609.28605#bib.bib1)\], with systems based on fine\-tuning, ensembling, and prompt arbitration\[[8](https://arxiv.org/html/2609.28605#bib.bib3),[11](https://arxiv.org/html/2609.28605#bib.bib4)\]\. Other CCL evaluation reports have explored retrieval\-augmented inference and multi\-round voting for fine\-grained Chinese prediction\[[14](https://arxiv.org/html/2609.28605#bib.bib16)\]\. FIE2026 adds confidence\-bearing intervals and banded utility\[[5](https://arxiv.org/html/2609.28605#bib.bib2)\]\.

Cost\-sensitive classification represents unequal mistakes with a cost matrix\[[4](https://arxiv.org/html/2609.28605#bib.bib13)\]; ordinal and label\-distribution methods encode rank or neighborhood structure\[[1](https://arxiv.org/html/2609.28605#bib.bib7),[6](https://arxiv.org/html/2609.28605#bib.bib8),[3](https://arxiv.org/html/2609.28605#bib.bib9)\]; and decision theory selects actions under application loss\[[7](https://arxiv.org/html/2609.28605#bib.bib14)\]\. UO\-FIE combines these perspectives through FIE’s score matrix\.

## 3Task Formulation

For a Chinese context–hypothesis pairx=\(c,h\)x=\(c,h\), a system predicts a factivity category and confidence\. Evaluation maps the gold and predicted judgments to ordered valuesy,y^∈\{0,…,8\}y,\\hat\{y\}\\in\\\{0,\\ldots,8\\\}, which we call*interval labels*\. Labels 0–3 represent four counter\-factive strengths, label 4 is non\-factive, and labels 5–8 represent four factive strengths\.

Letd=\|y−y^\|d=\|y\-\\hat\{y\}\|\. The instance utility is

Score⁡\(y,y^\)=\{1\.0000d=0,0\.9545d=1,0\.6827d=2,0d≥3\.\\operatorname\{Score\}\(y,\\hat\{y\}\)=\\begin\{cases\}1\.0000&d=0,\\\\ 0\.9545&d=1,\\\\ 0\.6827&d=2,\\\\ 0&d\\geq 3\.\\end\{cases\}\(1\)Equation \([1](https://arxiv.org/html/2609.28605#S3.E1)\) induces a symmetric score matrixSS, whereSi,jS\_\{i,j\}is the utility of predicting intervaljjfor gold intervalii\. The official ranking metric first averages utility within each gold class and then averages across the nine classes:

Mmacro=19∑c=081Nc∑n:yn=cSc,y^n\.M\_\{\\mathrm\{macro\}\}=\\frac\{1\}\{9\}\\sum\_\{c=0\}^\{8\}\\frac\{1\}\{N\_\{c\}\}\\sum\_\{n:y\_\{n\}=c\}S\_\{c,\\hat\{y\}\_\{n\}\}\.\(2\)The organizer also reports an instance\-level micro average as a reference statistic, but rankings are determined byMmacroM\_\{\\mathrm\{macro\}\}\. The diagonal ofSSrewards exact predictions, the first two off\-diagonal bands encode partial credit, and all more distant decisions receive zero\. This matrix connects the task definition to UO\-FIE\.

Figure 1:The two structures that motivate our design\. \(a\) Exact predictions receive full credit, nearby errors retain graded utility, and errors at distance three or more receive zero\. \(b\) The 566 training labels are highly imbalanced: interval 8 contains 363 examples, whereas several intermediate intervals contain fewer than ten\.Task input\(c,h\)\(c,h\)Qwen3\.5\-9BLoRA \+ final\-token poolingPredictive distributionpθ=softmax⁡\(z\)p\_\{\\theta\}=\\mathrm\{softmax\}\(z\)Decision alternativesanalysis: utility rulesubmission: calibrated ruleOrdered intervaly^∈\{0,…,8\}\\hat\{y\}\\in\\\{0,\\ldots,8\\\}Goldyy\+ score matrixSSscore\-shaped targetqyq\_\{y\}Training objectiveshard \+ score\-aware \+ CDF

Figure 2:Overview of UO\-FIE\. Hard supervision and the score matrix shape a categorical distribution over exact labels and graded alternatives\. Two decision alternatives operate on the same distribution: analysis uses expected utility, while the official submission uses prior correction and utility\-selected ordinal calibration\.
## 4Model Description: UO\-FIE

Figure[2](https://arxiv.org/html/2609.28605#S3.F2)summarizes the fine\-tuning pipeline\. Both decoders use the same predicted distribution: development comparisons use expected\-utility decoding, and the submitted system uses ordinal calibration\.

### 4\.1Utility\-Aligned Training

#### Keeping exact labels\.

We serialize the nine\-class rubric, context, and hypothesis into one discriminative input\. The fields follow this fixed order and are separated by blank lines\. Qwen3\.5\-9B is used as the backbone\[[12](https://arxiv.org/html/2609.28605#bib.bib12)\]; the preceding Qwen3 family is described by?\)\. After left truncation to 1,024 tokens, the final non\-padding hidden state passes through a linear nine\-class head\. The resulting predictive distribution ispθ=softmax⁡\(z\)p\_\{\\theta\}=\\operatorname\{softmax\}\(z\)\. Only the classifier and LoRA adapters are trained; the backbone parameters remain frozen\.

#### Turning scores into neighborhoods\.

LetS∈ℝ9×9S\\in\\mathbb\{R\}^\{9\\times 9\}contain the utilities from Equation \([1](https://arxiv.org/html/2609.28605#S3.E1)\)\. For gold classyy, we derive a soft target

qy,k=exp⁡\(Sy,k/Ts\)∑jexp⁡\(Sy,j/Ts\),q\_\{y,k\}=\\frac\{\\exp\(S\_\{y,k\}/T\_\{s\}\)\}\{\\sum\_\{j\}\\exp\(S\_\{y,j\}/T\_\{s\}\)\},\(3\)whereTs=0\.18T\_\{s\}=0\.18controls target sharpness\. Unlike generic label smoothing, Equation \([3](https://arxiv.org/html/2609.28605#S4.E3)\) assigns mass according to the evaluation neighborhood encoded bySS\. For an interior label, the gold interval remains the unique mode \(approximately 0\.34\), while each immediate neighbor receives approximately 0\.27\. The combined neighboring mass represents the metric’s credited band, and the hard\-label term encourages exact predictions\.

#### Balancing identity and graded error\.

Hard cross entropy encourages prediction of the correct interval\. Three score\-aware terms then organize probability mass by neighborhood utility and ordinal position:

ℒhard\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{hard\}\}=−log⁡py,\\displaystyle=\-\\log p\_\{y\},\(4\)ℒsoft\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{soft\}\}=−∑kqy,klogpk,\\displaystyle=\-\\sum\_\{k\}q\_\{y,k\}\\log p\_\{k\},\(5\)ℒutil\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{util\}\}=−log⁡\(max⁡\{∑kpk​Sy,k,ϵ\}\),\\displaystyle=\-\\log\\left\(\\max\\left\\\{\\sum\_\{k\}p\_\{k\}S\_\{y,k\},\\epsilon\\right\\\}\\right\),\(6\)ℒcdf\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{cdf\}\}=19​∑kSmoothL1⁡\(∑j≤kpj,∑j≤kqy,j\)\.\\displaystyle=\\frac\{1\}\{9\}\\sum\_\{k\}\\operatorname\{SmoothL1\}\\left\(\\sum\_\{j\\leq k\}p\_\{j\},\\sum\_\{j\\leq k\}q\_\{y,j\}\\right\)\.\(7\)Hereϵ=10−8\\epsilon=10^\{\-8\}keeps the logarithm finite\. Score\-shaped cross entropy encourages the prediction to follow the neighborhoods inSS\. The utility term places mass on labels that earn task credit; the lower\-weight CDF term compares cumulative mass along the ordinal axis, so larger shifts affect more thresholds\[[1](https://arxiv.org/html/2609.28605#bib.bib7)\]\.

Letw~y\\tilde\{w\}\_\{y\}denote the normalized full class weight and letγ⁡\(e\)\\gamma\(e\)be the class scale in Table[4](https://arxiv.org/html/2609.28605#A5.T4)\. The scheduled sample weight iswy​\(e\)=1\+γ⁡\(e\)​\(w~y−1\)w\_\{y\}\(e\)=1\+\\gamma\(e\)\(\\tilde\{w\}\_\{y\}\-1\), and the weighted objective is

ℒ=wy​\(e\)​\(λh​ℒhard\+λs​ℒsoft\+λu​ℒutil\+λc​ℒcdf\),\\mathcal\{L\}=w\_\{y\}\(e\)\\left\(\\lambda\_\{h\}\\mathcal\{L\}\_\{\\mathrm\{hard\}\}\+\\lambda\_\{s\}\\mathcal\{L\}\_\{\\mathrm\{soft\}\}\+\\lambda\_\{u\}\\mathcal\{L\}\_\{\\mathrm\{util\}\}\+\\lambda\_\{c\}\\mathcal\{L\}\_\{\\mathrm\{cdf\}\}\\right\),\(8\)where everyλ\\lambdafollows the schedule in Appendix Table[4](https://arxiv.org/html/2609.28605#A5.T4)\. Across training,λu\\lambda\_\{u\}rises from 0\.42 to 0\.55 whileλh\\lambda\_\{h\}falls from 0\.15 to 0\.04, shifting the objective toward utility alignment\. Exact\-label information remains in both hard supervision and the score\-shaped target, whose unique maximum is atyy\.

#### Gradually increasing rare\-class weights\.

Scheduled weighting handles label frequency under the class\-balanced ranking objective\. For class countnkn\_\{k\}, letuk=clip⁡\(N/\(9​nk\),0\.5,3\.0\)u\_\{k\}=\\operatorname\{clip\}\(\\sqrt\{N/\(9n\_\{k\}\)\},0\.5,3\.0\)\. We apply factorsf4=1\.40f\_\{4\}=1\.40,f3=f5=1\.25f\_\{3\}=f\_\{5\}=1\.25,f2=f6=1\.10f\_\{2\}=f\_\{6\}=1\.10, andfk=1f\_\{k\}=1otherwise, then definew~k=9​uk​fk/∑juj​fj\\tilde\{w\}\_\{k\}=9u\_\{k\}f\_\{k\}/\\sum\_\{j\}u\_\{j\}f\_\{j\}\. These modest factors emphasize the central boundary and taper across adjacent intermediate labels\. Clipping precedes boundary emphasis, and the final weights are renormalized to unit mean before their gradual introduction across training\.

### 4\.2Inference and Submission Calibration

Givenpθp\_\{\\theta\}, the analytical expected\-utility rule predicts

y^=arg⁡max⁡∑ij⁡pθ,i​Si,j\.\\hat\{y\}=\\arg\\max\_\{j\}\\sum\_\{i\}p\_\{\\theta,i\}S\_\{i,j\}\.\(9\)Whenpθp\_\{\\theta\}is a calibrated posterior estimate, Equation \([9](https://arxiv.org/html/2609.28605#S4.E9)\) is the Bayes action under task utilitySS\. It differs from argmax when probability mass spans neighboring intervals\. With only 566 highly skewed labels, however, the learned posterior can inherit class\-prior bias\. The official decision layer therefore applies a compact three\-stage calibration: it adjusts the class prior and posterior sharpness, projects the corrected distribution to an expected ordinal position, and maps that position through monotone thresholds\. Its low\-dimensional parameters are selected by macro utility on stratified OOF predictions and then held fixed\. Appendix[D](https://arxiv.org/html/2609.28605#A4)gives the decision rule, while Figure[3](https://arxiv.org/html/2609.28605#S5.F3)isolates the expected\-utility decoder before calibration\.

## 5Experiments

### 5\.1Data and Protocol

FIE2026 provides 566 labeled training instances\. Both tracks use the same 2,958 context–hypothesis pairs and official annotations for evaluation, but the organizer maintains separate leaderboards for the two resource settings\.

The fine\-tuned model uses Qwen3\.5\-9B with LoRA rank 32, learning rate10−410^\{\-4\}, batch size 4 with two gradient\-accumulation steps, three epochs, classifier dropout 0\.05, and LoRA dropout 0\. Training uses 8\-bit paged AdamW and 58\.2M trainable parameters \(0\.689% of 8\.45B\), taking 6\.6 minutes on a single NVIDIA H800 GPU\.

For development and calibration, we average per\-instance logits from five stratified folds under three split seeds\. These OOF predictions support model selection, calibration, and error analysis\. The cross\-validation models and the final model share the same architecture, schedule, and nine\-logit output\. A controlled single\-seed sweep selectsTs=0\.18T\_\{s\}=0\.18by relative OOF macro utility and far\-error rate \(Appendix Figure[4](https://arxiv.org/html/2609.28605#A1.F4)\)\. After selecting the configuration, the final LoRA model is retrained on all 566 examples and uses the fixed decoder\. We use OOF results for model selection and report final performance on the official evaluation set\.

### 5\.2Baseline Comparison: Exact Matches and Nearby Predictions

Table[1](https://arxiv.org/html/2609.28605#S5.T1)compares baseline accuracy for exact matches and nearby predictions\. The majority baseline always predicts label 8\. The two LinearSVM ensembles combine character TF–IDF, structural, and cue templates; the full\-cue variant expands uncertainty, negation, evidence, modality, and factuality cues\. The regression baseline fine\-tunes mDeBERTa\-v3\[[9](https://arxiv.org/html/2609.28605#bib.bib11)\]with Huber loss\. These early baselines use three\-fold OOF predictions; the final system follows the repeated five\-fold protocol in Section[5\.1](https://arxiv.org/html/2609.28605#S5.SS1)\.

Table 1:Development results on the 566 labeled instances\. Learned systems use three\-fold OOF evaluation\. Micro utility, Exact, and Within\-2 are instance\-level averages\. Majority prediction preserves apparent utility through class frequency; scalar regression reduces distance but lowers exact\-label accuracy\.The majority predictor reaches 0\.7652 micro utility despite using only one label\. In preliminary experiments, several categorical mDeBERTa objectives reproduce this shortcut\. Huber regression reverses the error profile: within\-two accuracy rises to 0\.8145, but exact accuracy falls to 0\.1325\. The linear systems show that lexical and structural evidence can improve utility without reducing the task to a scalar\. These results therefore favor a categorical model that also encodes distance\.

### 5\.3Controlled Configuration and Decoder Comparisons

We compare fine\-tuning configurations on repeated five\-fold OOF predictions using macro utility, matching the official class\-balanced aggregation\. Figure[3](https://arxiv.org/html/2609.28605#S5.F3)\(a\) applies the same raw expected\-utility decoder to every configuration\. Figure[3](https://arxiv.org/html/2609.28605#S5.F3)\(b\) then holds logits fixed and changes only the uncalibrated decoder, separating decision effects from model changes\.

Figure 3:Relative macro utility on repeated five\-fold OOF predictions; higher is better\. \(a\) Configuration variants use the same expected\-utility decoder and are normalized by the full configuration\. \(b\) Uncalibrated decoders use identical logits and are normalized by expected utility\. Appendix[D](https://arxiv.org/html/2609.28605#A4)describes the submitted calibrated decoder\.Under the common decoder, warm\-up, class\-weight, boundary\-factor, and lower\-rank variants retain 96\.54–98\.04% of the reference macro utility\. Last\-mean pooling, a sharper score target, and a larger hard\-label loss weight retain 83\.28–91\.99%\. With unadjusted logits fixed, expected\-utility decoding reaches 100\.00%, compared with 97\.39% for argmax and 99\.24% for ordinal\-mean rounding\. Appendix[D](https://arxiv.org/html/2609.28605#A4)defines the calibrated map used for the official submission\.

### 5\.4Official Results

As shown in Table[2](https://arxiv.org/html/2609.28605#S5.T2), UO\-FIE ranks first in the fine\-tuning track, and five\-prompt ordinal fusion ranks third in the prompt track\.

Table 2:Official results on the same 2,958 pairs under separate resource settings\. Rankings use macro utility; micro utility is shown for reference\.
### 5\.5Secondary Prompt\-Track System

The prompt system averages five 0–8 predictions from four semantic templates and two Gemini variants\. All use the same rubric and three demonstrations, with different emphases on source attribution, factual commitment, scope, and boundaries between adjacent labels\. The decoder assigns extreme labels only under unanimous agreement and preserves whether the mean prediction lies above or below the neutral label\. Appendix[C](https://arxiv.org/html/2609.28605#A3)gives the model configurations and decoding rule\.

## 6Analysis and Discussion

The majority baseline’s 0\.7652 micro utility shows how an instance average can obscure sparse labels; scheduled weights and macro\-utility analysis keep them visible\. Post\-hoc error analysis suggests that labels 2–6 and reported speech remain difficult \(Appendix[B](https://arxiv.org/html/2609.28605#A2)\)\. Reported speech requires separating the writer’s commitment from an attributed source, making intermediate boundaries and attribution clear targets for improvement\.

Figure[3](https://arxiv.org/html/2609.28605#S5.F3)\(b\) compares analytical rules on fixed logits; the official system separately calibrates prior skew, sharpness, and class boundaries\.

### 6\.1Limitations

This study uses 566 labeled examples and a single backbone, and the temperature sweep uses one random seed\. The calibration is specific to FIE2026; its applicability to other ordinal tasks remains to be evaluated\.

## 7Conclusion

UO\-FIE ranks first in the fine\-tuning track, while the independent prompt system ranks third in its track\. The results support combining exact\-label supervision with graded utility in both training and decoding\.

## Acknowledgements

The author is grateful to two friends who wish to remain anonymous: one for generously providing the computing resources and infrastructure needed for the experiments, and the other for their continued care, encouragement, and emotional support throughout this work\. Without their help and encouragement, this work would have been much more difficult to complete\.

## Appendix

The appendix provides supplementary analyses and implementation details: Section A examines score\-target temperature, Section B presents post\-hoc error analyses, Section C describes the prompt configurations and aggregation rule, Section D specifies the submitted decoder, and Section E summarizes the training settings\.

## Appendix AScore\-Target Temperature

Figure[4](https://arxiv.org/html/2609.28605#A1.F4)reports the controlled sweep used to setTsT\_\{s\}\. All six runs share the same base configuration, fixed seed, and raw expected\-utility decoder; only the score\-target temperature changes\. Relative macro utility peaks at 0\.18, which we select for the final configuration\.

Figure 4:Score\-target\-temperature sensitivity on stratified OOF development predictions\. Relative macro utility is normalized to the selectedTs=0\.18T\_\{s\}=0\.18run\. The far\-error rate counts predictions at least three intervals from the gold label\. All points use one seed and the same expected\-utility decoder\.
## Appendix BPost\-hoc Error Analysis

We analyze predictions from the submitted system using reference labels constructed separately for the shared evaluation inputs\. These labels are used only for post\-hoc analysis, not for system selection or official scoring\. Figure[5](https://arxiv.org/html/2609.28605#A2.F5)summarizes class behavior, with exact values in Table[3](https://arxiv.org/html/2609.28605#A2.T3)\.

Figure 5:Post\-hoc error analysis using the separately constructed reference labels\. \(a\) Row\-normalized confusion percentages\. \(b\) Exact recall and mean utility by class\. \(c\) Far\-error rate by class\.Table 3:Per\-class post\-hoc performance\. Recall is exact\-label recall; utility uses Equation \([1](https://arxiv.org/html/2609.28605#S3.E1)\)\.Labels 2–6 have the lowest exact recall, yet their mean utility remains moderate because nearby errors retain partial credit\. This pattern suggests that intermediate boundaries remain harder than the outer intervals\.

Figure 6:Differences from overall post\-hoc performance for overlapping subsets defined by lexical keywords\. Positive values indicate improvement in all three panels\. Reported speech is the weakest of the analyzed subsets\.
## Appendix CPrompt Configurations

We denote the five prompt configurations as R1–R5 according to their semantic roles\.

Figure[7](https://arxiv.org/html/2609.28605#A3.F7)summarizes organizer feedback for R1 and the successive fusion stages, showing the progression from a single configuration to the final ensemble\.

Figure 7:Prompt\-system evolution on the 2,958 shared inputs\. The online submission interface returned an aggregate score for each stage; Table[2](https://arxiv.org/html/2609.28605#S5.T2)separately reports the final leaderboard’s macro\- and micro\-average fields\.All configurations use temperature 0 and structured JSON output\. R5 returned 2,955 of 2,958 responses; the three missing outputs were filled by R1, which uses the same prompt with Gemini\-3\.5\-Flash\. Letar∈\{0,…,8\}a\_\{r\}\\in\\\{0,\\ldots,8\\\}be the interval label from configurationrrandm=15​∑r=15arm=\\frac\{1\}\{5\}\\sum\_\{r=1\}^\{5\}a\_\{r\}\. The fixed prompt decoding rule is

y^=\{0,m=0,1,0<m<1,⌊m\+0\.5⌋,1≤m<3\.5,3,3\.5≤m<4,4,m=4,5,4<m<4\.5,⌊m\+0\.5⌋,4\.5≤m<7,7,7≤m<8,8,m=8\.\\hat\{y\}=\\begin\{cases\}0,&m=0,\\\\ 1,&0<m<1,\\\\ \\lfloor m\+0\.5\\rfloor,&1\\leq m<3\.5,\\\\ 3,&3\.5\\leq m<4,\\\\ 4,&m=4,\\\\ 5,&4<m<4\.5,\\\\ \\lfloor m\+0\.5\\rfloor,&4\.5\\leq m<7,\\\\ 7,&7\\leq m<8,\\\\ 8,&m=8\.\\end\{cases\}\(10\)Round\-half\-up is used on1≤m<3\.51\\leq m<3\.5and4\.5≤m<74\.5\\leq m<7\. The thresholds around label 4 preserve whether the mean prediction is above or below the neutral label, while labels 0 and 8 require unanimity\.

## Appendix DOrdinal Calibration

Equation \([9](https://arxiv.org/html/2609.28605#S4.E9)\) assumes that the predictive distribution is calibrated\. The submitted decoder follows the decision\-theoretic distinction between probability estimation and action selection\[[7](https://arxiv.org/html/2609.28605#bib.bib14)\]\. To compensate for prior bias under sparse, imbalanced supervision, it applies a deterministic class\-prior adjustment followed by temperature scaling and monotone ordinal calibration\. Letzkz\_\{k\}be the UO\-FIE logit for interval labelkkandπk\\pi\_\{k\}its frequency in the 566\-instance training set\. It computes

p~k\\displaystyle\\tilde\{p\}\_\{k\}=exp⁡\(\(zk−τ​log⁡πk\)/T\)∑j=08exp⁡\(\(zj−τ​log⁡πj\)/T\),\\displaystyle=\{\}\\frac\{\\exp\\left\(\(z\_\{k\}\-\\tau\\log\\pi\_\{k\}\)/T\\right\)\}\{\\sum\_\{j=0\}^\{8\}\\exp\\left\(\(z\_\{j\}\-\\tau\\log\\pi\_\{j\}\)/T\\right\)\},\(11\)r\\displaystyle r=∑k=08k​p~k,\\displaystyle=\{\}\\sum\_\{k=0\}^\{8\}k\\tilde\{p\}\_\{k\},\(12\)y^\\displaystyle\\hat\{y\}=∑j=18𝟏\[r≥bj\],\\displaystyle=\{\}\\sum\_\{j=1\}^\{8\}\\mathbf\{1\}\[r\\geq b\_\{j\}\],\(13\)whereτ\\tau,TT, and𝐛\\mathbf\{b\}are the prior\-correction strength, calibration temperature, and ordered threshold vector\. The intermediaterris the expected ordinal position, and the ordered thresholds keep the final map monotone\. Calibration proceeds in two stages on repeated stratified OOF logits: a coarse search selectsτ\\tauandTT, followed by coordinate updates of𝐛\\mathbf\{b\}under monotonicity, minimum\-spacing, and class\-coverage constraints\. The final settings, rounded to two decimal places, areτ=0\.75\\tau=0\.75,T=0\.50T=0\.50, and

𝐛=\(1\.12,1\.68,2\.95,3\.29,4\.17,4\.95,5\.93,6\.72\)\.\\mathbf\{b\}=\(1\.12,1\.68,2\.95,3\.29,4\.17,4\.95,5\.93,6\.72\)\.The selected decoder is held fixed when applied to the model retrained on all labeled examples\. Figure[3](https://arxiv.org/html/2609.28605#S5.F3)\(b\) reports the corresponding decoder comparison with fixed logits\.

## Appendix ETraining Details

Table 4:Loss and class\-weight schedule for the final fine\-tuned model\.- •Data: five\-fold OOF under three split seeds; final training on all 566 released labeled instances\.
- •Backbone: Qwen3\.5\-9B\.
- •Training: LoRAr=32r=32,α=64\\alpha=64, adapter dropout 0, classifier dropout 0\.05, three epochs, learning rate10−410^\{\-4\}, effective batch size 8; a single NVIDIA H800 GPU, 6\.6 minutes\.

## References

- \[1\]W\. Cao, V\. Mirjalili, and S\. Raschka\(2020\)Rank consistent ordinal regression for neural networks with application to age estimation\.Pattern Recognition Letters140,pp\. 325–331\.External Links:[Document](https://dx.doi.org/10.1016/j.patrec.2020.11.008),[Link](https://doi.org/10.1016/j.patrec.2020.11.008)Cited by:[§2](https://arxiv.org/html/2609.28605#S2.p2.1),[§4\.1](https://arxiv.org/html/2609.28605#S4.SS1.SSS0.Px3.p1.2)\.
- \[2\]G\. Cong, J\. Wu, C\. Yang, T\. Xun, D\. F\. Wong, B\. Li, and Y\. Yuan\(2025\)Overview of CCL25\-eval task 4: factivity inference evaluation 2025\.InProceedings of the 24th China National Conference on Computational Linguistics \(CCL 2025\),H\. Lin, B\. Li, and H\. Tan \(Eds\.\),Jinan, China,pp\. 166–180\.External Links:[Link](https://aclanthology.org/2025.ccl-2.20/)Cited by:[§1](https://arxiv.org/html/2609.28605#S1.p1.1),[§2](https://arxiv.org/html/2609.28605#S2.p1.1)\.
- \[3\]R\. Díaz and A\. Marathe\(2019\)Soft labels for ordinal regression\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 4733–4742\.External Links:[Document](https://dx.doi.org/10.1109/CVPR.2019.00487),[Link](https://doi.org/10.1109/CVPR.2019.00487)Cited by:[§2](https://arxiv.org/html/2609.28605#S2.p2.1)\.
- \[4\]C\. Elkan\(2001\)The foundations of cost\-sensitive learning\.InProceedings of the 17th International Joint Conference on Artificial Intelligence,pp\. 973–978\.External Links:[Link](https://dblp.org/rec/conf/ijcai/Elkan01)Cited by:[§2](https://arxiv.org/html/2609.28605#S2.p2.1)\.
- \[5\]FIE2026 Organizing Committee\(2026\)Factivity inference evaluation 2026: task description and evaluation rules\.Note:Official task websiteAccessed 2026\-07\-17External Links:[Link](https://github.com/UM-FAH-Yuan/FIE2026)Cited by:[§1](https://arxiv.org/html/2609.28605#S1.p2.1),[§2](https://arxiv.org/html/2609.28605#S2.p1.1)\.
- \[6\]X\. Geng\(2016\)Label distribution learning\.IEEE Transactions on Knowledge and Data Engineering28\(7\),pp\. 1734–1748\.External Links:[Document](https://dx.doi.org/10.1109/TKDE.2016.2545658),[Link](https://doi.org/10.1109/TKDE.2016.2545658)Cited by:[§2](https://arxiv.org/html/2609.28605#S2.p2.1)\.
- \[7\]T\. Gneiting\(2011\)Making and evaluating point forecasts\.Journal of the American Statistical Association106\(494\),pp\. 746–762\.External Links:[Document](https://dx.doi.org/10.1198/jasa.2011.r10138),[Link](https://doi.org/10.1198/jasa.2011.r10138)Cited by:[Appendix D](https://arxiv.org/html/2609.28605#A4.p1.1),[§2](https://arxiv.org/html/2609.28605#S2.p2.1)\.
- \[8\]S\. Gu, T\. Lu, S\. Liu, K\. Guo, and Y\. Shao\(2025\)System report for CCL25\-eval task 4: factivity inference based on dynamic few\-shot learning\.InProceedings of the 24th China National Conference on Computational Linguistics \(CCL 2025\),H\. Lin, B\. Li, and H\. Tan \(Eds\.\),Jinan, China,pp\. 128–133\.External Links:[Link](https://aclanthology.org/2025.ccl-2.15/)Cited by:[§2](https://arxiv.org/html/2609.28605#S2.p1.1)\.
- \[9\]P\. He, J\. Gao, and W\. Chen\(2021\)DeBERTaV3: improving DeBERTa using ELECTRA\-style pre\-training with gradient\-disentangled embedding sharing\.arXiv preprint arXiv:2111\.09543\.External Links:[Link](https://arxiv.org/abs/2111.09543)Cited by:[§5\.2](https://arxiv.org/html/2609.28605#S5.SS2.p1.1)\.
- \[10\]E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen\(2021\)LoRA: low\-rank adaptation of large language models\.arXiv preprint arXiv:2106\.09685\.External Links:[Link](https://arxiv.org/abs/2106.09685)Cited by:[§1](https://arxiv.org/html/2609.28605#S1.p4.1)\.
- \[11\]D\. Liu, L\. Xia, Y\. Zhang, X\. Yang, and F\. Kong\(2025\)System report for CCL25\-eval task 4: prompting, scheduling, and arbitration strategies for Chinese factivity inference\.InProceedings of the 24th China National Conference on Computational Linguistics \(CCL 2025\),H\. Lin, B\. Li, and H\. Tan \(Eds\.\),Jinan, China,pp\. 146–151\.External Links:[Link](https://aclanthology.org/2025.ccl-2.17/)Cited by:[§2](https://arxiv.org/html/2609.28605#S2.p1.1)\.
- \[12\]Qwen Team\(2026\)Qwen3\.5\-9B model card\.Note:Hugging Face model repositoryRevision c2022362; accessed 2026\-07\-15External Links:[Link](https://huggingface.co/Qwen/Qwen3.5-9B)Cited by:[§4\.1](https://arxiv.org/html/2609.28605#S4.SS1.SSS0.Px1.p1.1)\.
- \[13\]R\. Saurí and J\. Pustejovsky\(2009\)FactBank: a corpus annotated with event factuality\.Language Resources and Evaluation43\(3\),pp\. 227–268\.External Links:[Document](https://dx.doi.org/10.1007/s10579-009-9089-9),[Link](https://doi.org/10.1007/s10579-009-9089-9)Cited by:[§1](https://arxiv.org/html/2609.28605#S1.p1.1),[§2](https://arxiv.org/html/2609.28605#S2.p1.1)\.
- \[14\]J\. Wang, R\. Liu, L\. Zhang, and J\. Li\(2025\)System report for CCL25\-Eval task 10: SRAG\-MAV for fine\-grained chinese hate speech recognition\.External Links:2507\.18580,[Link](https://arxiv.org/abs/2507.18580)Cited by:[§2](https://arxiv.org/html/2609.28605#S2.p1.1)\.
- \[15\]A\. S\. White, R\. Rudinger, K\. Rawlins, and B\. Van Durme\(2018\)Lexicosyntactic inference in neural models\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,Brussels, Belgium,pp\. 4717–4724\.External Links:[Link](https://aclanthology.org/D18-1501/),[Document](https://dx.doi.org/10.18653/v1/D18-1501)Cited by:[§1](https://arxiv.org/html/2609.28605#S1.p1.1),[§2](https://arxiv.org/html/2609.28605#S2.p1.1)\.
- \[16\]H\. Zhang, P\. Li, Z\. Qian, and X\. Zhu\(2023\)Incorporating factuality inference to identify document\-level event factuality\.InFindings of the Association for Computational Linguistics: ACL 2023,External Links:[Document](https://dx.doi.org/10.18653/v1/2023.findings-acl.879),[Link](https://aclanthology.org/2023.findings-acl.879/)Cited by:[§2](https://arxiv.org/html/2609.28605#S2.p1.1)\.

相似文章

FocuSFT:面向稀释感知长上下文微调的双层优化

Hugging Face Daily Papers

本文介绍了 FocuSFT,这是一种双层优化框架,它通过参数化记忆机制解决注意力稀释问题,从而提升长上下文语言模型的性能。在 BABILong 和 RULER 等基准测试中,该框架在准确性和上下文参与度方面均展现出显著提升。

FineSteer: 大规模语言模型推理时细粒度控制的统一框架

arXiv cs.CL

FineSteer 是一个新颖的推理时控制框架,将控制分解为条件控制和细粒度向量合成两个阶段,采用子空间引导条件控制(SCS)和混合控制专家(MoSE)机制来提高安全性和真实性,同时保持模型效用。实验表明在 TruthfulQA 上相比最新方法有 7.6% 的性能提升,且效用损失最小。