经典语义抽取式摘要能否在印地语中进行评估?一项复现研究
摘要
本文复现了一种针对印地语的分布式语义抽取式摘要方法,并在标准语料库上进行评估,发现句子位置是唯一有效的特征,且当前的印地语基准测试未能激励高级内容选择。
arXiv:2609.29090v1 Announce Type: new
Abstract: We replicate the distributional-semantics extractive summarisation method of Mohd, Jan and Shah (2020) and adapt it to Hindi, substituting a Devanagari-appropriate component at every language-specific step. The system is evaluated on two independent corpora --- the Hindi portion of XL-Sum and FIRE ILSUM 2.0 Hindi --- under a Devanagari-aware ROUGE implementation validated against the XL-Sum authors' own multilingual scorer, with all comparisons drawn as 1000-resample paired bootstraps. In its published equal-weight configuration the replicated system is significantly worse than a three-sentence lead baseline on both corpora, trailing Lead-3 by 0.042 ROUGE-1 Fon XL-Sum and by 0.265 on ILSUM. A feature ablation shows that sentenceposition is the only feature that contributes: position alone reproduces the lead baseline exactly, removing position gives the weakest configuration,and a validation-tuned weighting can at best equal Lead-3 and never exceed it. TextRank fails identically, making this a class-level rather than an implementation-level result. A selection analysis shows the remaining features steer extraction towards long, entity-dense body sentences while the references reuse the article lead.Current Hindi benchmarks therefore cannot reward non-lead content selection, motivating purpose-built evaluation resources.
查看缓存全文
缓存时间: 2026/09/25 09:16
# Can Classical Semantic-Extractive Summarization Be Evaluated in Hindi?A Replication Study Source: [https://arxiv.org/html/2609.29090](https://arxiv.org/html/2609.29090) Showket Ahmad Khan Mudasir Mohd Nasrullah Sheikh††thanks:ORCID: 0009\-0001\-8625\-8496††thanks:Corresponding author:[mudasir\.mohammad@kashmiruniversity\.ac\.in](mailto:[email protected])Affiliation:Department of Computer Science, South Campus, University of Kashmir, Anantnag, IndiaAffiliation:IBM Research, San Jose, CA, USAMohsin Altaf Wani Abid Hussain Wani Hilal Ahmad Khanday Niyaz Ahmad Wani††thanks:ORCID: 0000\-0003\-0094\-9921Affiliation:Department of Computer Science, South Campus, University of Kashmir, Anantnag, IndiaAffiliation:Manipal University Jaipur, Dehmi Kalan, Jaipur 303007, Rajasthan, India September 24, 2026 ###### Abstract We replicate the distributional\-semantics extractive summarisation method of Mohd, Jan and Shah \(2020\) and adapt it to Hindi, substituting a Devanagari\-appropriate component at every language\-specific step\. The system is evaluated on two independent corpora — the Hindi portion of XL\-Sum and FIRE ILSUM 2\.0 Hindi — under a Devanagari\-aware ROUGE implementation validated against the XL\-Sum authors’ own multilingual scorer, with all comparisons drawn as 1000\-resample paired bootstraps\. In its published equal\-weight configuration the replicated system is significantly worse than a three\-sentence lead baseline on both corpora, trailing Lead\-3 by 0\.042 ROUGE\-1 F on XL\-Sum and by 0\.265 on ILSUM\. A feature ablation shows that sentence position is the only feature that contributes: position alone reproduces the lead baseline exactly, removing position gives the weakest configuration, and a validation\-tuned weighting can at best equal Lead\-3 and never exceed it\. TextRank fails identically, making this a class\-level rather than an implementation\-level result\. A selection analysis shows the remaining features steer extraction towards long, entity\-dense body sentences while the references reuse the article lead\. Current Hindi benchmarks therefore cannot reward non\-lead content selection, motivating purpose\-built evaluation resources\. ## 1Setup We replicate the distributional\-semantics extractive summariser of Mohd, Jan and Shah\[[1](https://arxiv.org/html/2609.29090#bib.bib1)\]and adapt it to Hindi\. The system implements the paper’s pipeline unchanged in structure: each sentence is embedded through a word\-embedding “big vector”, the resulting sentence representations are clustered withkk\-means, and sentences are scored by an equally weighted sum of seven normalised features — sentence length, sentence position, TF–IDF mass, noun/verb count, proper\-noun count, aggregate embedding cosine to the rest of the document, and a cue\-phrase indicator — with a final extract assembled by a per\-cluster round\-robin over the ranked sentences\. Our Hindi instantiation \(available in the accompanying repository,hinexsum/\[[8](https://arxiv.org/html/2609.29090#bib.bib8)\]\) substitutes Devanagari\-appropriate components at each language\-specific step: NFC normalisation and danda\-aware sentence splitting for preprocessing, a Hindi stop\-word list and light suffix stemmer for lexical cleaning, Stanza’s Hindi model\[[5](https://arxiv.org/html/2609.29090#bib.bib5)\]for part\-of\-speech features, and 300\-dimensional fastText Common Crawl vectors \(cc\.hi\.300\.bin\)\[[4](https://arxiv.org/html/2609.29090#bib.bib4)\]for the embedding space\. Part\-of\-speech features are enabled throughout the results reported here, so every “our system” figure is the full seven\-feature configuration rather than an ablation of it\. We evaluate on two independent Hindi news\-summarisation corpora: the Hindi portion of XL\-Sum\[[2](https://arxiv.org/html/2609.29090#bib.bib2)\]\(single\-reference abstractive summaries of BBC articles\) and FIRE ILSUM 2\.0 Hindi\[[3](https://arxiv.org/html/2609.29090#bib.bib3)\]\(news article/summary pairs, obtained ungated from theILSUM/ILSUM\-2\.0Hugging Face distribution\)\. The evaluation protocol is held constant across both\. Every system produces a three\-sentence extract, and all systems — ours and the baselines — are scored on the identical danda\-aware sentence segmentation so that no comparison is confounded by tokenisation of the source\. Model selection and testing are separated: any weight chosen by tuning is selected on a validation slice and then evaluated once on a disjoint test slice \(for XL\-Sum, held\-out test documents 201–400 after tuning on a 200\-document validation split; for ILSUM, 200 training\-split documents for tuning and 200 test\-split documents for evaluation\)\. All confidence intervals are 95% intervals from a 1000\-resample bootstrap over documents, and system\-versus\-baseline gaps are reported as paired bootstraps drawn from a single shared resample matrix so that the difference intervals are internally valid\. Scores are reported for ROUGE\-1, ROUGE\-2, ROUGE\-L and ROUGE\-SU4 F\-measure\[[7](https://arxiv.org/html/2609.29090#bib.bib7)\]; ROUGE\-1 F is used as the primary quantity for significance because it is the least sparse at a three\-sentence budget\. Before drawing any conclusion from these numbers we validated our Devanagari\-aware ROUGE implementation \(hinexsum/rouge\.py\) against the XL\-Sum authors’ own multilingual scorer \(the csebuetnlp fork ofrouge\_scoreinvoked withlang=’hindi’and its pyonmttok tokeniser\)\. On twenty XL\-Sum documents scored under both implementations for three systems, the mean per\-document absolute difference in ROUGE\-1, ROUGE\-2 and ROUGE\-L F was 0\.0000, well inside the 0\.02 tolerance we set in advance\. The two scorers are genuinely independent code with different tokenisers — ours keeps the danda glued to a sentence\-final word whereas the external scorer strips punctuation — and on a constructed danda\-adjacent case they diverge by 0\.167, confirming that the external path is exercised rather than aliased; the difference simply never flips a content\-word match in real extractive summaries\. Because the external scorer does not compute ROUGE\-SU4, this cross\-check covers ROUGE\-1, ROUGE\-2 and ROUGE\-L only; we therefore treat those three metrics as instrument\-independent, while the ROUGE\-SU4 figures we report should be read as internally consistent but not externally cross\-validated\. ## 2Main result The central finding is that the replicated system, in its published equal\-weight configuration, is significantly worse than a lead baseline on both corpora\. On the XL\-Sum ablation set \(the first 200 test documents\), the equal\-weight system reaches ROUGE\-1 F of 0\.186 with clustering and round\-robin selection and 0\.185 without clustering, against 0\.228 for Lead\-3, a strong three\-sentence lead baseline\. On ILSUM the same equal\-weight configuration reaches only 0\.256 against 0\.522 for Lead\-3 — less than half the lead score\. Tables[1](https://arxiv.org/html/2609.29090#S2.T1)and[3](https://arxiv.org/html/2609.29090#S2.T3)give the full four\-metric breakdown with ROUGE\-1 confidence intervals; Tables[2](https://arxiv.org/html/2609.29090#S2.T2)and[4](https://arxiv.org/html/2609.29090#S2.T4)give the paired gaps against Lead\-3 that establish significance\. Table 1:XL\-Sum Hindi, 200 test documents, three\-sentence extracts\. ROUGE F\-measure; ROUGE\-1 with 95% bootstrap CI\.Table 2:XL\-Sum Hindi, paired ROUGE\-1 F gap against Lead\-3 \(1000\-resample paired bootstrap\)\. A gap is “real” when its 95% interval excludes zero\.Table 3:ILSUM 2\.0 Hindi, 200 held\-out test documents, three\-sentence extracts\. ROUGE F\-measure; ROUGE\-1 with 95% bootstrap CI\.Table 4:ILSUM 2\.0 Hindi, paired ROUGE\-1 F gap against Lead\-3 \(1000\-resample paired bootstrap\)\.The paired intervals show that the shortfall is not sampling noise\. On XL\-Sum the equal\-weight system trails Lead\-3 by 0\.042 to 0\.043 ROUGE\-1 F with intervals bounded well away from zero, and on ILSUM it trails by 0\.265 with an interval that does not approach zero\. It is important that TextRank\[[6](https://arxiv.org/html/2609.29090#bib.bib6)\], a second and entirely separate unsupervised extractive method, fails in exactly the same way: it trails Lead\-3 by 0\.025 on XL\-Sum and by 0\.254 on ILSUM, both significant\. Because our ROUGE agrees with the authors’ scorer to four decimal places, and because a method we did not write reproduces the same defeat, the result is best read as a class\-level failure of salience\-driven unsupervised extraction against a lead baseline on these corpora, not as a defect in our particular re\-implementation\. ## 3Ablation Isolating the contribution of each feature explains the shortfall precisely\. When the ranking uses the position feature alone, the system becomes identical to the lead baseline: on XL\-Sum “Ours\-position\-only” scores 0\.228 ROUGE\-1 F, the same value as Lead\-3 to three decimals, with a paired gap of\+\+0\.000 and an interval of \[\+\+0\.000,\+\+0\.000\] \(Tables[1](https://arxiv.org/html/2609.29090#S2.T1)and[2](https://arxiv.org/html/2609.29090#S2.T2)\)\. This equivalence is expected — the position score is monotone in sentence index, so selecting the top\-scoring sentences under position alone returns the article opening — but it is informative, because position\-only is the best\-performing configuration of our system, above the full seven\-feature model\. Conversely, removing position and retaining the other six features \(“Ours\-minus\-position”\) yields 0\.182, the weakest configuration in the table and significantly below Lead\-3 by 0\.046\. The equal\-weight full system sits between these poles at 0\.185–0\.186, indicating that when position is only one summand of seven its advantage is largely averaged away\. A validation\-set weight sweep confirms the pattern is monotone rather than incidental\. Increasing the position weight while holding the remaining features at unit weight raises ROUGE\-1 F monotonically on both corpora: on XL\-Sum the seven\-feature family rises from 0\.191 at unit position weight to 0\.235 at weight sixteen, and a position\-plus\-TF–IDF\-only family rises from 0\.211 to 0\.239 over the same range; on ILSUM the position\-plus\-TF–IDF family rises from 0\.347 at unit weight to 0\.560 at weight sixteen \(Table[5](https://arxiv.org/html/2609.29090#S3.T5)\)\. In every case the published equal\-weight setting is the worst point on the curve and heavier position weighting is monotonically better, up to a ceiling\. Table 5:Validation\-set ROUGE\-1 F under a position\-weight sweep \(three\-sentence budget\)\. XL\-Sum figures are the seven\-feature and position\+TF–IDF families; ILSUM figures are the position\+TF–IDF family\.The tuned ceiling is a win on neither corpus\. The single configuration selected on validation — position\+TF–IDF at position weight sixteen — was evaluated once on each held\-out test slice, and the outcome is summarised in Table[6](https://arxiv.org/html/2609.29090#S3.T6)\. On XL\-Sum it reaches 0\.231 ROUGE\-1 F against Lead\-3’s 0\.231, a paired gap of−\-0\.000 whose interval \[−\-0\.002,\+\+0\.001\] comfortably contains zero: a statistical tie, and a genuine equivalence rather than an underpowered comparison, given how narrow the interval is\. On ILSUM it reaches 0\.516 against 0\.522, a paired gap of−\-0\.005 whose interval \[−\-0\.011,−\-0\.000\] excludes zero only at its upper bound; the deficit is therefore statistically detectable but negligible \(≤\\leq0\.011 R1\-F\), leaving the tuned system at best equal to Lead\-3 and never above it\. The cross\-corpus reading is accordingly a tie on XL\-Sum, a marginal deficit on ILSUM, and a win on neither\. Careful weighting nonetheless recovers all of the ground the equal\-weight configuration loses — a real improvement over the published setting that decisively beats the equal\-weight, TextRank and random systems — so the best attainable behaviour of this feature family is to reproduce lead selection rather than to surpass it\. Table 6:Tuned configuration \(position\+TF–IDF, position weight sixteen\) against Lead\-3 on each held\-out test slice; ROUGE\-1 F with 95% bootstrap CIs and the paired gap\. Statistically detectable but negligible \(≤\\leq0\.011 R1\-F\) on ILSUM and a tie on XL\-Sum: at best equal, never above\. ## 4Mechanism The reason is visible in where each system reads\. Table[7](https://arxiv.org/html/2609.29090#S4.T7)gives, for the XL\-Sum ablation documents, the distribution of the source\-article positions of selected sentences, normalised so that zero is the article start and one its end, binned into deciles, together with the mean normalised position\. Position\-only and Lead\-3 are indistinguishable, drawing 0\.71 of their sentences from the first decile and reaching a mean position of 0\.070\. The minus\-position system is the most back\-loaded of all, with a mean position of 0\.501 and its single largest mass in the final decile, because once position is removed the embedding\-cosine, TF–IDF, length and proper\-noun features pull selection towards lexically dense sentences that in news prose sit in the body and tail\. The equal\-weight systems fall in between at a mean of roughly 0\.39, retaining only a faint lead tilt \(0\.20 in the first decile against the lead’s 0\.71\) — the quantitative signature of position being diluted to one\-seventh of the score\. Random selection is near\-uniform at 0\.488 and TextRank only slightly lead\-tilted at 0\.441\. Table 7:Fraction of selected sentences by source\-position decile on the XL\-Sum ablation set \(D1 = first tenth of the article, D10 = last tenth\), with mean normalised position\.A single document makes the failure mode concrete\. In XL\-Sum test document 179, a seventeen\-sentence report on an India–Bangladesh cricket Test, the reference summary is*“bharat aur bangladesh ke khilaf dusra Test shukravar se Chatgaon mein shuru ho raha hai\. pahla Test jitne ke bad bharatiya team do Test maichon ki series 2\-0 se jitne ke liye maidan par utregi\.”*\(“The second Test against India and Bangladesh begins on Friday in Chittagong\. After winning the first Test, the Indian team will take the field to win the two\-match Test series 2–0\.”\) — a two\-sentence framing of the fixture drawn from the article’s opening\. Lead\-3 selects sentences 0, 1 and 2, which introduce the match and its context\. The minus\-position system instead selects sentences 13, 15 and 16 \(normalised positions 0\.81, 0\.94 and 1\.00\), of which two are the squad lists: sentence 15 is*“bharatiya team \(inmen se chuni jayegi\) Saurabh Ganguly \(kaptan\), Virender Sehwag, Gautam Gambhir, Rahul Dravid, Sachin Tendulkar …”*\(“Indian team \(to be selected from among these\): Saurav Ganguly \(captain\), Virender Sehwag, Gautam Gambhir, Rahul Dravid, Sachin Tendulkar …”\) and sentence 16 the corresponding Bangladesh roster\. These sentences are exactly what the non\-position features reward: a roster is saturated with proper nouns, is long, and — because player names recur across the article — scores highly on both TF–IDF mass and aggregate embedding cosine\. Every feature the method treats as a proxy for salience is maximised by a list of names that carries almost none of the article’s summary\-worthy content, and the reference, being a lead paraphrase, shares almost no vocabulary with it\. The proper\-noun, length and TF–IDF features are not merely uninformative here; they are actively misdirected towards the least summary\-like sentences in the document\. ## 5Diagnosis The uncomfortable conclusion is that the benchmarks, not only the method, are the confounder\. The ILSUM Lead\-3 figures are the clearest evidence: a three\-sentence lead achieves ROUGE\-2 F of 0\.451 and ROUGE\-L F of 0\.499 \(Table[3](https://arxiv.org/html/2609.29090#S2.T3)\), which is only possible if the reference reuses the article’s opening almost verbatim at the bigram level\. An ILSUM reference is, in effect, a lightly edited copy of the lead, so any system that does not begin at the top is penalised regardless of whether its selection is informative\. XL\-Sum poses the same obstacle in a different form: its references are single lead\-style sentences, so a three\-sentence extract can match at most a fraction of the reference and the maximal\-overlap strategy is again to take the opening\. Under both designs the target rewards lead reproduction and offers no credit for correctly identifying salient non\-lead content, which is precisely the capability a semantic\-extractive method is intended to provide\. This is why heavier position weighting monotonically helps and why the ceiling is a tie with the lead: on these corpora the score is very nearly a measure of how lead\-like a system is, and no feature that pulls selection away from the opening can be rewarded\. We are careful not to overclaim\. What these experiments establish is narrow and specific: on XL\-Sum Hindi and ILSUM 2\.0 Hindi, at a three\-sentence budget under instrument\-validated ROUGE, the replicated equal\-weight semantic\-extractive method and a comparable TextRank baseline are significantly worse than a lead baseline, and the best weighting of this feature family only matches the lead\. We do not claim that distributional\-semantics extraction is without value in general, nor that these features cannot help on tasks whose references reward non\-lead content; our evidence speaks only to these benchmarks, whose reference style makes the lead nearly unbeatable and therefore cannot discriminate a good content selector from a positional heuristic\. Two further caveats bound the claim: we evaluate at a single three\-sentence budget, so the absolute scores could shift under a different summary length, even though the strength of the lead advantage makes a reversal unlikely\. Moreover, because ROUGE credits lexicalnn\-gram overlap rather than conveyed meaning, its design compounds the lead\-reference bias by rewarding token reuse over paraphrase, which is why the resource we propose below should be scored with complementary measures such as ChrF and BERTScore alongside human references, and not with ROUGE alone\. The failure we document is a failure of measurability as much as of method\. This motivates a purpose\-built Hindi resource\. A benchmark able to reward semantic content selection would need human\-written summaries that are longer than a single lead sentence and deliberately draw on material from across the article rather than its opening, ideally with multiple references per document so that legitimate paraphrase variation is credited rather than penalised\. It would further benefit from sentence\-level faithfulness labels, so that a system choosing an informative but non\-lead sentence can be distinguished from one that has simply drifted into the article’s tail\. Only against such a resource can the question in our title be answered on its own terms, rather than settled in advance by the reference style of the available data\. ## Software and data availability The Hindi summariser, the Devanagari\-aware ROUGE module and the full evaluation harness used for every number in this paper are released as HinExSum\[[8](https://arxiv.org/html/2609.29090#bib.bib8)\], an open\-source Python package under the MIT licence, at[https://github\.com/mudasirmohd/HinExSum](https://github.com/mudasirmohd/HinExSum)\. A reproducible Code Ocean capsule containing the code, the environment specification and the evaluation entry points is archived at[https://doi\.org/10\.24433/CO\.6462433\.v1](https://doi.org/10.24433/CO.6462433.v1)\. The corpora are third\-party resources and are not redistributed here: the Hindi portion of XL\-Sum\[[2](https://arxiv.org/html/2609.29090#bib.bib2)\]and FIRE ILSUM 2\.0 Hindi\[[3](https://arxiv.org/html/2609.29090#bib.bib3)\]are obtained from their public Hugging Face distributions, and the fastTextcc\.hi\.300\.binvectors\[[4](https://arxiv.org/html/2609.29090#bib.bib4)\]are downloaded from the fastText site as described in the repositoryREADME\. ## References - \[1\]M\. Mohd, R\. Jan, and M\. Shah\.Text document summarization using word embedding\.*Expert Systems with Applications*, 143:112958, 2020\.doi:[10\.1016/j\.eswa\.2019\.112958](https://doi.org/10.1016/j.eswa.2019.112958)\. - \[2\]T\. Hasan, A\. Bhattacharjee, M\. S\. Islam, K\. Mubasshir, Y\.\-F\. Li, Y\.\-B\. Kang, M\. S\. Rahman, and R\. Shahriyar\.XL\-Sum: Large\-scale multilingual abstractive summarization for 44 languages\.In*Findings of the Association for Computational Linguistics: ACL\-IJCNLP 2021*, pages 4693–4703, 2021\. - \[3\]S\. Satapara, B\. Modha, S\. Modha, and P\. Mehta\.FIRE 2022 ILSUM track: Indian language summarization\.In*Proceedings of the 14th Annual Meeting of the Forum for Information Retrieval Evaluation \(FIRE\)*, 2022\. - \[4\]E\. Grave, P\. Bojanowski, P\. Gupta, A\. Joulin, and T\. Mikolov\.Learning word vectors for 157 languages\.In*Proceedings of the 11th International Conference on Language Resources and Evaluation \(LREC\)*, 2018\. - \[5\]P\. Qi, Y\. Zhang, Y\. Zhang, J\. Bolton, and C\. D\. Manning\.Stanza: A Python natural language processing toolkit for many human languages\.In*Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations*, pages 101–108, 2020\. - \[6\]R\. Mihalcea and P\. Tarau\.TextRank: Bringing order into text\.In*Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing \(EMNLP\)*, pages 404–411, 2004\. - \[7\]C\.\-Y\. Lin\.ROUGE: A package for automatic evaluation of summaries\.In*Text Summarization Branches Out: Proceedings of the ACL\-04 Workshop*, pages 74–81, 2004\. - \[8\]M\. Mohd\.HinExSum: A Hindi extractive summariser using distributional semantics, version 1\.0\.0 \[computer software\], 2026\.[https://github\.com/mudasirmohd/HinExSum](https://github.com/mudasirmohd/HinExSum)\.Code Ocean capsule: doi:[10\.24433/CO\.6462433\.v1](https://doi.org/10.24433/CO.6462433.v1)\.
相似文章
使用语法与语义上下文评估汇总(SSAS)的情感预测一致性分析
本论文提出了SSAS(语法与语义上下文评估汇总)框架,旨在通过分层分类和迭代汇总来减少噪声和方差,提高基于大语言模型的情感预测的一致性。在三个行业标准数据集上的实证评估显示,数据质量和企业决策可靠性可提升30%。
一种用于生物医学与临床文本抽取式摘要的混合分层1D-CNN-BiLSTM框架
本文介绍了一种混合分层1D-CNN-BiLSTM框架,用于生物医学与临床文本的抽取式摘要,旨在通过直接从源文档中选择句子而不是生成新文本来保持事实性。
LexLattice:基于文档层次结构的神经元胞自动机多语言抽取式摘要
LexLattice 提出了一种多语言抽取式摘要方法,该方法利用文档层次结构上的神经元胞自动机,在 EUR-Lex-Sum 数据集上实现了最先进的 ROUGE 分数,并通过紧凑模型超越了大型指令调优基线。
抽象摘要中关系级幻觉评估的基于依据的分解框架
本文提出了一个基于依据的、可分解的框架,用于评估抽象式摘要中的关系级幻觉,引入了标准化的关系幻觉指数(RHI),并带有语言感知的关系抽取增强。
评估文本摘要的抽象性度量:一种改进的公式及其经验验证
本文引入了参考抽象度(RA)、摘要抽象度(SA)和抽象比率(AR)等度量指标,通过使用文档长度的调和均值与三次非重叠因子来量化文本摘要中的抽象性。在XSUM数据集上对四种模型进行的经验评估表明,这些度量能够有效区分抽取式摘要和生成式摘要,并能识别潜在的幻觉问题。