Can Classical Semantic-Extractive Summarization Be Evaluated in Hindi? A Replication Study
Summary
This paper replicates a distributional-semantics extractive summarization method for Hindi and evaluates it on standard corpora, finding that sentence position is the only contributing feature and current Hindi benchmarks fail to incentivize advanced content selection.
View Cached Full Text
Cached at: 09/25/26, 09:16 AM
# Can Classical Semantic-Extractive Summarization Be Evaluated in Hindi?A Replication Study Source: [https://arxiv.org/html/2609.29090](https://arxiv.org/html/2609.29090) Showket Ahmad Khan Mudasir Mohd Nasrullah Sheikh††thanks:ORCID: 0009\-0001\-8625\-8496††thanks:Corresponding author:[mudasir\.mohammad@kashmiruniversity\.ac\.in](mailto:[email protected])Affiliation:Department of Computer Science, South Campus, University of Kashmir, Anantnag, IndiaAffiliation:IBM Research, San Jose, CA, USAMohsin Altaf Wani Abid Hussain Wani Hilal Ahmad Khanday Niyaz Ahmad Wani††thanks:ORCID: 0000\-0003\-0094\-9921Affiliation:Department of Computer Science, South Campus, University of Kashmir, Anantnag, IndiaAffiliation:Manipal University Jaipur, Dehmi Kalan, Jaipur 303007, Rajasthan, India September 24, 2026 ###### Abstract We replicate the distributional\-semantics extractive summarisation method of Mohd, Jan and Shah \(2020\) and adapt it to Hindi, substituting a Devanagari\-appropriate component at every language\-specific step\. The system is evaluated on two independent corpora — the Hindi portion of XL\-Sum and FIRE ILSUM 2\.0 Hindi — under a Devanagari\-aware ROUGE implementation validated against the XL\-Sum authors’ own multilingual scorer, with all comparisons drawn as 1000\-resample paired bootstraps\. In its published equal\-weight configuration the replicated system is significantly worse than a three\-sentence lead baseline on both corpora, trailing Lead\-3 by 0\.042 ROUGE\-1 F on XL\-Sum and by 0\.265 on ILSUM\. A feature ablation shows that sentence position is the only feature that contributes: position alone reproduces the lead baseline exactly, removing position gives the weakest configuration, and a validation\-tuned weighting can at best equal Lead\-3 and never exceed it\. TextRank fails identically, making this a class\-level rather than an implementation\-level result\. A selection analysis shows the remaining features steer extraction towards long, entity\-dense body sentences while the references reuse the article lead\. Current Hindi benchmarks therefore cannot reward non\-lead content selection, motivating purpose\-built evaluation resources\. ## 1Setup We replicate the distributional\-semantics extractive summariser of Mohd, Jan and Shah\[[1](https://arxiv.org/html/2609.29090#bib.bib1)\]and adapt it to Hindi\. The system implements the paper’s pipeline unchanged in structure: each sentence is embedded through a word\-embedding “big vector”, the resulting sentence representations are clustered withkk\-means, and sentences are scored by an equally weighted sum of seven normalised features — sentence length, sentence position, TF–IDF mass, noun/verb count, proper\-noun count, aggregate embedding cosine to the rest of the document, and a cue\-phrase indicator — with a final extract assembled by a per\-cluster round\-robin over the ranked sentences\. Our Hindi instantiation \(available in the accompanying repository,hinexsum/\[[8](https://arxiv.org/html/2609.29090#bib.bib8)\]\) substitutes Devanagari\-appropriate components at each language\-specific step: NFC normalisation and danda\-aware sentence splitting for preprocessing, a Hindi stop\-word list and light suffix stemmer for lexical cleaning, Stanza’s Hindi model\[[5](https://arxiv.org/html/2609.29090#bib.bib5)\]for part\-of\-speech features, and 300\-dimensional fastText Common Crawl vectors \(cc\.hi\.300\.bin\)\[[4](https://arxiv.org/html/2609.29090#bib.bib4)\]for the embedding space\. Part\-of\-speech features are enabled throughout the results reported here, so every “our system” figure is the full seven\-feature configuration rather than an ablation of it\. We evaluate on two independent Hindi news\-summarisation corpora: the Hindi portion of XL\-Sum\[[2](https://arxiv.org/html/2609.29090#bib.bib2)\]\(single\-reference abstractive summaries of BBC articles\) and FIRE ILSUM 2\.0 Hindi\[[3](https://arxiv.org/html/2609.29090#bib.bib3)\]\(news article/summary pairs, obtained ungated from theILSUM/ILSUM\-2\.0Hugging Face distribution\)\. The evaluation protocol is held constant across both\. Every system produces a three\-sentence extract, and all systems — ours and the baselines — are scored on the identical danda\-aware sentence segmentation so that no comparison is confounded by tokenisation of the source\. Model selection and testing are separated: any weight chosen by tuning is selected on a validation slice and then evaluated once on a disjoint test slice \(for XL\-Sum, held\-out test documents 201–400 after tuning on a 200\-document validation split; for ILSUM, 200 training\-split documents for tuning and 200 test\-split documents for evaluation\)\. All confidence intervals are 95% intervals from a 1000\-resample bootstrap over documents, and system\-versus\-baseline gaps are reported as paired bootstraps drawn from a single shared resample matrix so that the difference intervals are internally valid\. Scores are reported for ROUGE\-1, ROUGE\-2, ROUGE\-L and ROUGE\-SU4 F\-measure\[[7](https://arxiv.org/html/2609.29090#bib.bib7)\]; ROUGE\-1 F is used as the primary quantity for significance because it is the least sparse at a three\-sentence budget\. Before drawing any conclusion from these numbers we validated our Devanagari\-aware ROUGE implementation \(hinexsum/rouge\.py\) against the XL\-Sum authors’ own multilingual scorer \(the csebuetnlp fork ofrouge\_scoreinvoked withlang=’hindi’and its pyonmttok tokeniser\)\. On twenty XL\-Sum documents scored under both implementations for three systems, the mean per\-document absolute difference in ROUGE\-1, ROUGE\-2 and ROUGE\-L F was 0\.0000, well inside the 0\.02 tolerance we set in advance\. The two scorers are genuinely independent code with different tokenisers — ours keeps the danda glued to a sentence\-final word whereas the external scorer strips punctuation — and on a constructed danda\-adjacent case they diverge by 0\.167, confirming that the external path is exercised rather than aliased; the difference simply never flips a content\-word match in real extractive summaries\. Because the external scorer does not compute ROUGE\-SU4, this cross\-check covers ROUGE\-1, ROUGE\-2 and ROUGE\-L only; we therefore treat those three metrics as instrument\-independent, while the ROUGE\-SU4 figures we report should be read as internally consistent but not externally cross\-validated\. ## 2Main result The central finding is that the replicated system, in its published equal\-weight configuration, is significantly worse than a lead baseline on both corpora\. On the XL\-Sum ablation set \(the first 200 test documents\), the equal\-weight system reaches ROUGE\-1 F of 0\.186 with clustering and round\-robin selection and 0\.185 without clustering, against 0\.228 for Lead\-3, a strong three\-sentence lead baseline\. On ILSUM the same equal\-weight configuration reaches only 0\.256 against 0\.522 for Lead\-3 — less than half the lead score\. Tables[1](https://arxiv.org/html/2609.29090#S2.T1)and[3](https://arxiv.org/html/2609.29090#S2.T3)give the full four\-metric breakdown with ROUGE\-1 confidence intervals; Tables[2](https://arxiv.org/html/2609.29090#S2.T2)and[4](https://arxiv.org/html/2609.29090#S2.T4)give the paired gaps against Lead\-3 that establish significance\. Table 1:XL\-Sum Hindi, 200 test documents, three\-sentence extracts\. ROUGE F\-measure; ROUGE\-1 with 95% bootstrap CI\.Table 2:XL\-Sum Hindi, paired ROUGE\-1 F gap against Lead\-3 \(1000\-resample paired bootstrap\)\. A gap is “real” when its 95% interval excludes zero\.Table 3:ILSUM 2\.0 Hindi, 200 held\-out test documents, three\-sentence extracts\. ROUGE F\-measure; ROUGE\-1 with 95% bootstrap CI\.Table 4:ILSUM 2\.0 Hindi, paired ROUGE\-1 F gap against Lead\-3 \(1000\-resample paired bootstrap\)\.The paired intervals show that the shortfall is not sampling noise\. On XL\-Sum the equal\-weight system trails Lead\-3 by 0\.042 to 0\.043 ROUGE\-1 F with intervals bounded well away from zero, and on ILSUM it trails by 0\.265 with an interval that does not approach zero\. It is important that TextRank\[[6](https://arxiv.org/html/2609.29090#bib.bib6)\], a second and entirely separate unsupervised extractive method, fails in exactly the same way: it trails Lead\-3 by 0\.025 on XL\-Sum and by 0\.254 on ILSUM, both significant\. Because our ROUGE agrees with the authors’ scorer to four decimal places, and because a method we did not write reproduces the same defeat, the result is best read as a class\-level failure of salience\-driven unsupervised extraction against a lead baseline on these corpora, not as a defect in our particular re\-implementation\. ## 3Ablation Isolating the contribution of each feature explains the shortfall precisely\. When the ranking uses the position feature alone, the system becomes identical to the lead baseline: on XL\-Sum “Ours\-position\-only” scores 0\.228 ROUGE\-1 F, the same value as Lead\-3 to three decimals, with a paired gap of\+\+0\.000 and an interval of \[\+\+0\.000,\+\+0\.000\] \(Tables[1](https://arxiv.org/html/2609.29090#S2.T1)and[2](https://arxiv.org/html/2609.29090#S2.T2)\)\. This equivalence is expected — the position score is monotone in sentence index, so selecting the top\-scoring sentences under position alone returns the article opening — but it is informative, because position\-only is the best\-performing configuration of our system, above the full seven\-feature model\. Conversely, removing position and retaining the other six features \(“Ours\-minus\-position”\) yields 0\.182, the weakest configuration in the table and significantly below Lead\-3 by 0\.046\. The equal\-weight full system sits between these poles at 0\.185–0\.186, indicating that when position is only one summand of seven its advantage is largely averaged away\. A validation\-set weight sweep confirms the pattern is monotone rather than incidental\. Increasing the position weight while holding the remaining features at unit weight raises ROUGE\-1 F monotonically on both corpora: on XL\-Sum the seven\-feature family rises from 0\.191 at unit position weight to 0\.235 at weight sixteen, and a position\-plus\-TF–IDF\-only family rises from 0\.211 to 0\.239 over the same range; on ILSUM the position\-plus\-TF–IDF family rises from 0\.347 at unit weight to 0\.560 at weight sixteen \(Table[5](https://arxiv.org/html/2609.29090#S3.T5)\)\. In every case the published equal\-weight setting is the worst point on the curve and heavier position weighting is monotonically better, up to a ceiling\. Table 5:Validation\-set ROUGE\-1 F under a position\-weight sweep \(three\-sentence budget\)\. XL\-Sum figures are the seven\-feature and position\+TF–IDF families; ILSUM figures are the position\+TF–IDF family\.The tuned ceiling is a win on neither corpus\. The single configuration selected on validation — position\+TF–IDF at position weight sixteen — was evaluated once on each held\-out test slice, and the outcome is summarised in Table[6](https://arxiv.org/html/2609.29090#S3.T6)\. On XL\-Sum it reaches 0\.231 ROUGE\-1 F against Lead\-3’s 0\.231, a paired gap of−\-0\.000 whose interval \[−\-0\.002,\+\+0\.001\] comfortably contains zero: a statistical tie, and a genuine equivalence rather than an underpowered comparison, given how narrow the interval is\. On ILSUM it reaches 0\.516 against 0\.522, a paired gap of−\-0\.005 whose interval \[−\-0\.011,−\-0\.000\] excludes zero only at its upper bound; the deficit is therefore statistically detectable but negligible \(≤\\leq0\.011 R1\-F\), leaving the tuned system at best equal to Lead\-3 and never above it\. The cross\-corpus reading is accordingly a tie on XL\-Sum, a marginal deficit on ILSUM, and a win on neither\. Careful weighting nonetheless recovers all of the ground the equal\-weight configuration loses — a real improvement over the published setting that decisively beats the equal\-weight, TextRank and random systems — so the best attainable behaviour of this feature family is to reproduce lead selection rather than to surpass it\. Table 6:Tuned configuration \(position\+TF–IDF, position weight sixteen\) against Lead\-3 on each held\-out test slice; ROUGE\-1 F with 95% bootstrap CIs and the paired gap\. Statistically detectable but negligible \(≤\\leq0\.011 R1\-F\) on ILSUM and a tie on XL\-Sum: at best equal, never above\. ## 4Mechanism The reason is visible in where each system reads\. Table[7](https://arxiv.org/html/2609.29090#S4.T7)gives, for the XL\-Sum ablation documents, the distribution of the source\-article positions of selected sentences, normalised so that zero is the article start and one its end, binned into deciles, together with the mean normalised position\. Position\-only and Lead\-3 are indistinguishable, drawing 0\.71 of their sentences from the first decile and reaching a mean position of 0\.070\. The minus\-position system is the most back\-loaded of all, with a mean position of 0\.501 and its single largest mass in the final decile, because once position is removed the embedding\-cosine, TF–IDF, length and proper\-noun features pull selection towards lexically dense sentences that in news prose sit in the body and tail\. The equal\-weight systems fall in between at a mean of roughly 0\.39, retaining only a faint lead tilt \(0\.20 in the first decile against the lead’s 0\.71\) — the quantitative signature of position being diluted to one\-seventh of the score\. Random selection is near\-uniform at 0\.488 and TextRank only slightly lead\-tilted at 0\.441\. Table 7:Fraction of selected sentences by source\-position decile on the XL\-Sum ablation set \(D1 = first tenth of the article, D10 = last tenth\), with mean normalised position\.A single document makes the failure mode concrete\. In XL\-Sum test document 179, a seventeen\-sentence report on an India–Bangladesh cricket Test, the reference summary is*“bharat aur bangladesh ke khilaf dusra Test shukravar se Chatgaon mein shuru ho raha hai\. pahla Test jitne ke bad bharatiya team do Test maichon ki series 2\-0 se jitne ke liye maidan par utregi\.”*\(“The second Test against India and Bangladesh begins on Friday in Chittagong\. After winning the first Test, the Indian team will take the field to win the two\-match Test series 2–0\.”\) — a two\-sentence framing of the fixture drawn from the article’s opening\. Lead\-3 selects sentences 0, 1 and 2, which introduce the match and its context\. The minus\-position system instead selects sentences 13, 15 and 16 \(normalised positions 0\.81, 0\.94 and 1\.00\), of which two are the squad lists: sentence 15 is*“bharatiya team \(inmen se chuni jayegi\) Saurabh Ganguly \(kaptan\), Virender Sehwag, Gautam Gambhir, Rahul Dravid, Sachin Tendulkar …”*\(“Indian team \(to be selected from among these\): Saurav Ganguly \(captain\), Virender Sehwag, Gautam Gambhir, Rahul Dravid, Sachin Tendulkar …”\) and sentence 16 the corresponding Bangladesh roster\. These sentences are exactly what the non\-position features reward: a roster is saturated with proper nouns, is long, and — because player names recur across the article — scores highly on both TF–IDF mass and aggregate embedding cosine\. Every feature the method treats as a proxy for salience is maximised by a list of names that carries almost none of the article’s summary\-worthy content, and the reference, being a lead paraphrase, shares almost no vocabulary with it\. The proper\-noun, length and TF–IDF features are not merely uninformative here; they are actively misdirected towards the least summary\-like sentences in the document\. ## 5Diagnosis The uncomfortable conclusion is that the benchmarks, not only the method, are the confounder\. The ILSUM Lead\-3 figures are the clearest evidence: a three\-sentence lead achieves ROUGE\-2 F of 0\.451 and ROUGE\-L F of 0\.499 \(Table[3](https://arxiv.org/html/2609.29090#S2.T3)\), which is only possible if the reference reuses the article’s opening almost verbatim at the bigram level\. An ILSUM reference is, in effect, a lightly edited copy of the lead, so any system that does not begin at the top is penalised regardless of whether its selection is informative\. XL\-Sum poses the same obstacle in a different form: its references are single lead\-style sentences, so a three\-sentence extract can match at most a fraction of the reference and the maximal\-overlap strategy is again to take the opening\. Under both designs the target rewards lead reproduction and offers no credit for correctly identifying salient non\-lead content, which is precisely the capability a semantic\-extractive method is intended to provide\. This is why heavier position weighting monotonically helps and why the ceiling is a tie with the lead: on these corpora the score is very nearly a measure of how lead\-like a system is, and no feature that pulls selection away from the opening can be rewarded\. We are careful not to overclaim\. What these experiments establish is narrow and specific: on XL\-Sum Hindi and ILSUM 2\.0 Hindi, at a three\-sentence budget under instrument\-validated ROUGE, the replicated equal\-weight semantic\-extractive method and a comparable TextRank baseline are significantly worse than a lead baseline, and the best weighting of this feature family only matches the lead\. We do not claim that distributional\-semantics extraction is without value in general, nor that these features cannot help on tasks whose references reward non\-lead content; our evidence speaks only to these benchmarks, whose reference style makes the lead nearly unbeatable and therefore cannot discriminate a good content selector from a positional heuristic\. Two further caveats bound the claim: we evaluate at a single three\-sentence budget, so the absolute scores could shift under a different summary length, even though the strength of the lead advantage makes a reversal unlikely\. Moreover, because ROUGE credits lexicalnn\-gram overlap rather than conveyed meaning, its design compounds the lead\-reference bias by rewarding token reuse over paraphrase, which is why the resource we propose below should be scored with complementary measures such as ChrF and BERTScore alongside human references, and not with ROUGE alone\. The failure we document is a failure of measurability as much as of method\. This motivates a purpose\-built Hindi resource\. A benchmark able to reward semantic content selection would need human\-written summaries that are longer than a single lead sentence and deliberately draw on material from across the article rather than its opening, ideally with multiple references per document so that legitimate paraphrase variation is credited rather than penalised\. It would further benefit from sentence\-level faithfulness labels, so that a system choosing an informative but non\-lead sentence can be distinguished from one that has simply drifted into the article’s tail\. Only against such a resource can the question in our title be answered on its own terms, rather than settled in advance by the reference style of the available data\. ## Software and data availability The Hindi summariser, the Devanagari\-aware ROUGE module and the full evaluation harness used for every number in this paper are released as HinExSum\[[8](https://arxiv.org/html/2609.29090#bib.bib8)\], an open\-source Python package under the MIT licence, at[https://github\.com/mudasirmohd/HinExSum](https://github.com/mudasirmohd/HinExSum)\. A reproducible Code Ocean capsule containing the code, the environment specification and the evaluation entry points is archived at[https://doi\.org/10\.24433/CO\.6462433\.v1](https://doi.org/10.24433/CO.6462433.v1)\. The corpora are third\-party resources and are not redistributed here: the Hindi portion of XL\-Sum\[[2](https://arxiv.org/html/2609.29090#bib.bib2)\]and FIRE ILSUM 2\.0 Hindi\[[3](https://arxiv.org/html/2609.29090#bib.bib3)\]are obtained from their public Hugging Face distributions, and the fastTextcc\.hi\.300\.binvectors\[[4](https://arxiv.org/html/2609.29090#bib.bib4)\]are downloaded from the fastText site as described in the repositoryREADME\. ## References - \[1\]M\. Mohd, R\. Jan, and M\. Shah\.Text document summarization using word embedding\.*Expert Systems with Applications*, 143:112958, 2020\.doi:[10\.1016/j\.eswa\.2019\.112958](https://doi.org/10.1016/j.eswa.2019.112958)\. - \[2\]T\. Hasan, A\. Bhattacharjee, M\. S\. Islam, K\. Mubasshir, Y\.\-F\. Li, Y\.\-B\. Kang, M\. S\. Rahman, and R\. Shahriyar\.XL\-Sum: Large\-scale multilingual abstractive summarization for 44 languages\.In*Findings of the Association for Computational Linguistics: ACL\-IJCNLP 2021*, pages 4693–4703, 2021\. - \[3\]S\. Satapara, B\. Modha, S\. Modha, and P\. Mehta\.FIRE 2022 ILSUM track: Indian language summarization\.In*Proceedings of the 14th Annual Meeting of the Forum for Information Retrieval Evaluation \(FIRE\)*, 2022\. - \[4\]E\. Grave, P\. Bojanowski, P\. Gupta, A\. Joulin, and T\. Mikolov\.Learning word vectors for 157 languages\.In*Proceedings of the 11th International Conference on Language Resources and Evaluation \(LREC\)*, 2018\. - \[5\]P\. Qi, Y\. Zhang, Y\. Zhang, J\. Bolton, and C\. D\. Manning\.Stanza: A Python natural language processing toolkit for many human languages\.In*Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations*, pages 101–108, 2020\. - \[6\]R\. Mihalcea and P\. Tarau\.TextRank: Bringing order into text\.In*Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing \(EMNLP\)*, pages 404–411, 2004\. - \[7\]C\.\-Y\. Lin\.ROUGE: A package for automatic evaluation of summaries\.In*Text Summarization Branches Out: Proceedings of the ACL\-04 Workshop*, pages 74–81, 2004\. - \[8\]M\. Mohd\.HinExSum: A Hindi extractive summariser using distributional semantics, version 1\.0\.0 \[computer software\], 2026\.[https://github\.com/mudasirmohd/HinExSum](https://github.com/mudasirmohd/HinExSum)\.Code Ocean capsule: doi:[10\.24433/CO\.6462433\.v1](https://doi.org/10.24433/CO.6462433.v1)\.
Similar Articles
Consistency Analysis of Sentiment Predictions using Syntactic & Semantic Context Assessment Summarization (SSAS)
This paper presents SSAS (Syntactic & Semantic Context Assessment Summarization), a framework designed to improve consistency in LLM-based sentiment prediction by reducing noise and variance through hierarchical classification and iterative summarization. Empirical evaluation on three industry-standard datasets shows up to 30% improvement in data quality and reliability for enterprise decision-making.
A Hybrid Hierarchical 1D-CNN-BiLSTM Framework for Extractive Summarization of Biomedical and Clinical Text
The paper introduces a hybrid hierarchical 1D-CNN-BiLSTM framework for extractive summarization of biomedical and clinical text, designed to preserve factuality by selecting sentences directly from source documents rather than generating new text.
LexLattice: Multilingual Extractive Summarization via Neural Cellular Automata on Document Hierarchies
LexLattice introduces a multilingual extractive summarization method using neural cellular automata on document hierarchies, achieving state-of-the-art ROUGE scores on the EUR-Lex-Sum dataset with a compact model surpassing large instruction-tuned baselines.
A Grounded and Decomposed Framework for Relation-Level Hallucination Evaluation in Abstractive Summarization
This paper presents a grounded and decomposed framework for evaluating relation-level hallucinations in abstractive summarization, introducing a normalized Relation Hallucination Index (RHI) with linguistically informed relation extraction enhancements.
Abstractiveness Metrics for Evaluating Text Summarization: A Refined Formulation with Empirical Validation
This paper introduces Reference Abstraction (RA), Summary Abstraction (SA), and Abstraction Ratio (AR) metrics to quantify abstractiveness in text summarization, using harmonic mean of document lengths and cubic non-overlap factor. Empirical evaluation on XSUM with four models shows the metrics effectively discriminate between extractive and abstractive summaries, and flag potential hallucination.