nlp-research

Tag

Cards List
#nlp-research

Beyond Unsafe Detection: Counterfactually Anchored Evidence Attribution for Multi-Turn LLM Safety Failures

arXiv cs.CL ↗ · 5d ago Cached

This paper proposes a counterfactually anchored evidence attribution method for identifying safety failures in multi-turn LLM conversations, achieving high detection performance and low false positives.

0 favorites 0 likes
#nlp-research

Do small language models know what they don't know?

arXiv cs.CL ↗ · 2026-09-21 Cached

This paper evaluates entropy-based confidence signals for small language models, finding that semantic entropy provides a viable method to improve accuracy by up to 50 percentage points through selective routing to larger models.

0 favorites 0 likes
#nlp-research

Nepali Legal Expertise through Generative and Extractive Pre-trained Transformers (NepLEGiT)

arXiv cs.CL ↗ · 2026-09-16 Cached

The paper introduces NepLEGiT, a specialized small language model pre-trained from scratch on Nepali legal text to enhance legal knowledge accessibility and service delivery in Nepal.

0 favorites 0 likes
#nlp-research

Bangla Sentence Function Classification: Corpus Development, Model Benchmarking, and Interpretability

arXiv cs.CL ↗ · 2026-09-15 Cached

This paper introduces a corpus of 10,000 annotated Bangla sentences for function classification and benchmarks various models, achieving 0.95 accuracy with a double-level ensemble using TF-IDF features.

0 favorites 0 likes
#nlp-research

5-Dialects-BN: Unmasking the Impact of Transliteration on Bangla Dialectal LLMs

arXiv cs.CL ↗ · 2026-09-10 Cached

The paper presents 5-Dialects-BN, a manually annotated benchmark dataset for five Bangla dialects with aligned transliterations and translations, designed to improve evaluation of dialect-aware LLMs.

0 favorites 0 likes
#nlp-research

The Curse of Multilinguality in Lexical Normalization

arXiv cs.CL ↗ · 2026-09-02 Cached

This paper investigates the curse of multilinguality in lexical normalization, finding that training a single model on multiple languages leads to decreased per-language accuracy, with optimal performance when languages are trained in small groups.

0 favorites 0 likes
#nlp-research

Attribute-Based Activation Steering of LLMs for Group-Specific Explanation Generation

arXiv cs.CL ↗ · 2026-09-01 Cached

This paper proposes attribute-based activation steering to tailor LLM explanations to specific groups, achieving better specificity and factuality compared to prompting and state-of-the-art baselines.

0 favorites 0 likes
#nlp-research

Why Current XAI Is Not Enough for Arabic NLP: A Critical Survey of the Explainability Gap

arXiv cs.CL ↗ · 2026-08-28 Cached

This survey identifies three critical gaps in explainable AI for Arabic NLP—method, task, and linguistic—and proposes a taxonomy and research agenda for linguistically grounded explanations.

0 favorites 0 likes
#nlp-research

Padamitra: Grounded Glossary Generation for Classical Sanskrit

arXiv cs.CL ↗ · 2026-08-27 Cached

This paper introduces grounded glossary generation for Classical Sanskrit, a task involving recovering Sanskrit phrases and producing translation-grounded meanings from sloka-translation pairs. It constructs a benchmark from Hindu texts and evaluates various AI models, finding that instruction fine-tuning improves performance, with morphological modeling identified as a key challenge.

0 favorites 0 likes
#nlp-research

More Computational Resources Do Not Ensure Higher Scholarly Impact: Evidence from Leading NLP Conference Papers

arXiv cs.CL ↗ · 2026-08-25 Cached

This paper analyzes NLP conference papers from 2020 to 2025 to examine the relationship between reported GPU resources and scholarly impact, finding that while resources are associated with higher citations, they explain little of the variance in impact.

0 favorites 0 likes
#nlp-research

@tomaarsen: At 1400x cheaper, I know I'm sticking to embeddings (dense, sparse, multi-vector), plus hybrid (incl. bm25) and reranke…

X AI KOLs Following ↗ · 2026-08-21 Cached

A discussion on the cost-effectiveness of embeddings versus LLMs, referencing a study called 'embedder's dilemma' that finds LLMs outperform embedding models at significantly higher cost.

0 favorites 0 likes
#nlp-research

SuTRA : Structurally-Unified Tokenization with Root Awareness

arXiv cs.CL ↗ · 2026-08-20 Cached

SuTRA is a morphology-aware tokenization algorithm that preserves akshara indivisibility for Indic languages, reducing morphological shattering and achieving improvements in machine translation metrics over standard BPE methods.

0 favorites 0 likes
#nlp-research

TokEval: A Tokenizer Evaluation Suite

arXiv cs.CL ↗ · 2026-08-19 Cached

This paper introduces TokEval, a framework for evaluating language model tokenizers using intrinsic metrics that correlate with downstream task performance.

0 favorites 0 likes
#nlp-research

There is No Theoretical Curse of Multilinguality For Embedding Space Structure

arXiv cs.CL ↗ · 2026-08-19 Cached

This paper proves that there is no theoretical curse of multilinguality for embedding space structure, showing that the minimum dimensionality required grows only logarithmically with the number of languages, suggesting empirical issues stem from data and training conditions.

0 favorites 0 likes
#nlp-research

When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text

arXiv cs.CL ↗ · 2026-08-18 Cached

This paper presents an empirical study on sentiment classifier behavior with sarcastic and AI-paraphrased social text, revealing lower confidence on sarcasm, higher accuracy on AI paraphrases, and an abstention method that improves performance by handling low-confidence inputs.

0 favorites 0 likes
#nlp-research

When More Becomes Less: Position-Dependent Repetition Effects in Language Models

arXiv cs.CL ↗ · 2026-08-06 Cached

This paper shows that repetition effects in language models depend on readout position: adjacent repetition boosts target probability, while displaced repetition produces an inverted-U curve. The finding challenges assumptions in cloze-style probing and is validated across multiple models and languages.

0 favorites 0 likes
#nlp-research

Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm

arXiv cs.CL ↗ · 2026-07-31 Cached

This arXiv paper proposes capability-sustaining emotional dialogue (CSED) as a longitudinal research paradigm for emotional support systems, arguing that current approaches focus on immediate relief and neglect long-term user capabilities. A literature audit shows most systems overlook longitudinal outcomes, suggesting a new agenda for data, models, evaluation, and governance.

0 favorites 0 likes
#nlp-research

Characterizing Narrative Content in Web-scale LLM Pretraining Data

Hugging Face Daily Papers ↗ · 2026-06-17 Cached

A fine-grained study of narrative features in web-scale LLM pretraining data, introducing NarraBERT and NarraDolma to measure narrative patterns and their distribution across sources.

0 favorites 0 likes
#nlp-research

The First Token Knows: Single-Decode Confidence for Hallucination Detection

Hugging Face Daily Papers ↗ · 2026-05-06 Cached

This paper introduces a method for detecting hallucinations in large language models by leveraging the confidence of the first generated token, requiring only a single decode step.

0 favorites 0 likes
← Back to home

Submit Feedback