nlp

Tag

Cards List
#nlp

Loanword or Switch? The Annotation Boundary, Not the Model, Drives Kazakh-Russian Code-Switching Identification

arXiv cs.CL · 5d ago Cached

This paper argues that the loanword-vs-switch annotation boundary, rather than the choice of model, drives Kazakh-Russian code-switching identification. The authors present a document-level gold LID dataset with an explicit annotation rule and show that naive heuristics and off-the-shelf LID tools fail to distinguish integrated borrowings from genuine code-switches.

0 favorites 0 likes
#nlp

DE-NER : Zero-shot Named Entity Recognition via Dialogue Elicitation of Large Language Models

arXiv cs.CL · 5d ago Cached

Introduces DE-NER, a dialogue elicitation framework for zero-shot named entity recognition that uses self-play between questioner and roleplayer LLMs to clarify entity boundaries, achieving an average 3.75% F1 improvement over baselines.

0 favorites 0 likes
#nlp

The methodology of Constructing the Large-Scale Dataset for Detecting Presuicidal and Anti-Suicidal Signals in Social Media Texts in Russian

arXiv cs.CL · 5d ago Cached

The paper presents a methodology for building a large-scale Russian dataset from social media texts to detect presuicidal and anti-suicidal signals, including annotation guidelines and baseline classification experiments.

0 favorites 0 likes
#nlp

Predicting Startup Exit from Textual Descriptors - A Computational Linguistics Framework

arXiv cs.CL · 5d ago Cached

This study applies computational linguistics and supervised machine learning to predict early-stage startup exits from textual descriptors alone, finding that founder narratives carry predictive signal and introducing a quantifiable Hyping Score.

0 favorites 0 likes
#nlp

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

Hugging Face Daily Papers · 5d ago Cached

Introduces DataSpace, a benchmark for evaluating data agents on verifiable tabular analytics over heterogeneous workspaces, containing 410 cross-language tasks and 7,439 artifacts. Current frontier models achieve only 66.34% accuracy, indicating headroom.

0 favorites 0 likes
#nlp

ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels

Hugging Face Daily Papers · 5d ago Cached

ChronoLens is a framework using multilingual language models and crosscoders to measure historical language change across linguistic levels in 44.98 million documents from five parliamentary traditions (1803–2026), showing that change magnitude and direction vary across languages and time.

0 favorites 0 likes
#nlp

Categorization with NLP

Lobsters Hottest · 5d ago Cached

A developer deep dive into an NLP-based grocery categorization tool, covering input lexing, stemming, and a unigram database approach to classify products into categories.

0 favorites 0 likes
#nlp

@DanKornas: Turning unstructured news from websites, social platforms, email, and feeds into publishable intelligence reports is di…

X AI KOLs Timeline · 6d ago Cached

Taranis AI is an open-source OSINT tool that uses AI and NLP to gather, enrich, and structure unstructured news from multiple sources into publishable intelligence reports.

0 favorites 0 likes
#nlp

Cross-Lingual Transfer for Machine Translation in Turkic Languages

arXiv cs.CL · 6d ago Cached

This paper studies cross-lingual transfer for machine translation among five Turkic languages using pairwise transfer matrices with mT5, finding that transfer is strongest between closely related pairs and that Latinization helps in script-mismatched settings.

0 favorites 0 likes
#nlp

Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives

arXiv cs.LG · 6d ago Cached

This paper benchmarks six frontier LLMs on coding crash attributes from police crash narratives against an official fatal-crash database, finding that while GPT-5.5 High leads among LLMs, simple baselines rival or beat LLM performance and attribute-specific differences outweigh model differences.

0 favorites 0 likes
#nlp

Detecting Experiential Intertextuality Across Migration Routes: Beyond Surface Similarity in French Narratives

arXiv cs.CL · 6d ago Cached

This paper introduces experiential intertextuality detection, using annotation-free methods including zero-shot LLM scoring to identify shared experiential echoes across French migration narratives from different routes.

0 favorites 0 likes
#nlp

M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models

arXiv cs.CL · 6d ago Cached

M3-DuplexBench is a new multi-turn, multilingual, multidomain benchmark for evaluating full-duplex spoken dialogue systems, supporting English and Japanese across casual conversation and question answering domains.

0 favorites 0 likes
#nlp

From Inline Notes to Collected Commentaries: Toward Context-Preserving Organization of Exegetical Knowledge in Classical Chinese Texts

arXiv cs.CL · 6d ago Cached

This paper presents a computational framework for automatically compiling collected commentaries on classical Chinese texts, preserving contextual dependencies of inline notes via prompt chaining and cross-source clustering.

0 favorites 0 likes
#nlp

TELLER: Dual-Path Iterative Preference Optimization for Table Entity Linking

arXiv cs.CL · 6d ago Cached

Presents TELLER, a dual-path iterative preference optimization approach for table entity linking, with direct-answer and reasoning paths that improve accuracy on TableInstruct and MammoTab V2 benchmarks.

0 favorites 0 likes
#nlp

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

arXiv cs.CL · 6d ago Cached

Introduces FinIndices, a large-scale benchmark evaluating LLM data-processing fidelity on uncropped financial statements, revealing knowledge and structural bottlenecks in financial reasoning.

0 favorites 0 likes
#nlp

Imbalanced Data Clustering via Targeted Data Augmentation Using GMM and LLM

arXiv cs.CL · 6d ago Cached

This paper presents a novel unsupervised data augmentation method combining Gaussian Mixture Models and Large Language Models to improve clustering on imbalanced text datasets by generating synthetic documents for underrepresented clusters.

0 favorites 0 likes
#nlp

Hierarchical Copula-Gumbel-Top-\texorpdfstring{$K$}{K} Routing: Two-Sided Dependence Control for Frozen Mixture-of-Experts at Fixed Per-Token Routing Laws

arXiv cs.LG · 6d ago Cached

The paper introduces Hierarchical Copula-Gumbel-Top-K (H-CGA) routing, a method to control joint dependence among token routing choices in frozen Mixture-of-Experts models while keeping each token's routing law exactly fixed. It provides theoretical trade-offs between coherence and load dispersion and validates the mechanism with a small-scale pilot.

0 favorites 0 likes
#nlp

Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents

Hugging Face Daily Papers · 6d ago Cached

This paper introduces SIEVE, a search-inspect-fetch strategy that uses Boolean Query Language to make deep-research agents retrieve only relevant document sections, achieving higher accuracy with 20.7–50.6% fewer tokens across multiple benchmark datasets and agent backbones.

0 favorites 0 likes
#nlp

@antoniolupetti: "Understanding Transformers and Attention Mechanisms" is a very interesting paper that presents the Transformer archite…

X AI KOLs Timeline · 6d ago Cached

A tweet highlights an arxiv paper by Michel Fabrice Serret that introduces Transformers and attention mechanisms from an applied mathematics perspective, covering vectorization, multi-head attention, and methods to reduce attention costs like KV caching and latent attention.

0 favorites 0 likes
#nlp

@stanfordnlp: Throwback Thursday – but you should be watching the 2024 CS224N videos these days (https://youtube.com/playlist?list=PL…

X AI KOLs Following · 2026-07-31 Cached

Stanford NLP Group points to the 2024 CS224N video playlist, reminding viewers to watch the latest lectures on natural language processing with deep learning.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback