chinese-nlp

Tag

Cards List
#chinese-nlp

CMNIE: An Information Extraction Benchmark for Chinese Military News

arXiv cs.CL · 2026-09-11 Cached

CMNIE is a new information extraction benchmark for Chinese military news that jointly annotates events, entities, and relations to evaluate models on schema adherence and exact span matching.

0 favorites 0 likes
#chinese-nlp

GUIDE: Generative Unsupervised Chinese Query Correction via Phonetic and Visual Shared-ID Encoding

arXiv cs.CL · 2026-08-27 Cached

GUIDE is a generative unsupervised framework for Chinese query correction that uses phonetic and visual shared-ID encoding to constrain corrections and adapt to changing vocabularies, outperforming baselines in experiments and online A/B testing.

0 favorites 0 likes
#chinese-nlp

New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs

arXiv cs.CL · 2026-08-14 Cached

This paper proposes SeTox, a search-augmented LLM framework for detecting implicit toxicity in Chinese neologisms by leveraging real-time web context and public consensus. Experiments show that even 3B-scale models outperform larger recent models on this task.

0 favorites 0 likes
#chinese-nlp

CNM-BERT: A Drop-In Structural Embedding for Chinese Characters via Ideographic Description Sequences

arXiv cs.CL · 2026-08-07 Cached

This paper proposes CNM, a lightweight augmentation that injects discrete compositional structure of Chinese characters into BERT via Ideographic Description Sequences, improving performance on rare and out-of-vocabulary characters while preserving general NLU accuracy.

0 favorites 0 likes
#chinese-nlp

Cross-Platform Chinese Offensive Comment Detection via Dual-Threshold Hard Example Mining

arXiv cs.CL · 2026-06-29 Cached

This paper proposes a dual-threshold hard example mining strategy for cross-platform Chinese offensive comment detection, addressing performance degradation due to domain shift. The method fine-tunes a RoBERTa model on the COLD dataset and adapts it to four Chinese social media platforms with minimal labeled data.

0 favorites 0 likes
#chinese-nlp

A Reproducible Multi-Architecture Baseline for Token-Level Chinese Metaphor Identification under the MIPVU Framework

arXiv cs.CL · 2026-05-11 Cached

This paper establishes a reproducible multi-architecture baseline for token-level Chinese metaphor identification using the MIPVU framework and the PSU Chinese Metaphor Corpus. It compares encoder models like RoBERTa and MelBERT against the Qwen3.5-9B generative model, releasing code and data to facilitate future research.

0 favorites 0 likes
#chinese-nlp

CFMS: Towards Explainable and Fine-Grained Chinese Multimodal Sarcasm Detection Benchmark

arXiv cs.CL · 2026-04-21 Cached

Researchers from Peking University introduce CFMS, the first fine-grained Chinese multimodal sarcasm detection benchmark with 2,796 image-text pairs and a triple-level annotation framework (sarcasm identification, target recognition, explanation generation), along with a novel RL-augmented in-context learning method (PGDS) that significantly outperforms existing baselines.

0 favorites 0 likes
← Back to home

Submit Feedback