transformer-models

Tag

Cards List
#transformer-models

An Expanded Synthetic Conversation Dataset for Multi-Turn Smishing Detection

arXiv cs.CL · 2026-06-08 Cached

This paper presents COVA-X, an expanded synthetic multi-turn conversation dataset for smishing detection, and shows that Longformer now outperforms XGBoost, confirming that transformer models benefit from larger training corpora.

0 favorites 0 likes
#transformer-models

Long Live Fine-Tuning: Task-Specific Transformers Outperform Zero-Shot LLMs for Misinformation Response Classification on Reddit

arXiv cs.CL · 2026-06-04 Cached

Researchers from University of Technology Sydney compare fine-tuned transformers (DistilBERT, RoBERTa) against zero-shot LLMs (Llama variants, Claude, Gemini) for classifying misinformation responses on Reddit, finding that fine-tuned RoBERTa achieves 0.62 macro-F1 versus 0.50 for the best zero-shot model. The study shows that task-specific fine-tuning outperforms larger generalist models, particularly for detecting belief propagation, and that safety-alignment artifacts in frontier models can degrade performance.

0 favorites 0 likes
#transformer-models

Why Financial Institutions Are Converging on Transaction Foundation Models to Build Their Own Intelligence

NVIDIA Blog · 2026-06-02 Cached

Financial institutions are shifting from siloed AI models to unified transaction foundation models built on transformer architectures, as demonstrated by NVIDIA's report and Revolut's PRAGMA model, which improves fraud detection, credit scoring, and recommendations while reducing feature engineering effort.

0 favorites 0 likes
#transformer-models

One Mask to Rule Them All: On Hidden Facts after Editing and How to Find Them

arXiv cs.LG · 2026-05-29 Cached

This paper investigates the internal mechanisms of knowledge editing methods ROME and MEMIT, revealing that edits rely on a common functional subspace of weights and suppress rather than overwrite knowledge, explaining why edits fail to propagate to related facts.

0 favorites 0 likes
#transformer-models

Temporal Concept Drift in Legal Judgment Prediction: Neural Baselines Across Three Epochs of Ukrainian Court Decisions

arXiv cs.CL · 2026-05-26 Cached

This paper investigates temporal concept drift in legal judgment prediction by fine-tuning transformer models on Ukrainian court decisions from three epochs defined by geopolitical disruptions. Findings show severe forward degradation, asymmetry in backward transfer, and that chronological continual learning effectively mitigates forgetting while domain pretraining reduces degradation magnitude.

0 favorites 0 likes
#transformer-models

Language Models Need Sleep

Hugging Face Daily Papers · 2026-05-25 Cached

This paper proposes a sleep-like consolidation mechanism for transformer models that uses fast weights and recurrent passes to improve long-context processing while maintaining inference speed.

0 favorites 0 likes
#transformer-models

Findings of the Counter Turing Test: AI-Generated Text Detection

arXiv cs.CL · 2026-05-21 Cached

This paper presents findings from the Counter Turing Test shared task on AI-generated text detection, with top systems achieving perfect binary classification but significantly lower performance in model attribution, highlighting the difficulty of distinguishing outputs from different large language models.

0 favorites 0 likes
#transformer-models

Ideology Prediction of German Political Texts

arXiv cs.CL · 2026-05-15 Cached

The paper proposes a transformer-based model to predict political ideology of German political texts on a continuous left-to-right spectrum. The study compares 13 models and finds DeBERTa-large and Gemma2-2B perform best on different tasks.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback