bert

Tag

Cards List
#bert

Reliable Financial Named Entity Recognition under Domain Shift

arXiv cs.CL · 2026-08-21 Cached

This paper studies confidence estimation and selective prediction for financial named entity recognition under domain shift, evaluating BERT and LoRA-tuned Qwen models to enhance reliability across different input distributions like SEC filings and social media.

0 favorites 0 likes
#bert

Transformer Models for Text Summarization: A Comparative Study of BART, BERT, and RoBERTa

arXiv cs.CL · 2026-08-21 Cached

A comparative study of BART, BERT, and RoBERTa for text summarization, examining their architectures and suitability for extractive and abstractive summarization tasks.

0 favorites 0 likes
#bert

Assessing Reliability of BERT-Based Models on Question Answering Tasks

arXiv cs.CL · 2026-08-12 Cached

This paper evaluates the reliability of BERT-based QA models (RoBERTa, ALBERT, DistilBERT) under Monte Carlo Dropout and input paraphrasing, finding RoBERTa more consistent and validating MCD as a reliability metric.

0 favorites 0 likes
#bert

STCAD: Scalable Trajectory Clustering and Anomaly Detection on Terabyte-Scale AIS Data

arXiv cs.LG · 2026-08-12 Cached

Presents STCAD, a scalable framework using BERT-based encoding and CURE clustering to perform trajectory clustering and anomaly detection on terabyte-scale AIS maritime data, demonstrating stable clusters and clear separation of anomalous vessel behavior.

0 favorites 0 likes
#bert

Detection of Self-Introductions in Legislative Testimony

arXiv cs.CL · 2026-08-11 Cached

This paper presents a machine learning pipeline for detecting self-introductions in legislative committee testimony, using features like bag-of-words and BERT probabilities. XGBoost achieves the best F1 score of 0.9747, improving further with BERT-augmented features.

0 favorites 0 likes
#bert

CNM-BERT: A Drop-In Structural Embedding for Chinese Characters via Ideographic Description Sequences

arXiv cs.CL · 2026-08-07 Cached

This paper proposes CNM, a lightweight augmentation that injects discrete compositional structure of Chinese characters into BERT via Ideographic Description Sequences, improving performance on rare and out-of-vocabulary characters while preserving general NLU accuracy.

0 favorites 0 likes
#bert

Joint Affine Spectral Shaping: Coupling Weight and Bias Updates Beyond Weight-Only Muon

arXiv cs.LG · 2026-08-05 Cached

This paper proposes Joint Affine Spectral Shaping (JRI), which extends weight-only spectral optimizers like Muon to jointly update weight and bias in affine layers, showing small but consistent accuracy improvements on a BERT-mini IMDb classification task.

0 favorites 0 likes
#bert

@jxmnop: we didn't ever need to invent Masked Language Modeling, I don't think. it was a bit silly by construction. in most alte…

X AI KOLs Following · 2026-08-04

A tweet argues that masked language modeling was unnecessary and that autoregressive models would have sufficed, with a nod to BERT.

0 favorites 0 likes
#bert

Automated Multilabel Mpox Research Classification with Explainable Transformer Models

arXiv cs.CL · 2026-07-30 Cached

This paper proposes an automated multilabel classification system for Mpox research articles using BERT, achieving 97% accuracy, and employs SHAP for explainability. The system aims to help researchers and healthcare workers quickly find relevant information.

0 favorites 0 likes
#bert

BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi

arXiv cs.CL · 2026-07-28 Cached

This paper compares fine-tuned MahaBERT-based models with large language models (Gemini, LLaMA-3.3-70B, Gemma) for Marathi named entity recognition, finding that the specialized BERT models significantly outperform the LLMs, achieving F1-scores of 0.88–0.91 versus 0.57–0.69.

0 favorites 0 likes
#bert

@HuggingModels: Ever needed to pull rich meaning from text without fine tuning? Meet MyAwesomeModel, a BERT based feature extractor tha…

X AI KOLs Timeline · 2026-07-23 Cached

MyAwesomeModel, a BERT-based feature extractor, converts sentences into dense vectors for semantic search, clustering, and NLP pipelines.

0 favorites 0 likes
#bert

Translation as Augmentation: Effect of Translated Data on Assessment of Difficulty

arXiv cs.CL · 2026-07-22 Cached

This paper proposes a cross-lingual data augmentation strategy that uses machine translation to transfer expert-annotated difficulty labels from high-resource languages to low-resource languages. Experiments with BERT-based regression models show that augmenting scarce native data with translated corpora significantly improves the accuracy of text difficulty assessment.

0 favorites 0 likes
#bert

Team DACTYL at PAN 2026: Bayesian Data Mixing and Empirical X-risk Minimization for AI-text Detection

arXiv cs.CL · 2026-07-21 Cached

This paper presents methods for detecting AI-generated text using Bayesian data mixing and empirical X-risk minimization, achieving high performance on OOD detection with ModernBERT-large and MCGrad classifiers.

0 favorites 0 likes
#bert

Candidate Attended Dialogue State Tracking Using BERT

arXiv cs.CL · 2026-07-20 Cached

The paper presents a scalable framework for multi-domain dialogue state tracking using BERT, achieving zero-shot generalization and improving performance on the SGD dataset.

0 favorites 0 likes
#bert

Translation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource Languages

arXiv cs.CL · 2026-07-15 Cached

This paper investigates the feasibility of using translation-based fine-tuning as a resource-efficient alternative to native-language BERT models for low-resource languages, finding it comparable or superior in 53.3% of cases across six NLP tasks.

0 favorites 0 likes
#bert

Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoders

arXiv cs.CL · 2026-07-10 Cached

This paper proposes a Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoder to extract universal features across independently trained BERT models, improving cross-seed feature alignment beyond post-hoc methods.

0 favorites 0 likes
#bert

Separating Representation from Reconstruction Enables Scalable Text Encoders

arXiv cs.CL · 2026-07-07 Cached

CrossBERT decouples representation learning from token reconstruction, enabling higher masking ratios and better sample efficiency, outperforming BERT on MTEB and GLUE benchmarks.

0 favorites 0 likes
#bert

@goyalayus: Actually, you don't need any of these but just these four books, then you can go into any niche and pick up things. (I …

X AI KOLs Timeline · 2026-07-04 Cached

A tweet from @goyalayus suggests that four unspecified books are sufficient to learn any niche, followed by a quoted recommendation from Praveen Kumar Verma to read foundational LLM papers in a specific order, including Attention Is All You Need, BERT, GPT, GPT-2, Scaling Laws, and GPT-3.

1 favorites 0 likes
#bert

BamiBERT: A New BERT-based Language Model for Vietnamese

arXiv cs.CL · 2026-07-03 Cached

BamiBERT is a new BERT-based pre-trained language model for Vietnamese that addresses limitations of PhoBERT, supporting longer context and operating without word segmentation, achieving state-of-the-art results on multiple Vietnamese benchmarks.

0 favorites 0 likes
#bert

@TheTuringPost: A great source to understand or refresh Transformer architecture It explains how transformers process text token by tok…

X AI KOLs Timeline · 2026-07-03 Cached

Promotes an educational resource explaining Transformer architecture, covering token embeddings, self-attention, residual connections, and connections to GPT and BERT.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback