modernbert

Tag

Cards List
#modernbert

My frozen-encoder decision heads were reading 3 tokens per label: fixing the option budget took banking77 from 67.5% to 76.0% [P]

Reddit r/MachineLearning ↗ · 5d ago

The author distills decisions from LLMs into small heads on a frozen ModernBERT encoder, improving accuracy on the banking77 task from 67.5% to 76.0% by fixing the option budget, and provides insights on token reading and teacher agreement.

0 favorites 0 likes
#modernbert

Why Advanced Encoders Lag on Sparse Retrieval? The Answer and an Approach to Bridging Vocabulary Gaps

arXiv cs.AI ↗ · 2026-07-02 Cached

This paper identifies a vocabulary gap as the root cause why advanced encoders like ModernBERT underperform in learned sparse retrieval, and proposes Vocabulary Transfer (VT), a model-agnostic framework that migrates encoders to sparse-friendly vocabularies, achieving state-of-the-art on the BEIR benchmark.

0 favorites 0 likes
#modernbert

BERTomelo: Your Portuguese Encoder Best Friend

arXiv cs.CL ↗ · 2026-06-30 Cached

This paper introduces BERTomelo, a next-generation monolingual encoder pre-trained for Portuguese using the ModernBERT architecture, achieving superior performance on downstream tasks like STS and NER compared to previous Portuguese and multilingual models.

0 favorites 0 likes
#modernbert

Legal Domain Adaptation of Modern BERT Models

arXiv cs.CL ↗ · 2026-06-30 Cached

This paper explores domain adaptation of ModernBERT models in the legal domain by further pre-training on US court opinions, achieving significant improvements over the vanilla model and releasing the checkpoints publicly.

0 favorites 0 likes
#modernbert

Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM Adversarial Evaluation

arXiv cs.CL ↗ · 2026-06-25 Cached

This paper systematically compares fine-tuned encoder classifiers (ModernBERT family) against decoder-based safety judges for LLM adversarial evaluation, finding that encoders can offer a cost- and latency-efficient alternative without significant performance loss.

0 favorites 0 likes
#modernbert

Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States

Hugging Face Daily Papers ↗ · 2026-06-17 Cached

Introduces LOCUS, a comprehensive corpus of U.S. local ordinance codes designed to enable machine-readable legal AI research, covering codes from 9,239 cities and counties with ModernBERT-based classifiers for analysis.

0 favorites 0 likes
#modernbert

Show HN: Context-aware Japanese furigana using Sudachi and ModernBERT

Hacker News Top ↗ · 2026-05-29 Cached

EZFurigana is a free, privacy-focused tool that uses Sudachi and ModernBERT to add context-aware furigana to Japanese text, supporting various input formats and customization options.

0 favorites 0 likes
#modernbert

ACL-Verbatim: hallucination-free question answering for research

Hugging Face Daily Papers ↗ · 2026-05-20 Cached

ACL-Verbatim introduces a family of lightweight extractive models for grounded RAG that return exact text spans from source, outperforming larger LLM-based extractors.

0 favorites 0 likes
#modernbert

Introducing the Ettin Reranker Family

Reddit r/LocalLLaMA ↗ · 2026-05-19 Cached

Introducing the Ettin Reranker family: six new state-of-the-art CrossEncoder rerankers at various sizes, built on ModernBERT encoders, with open-source data and training recipe.

0 favorites 0 likes
#modernbert

HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools

arXiv cs.CL ↗ · 2026-05-19 Cached

HyDRA is a hybrid dynamic routing architecture for heterogeneous LLM pools that predicts fine-grained capability requirements per query and selects the cheapest capable model via shortfall matching, achieving up to 72.5% cost savings with quality maintained. It is deployed in GitHub Copilot's VS Code Chat auto-mode and decouples routing from model catalog, requiring no retraining when models change.

0 favorites 0 likes
#modernbert

A Causal Language Modeling Detour Improves Encoder Continued Pretraining

Hugging Face Daily Papers ↗ · 2026-05-12 Cached

This paper demonstrates that switching from Masked Language Modeling to Causal Language Modeling during encoder adaptation improves downstream performance on biomedical texts. The authors release ModernBERT-bio and ModernCamemBERT-bio as state-of-the-art biomedical encoders.

0 favorites 0 likes
← Back to home

Submit Feedback