preference-alignment

Tag

Cards List
#preference-alignment

LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation

arXiv cs.CL · 3d ago Cached

LOCUS is a task-aware low-rank post-training method that reduces output token length in language models while maintaining preference alignment, achieving up to 39.84% reduction on Pythia-2.8B with minimal parameter updates.

0 favorites 0 likes
#preference-alignment

BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences

arXiv cs.CL · 2026-07-31 Cached

BridgeAlign proposes a preference-alignment pipeline for humanities and social sciences, generating 210k synthetic preference samples and enabling Qwen3-8B to achieve strong results across 17 benchmarks.

0 favorites 0 likes
#preference-alignment

On the Limits of Steering Vectors for Preference-Aligned Generation

arXiv cs.CL · 2026-07-03 Cached

This paper systematically studies the limitations of steering vectors for controlled text generation, finding that their effectiveness varies across traits, degrades on task transfer, and suffers from composition tradeoffs.

0 favorites 0 likes
#preference-alignment

PAPA: Online Personalized Active Preference Alignment

arXiv cs.LG · 2026-07-02 Cached

This paper introduces Personalized Active Preference Alignment (PAPA), a method for fine-tuning diffusion models using real-time user feedback without a parameterized reward model, enhancing efficiency in personalized tasks like recommendations and image generation.

0 favorites 0 likes
#preference-alignment

\textsc{DiARC}: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models

arXiv cs.CL · 2026-06-26 Cached

DiARC is a method that improves the reasoning ability of large language models on ARC-like tasks by constructing preference pairs from positive and negative samples, outperforming baselines across multiple benchmarks.

0 favorites 0 likes
#preference-alignment

CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment

arXiv cs.CL · 2026-06-16 Cached

This paper introduces CHILLGuard, a fine-grained Chinese LLM content safety guardrail built on a new 5-macro, 31-micro category risk taxonomy and a scalable multi-stage data construction pipeline. The model achieves state-of-the-art performance, improving F1 score by 15.92% over existing baselines.

0 favorites 0 likes
#preference-alignment

TuneJury: An Open Metric for Improving Music Generation Preference Alignment

Hugging Face Daily Papers · 2026-06-15 Cached

TuneJury is an open-source pairwise reward model for text-to-music generation that provides calibrated preference scoring and generalizes across multiple downstream applications.

0 favorites 0 likes
#preference-alignment

Spectral Souping: A Unified Framework for Online Preference Alignment

arXiv cs.LG · 2026-05-21 Cached

This paper introduces Spectral Souping, a framework for efficiently aligning LLMs with individual user preferences by discovering a universal spectral representation that enables merging of specialized policies at inference time without costly retraining.

0 favorites 0 likes
#preference-alignment

Implicit Preference Alignment for Human Image Animation

Hugging Face Daily Papers · 2026-05-08 Cached

This paper introduces Implicit Preference Alignment (IPA), a data-efficient post-training framework that improves hand motion generation in human image animation without requiring paired preference data. It utilizes implicit reward maximization and hand-aware local optimization to enhance generation quality while reducing data curation costs.

0 favorites 0 likes
#preference-alignment

Arch-Router: Aligning LLM Routing with Human Preferences

Papers with Code Trending · 2025-06-19 Cached

Arch-Router is a compact 1.5B model that aligns LLM routing with human preferences by mapping queries to user-defined domains and action types, outperforming proprietary models in subjective evaluations.

0 favorites 0 likes
← Back to home

Submit Feedback