behavioral-alignment

Tag

Cards List
#behavioral-alignment

When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety

Hugging Face Daily Papers ↗ · 3d ago Cached

The paper compares behavioral safeguards like DPO with representation engineering methods for LLM safety, finding that representation engineering offers practical advantages in specific conditions and complements behavioral approaches.

0 favorites 0 likes
#behavioral-alignment

From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation

arXiv cs.AI ↗ · 2026-08-13 Cached

The paper introduces a behavioral alignment framework for personalized LLM judges in recommendation evaluation, addressing bidirectional rationalization where off-the-shelf LLMs argue both for and against user engagement on the same item. Their fine-tuned and preference-optimized approach achieves a 32.19% Macro-F1 lift over zero-shot and matches production feature-engineered baselines.

0 favorites 0 likes
#behavioral-alignment

On the use of foundation models in cognitive science

arXiv cs.CL ↗ · 2026-08-11 Cached

This perspective paper from arXiv articulates a four-stage inferential framework for evaluating foundation models as cognitive and developmental models, emphasizing that behavioral alignment alone is insufficient and must be embedded within theoretical commitments and contrastive evaluation.

0 favorites 0 likes
#behavioral-alignment

Improving language model behavior by training on a curated dataset

OpenAI Blog ↗ · 2021-06-10 Cached

OpenAI research demonstrates that language model behavior can be significantly improved through fine-tuning on small, curated datasets (<100 examples) targeting specific behavioral values, with effectiveness increasing at larger model scales. The approach provides users with tools to align models with Charter-compatible values for their specific applications.

0 favorites 0 likes
← Back to home

Submit Feedback