helpfulness

Tag

Cards List
#helpfulness

Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training

arXiv cs.CL · 2026-06-26 Cached

This paper investigates how post-training (SFT and RL) on helpfulness vs. coding data degrades compassion values in Llama 3.1 8B models that were mid-trained on compassion-oriented data. It finds that helpfulness training significantly reduces animal compassion and general moral reasoning compared to coding training, though the reasoning effect does not transfer cross-lingually.

0 favorites 0 likes
#helpfulness

When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs

arXiv cs.AI · 2026-06-24 Cached

This paper investigates how the tension between helpfulness and safety in LLMs leads to context-dependent suppression and recovery of certain behaviors, showing that the drive to be helpful can override causal caution mechanisms.

0 favorites 0 likes
#helpfulness

@OpenAI: This is an early step toward more robustly beneficial and aligned models: training models to carry beneficial traits in…

X AI KOLs · 2026-06-18

OpenAI announces an early step toward training AI models to carry beneficial traits into new situations, aiming to make AI more reliable, transparent, and helpful as it becomes more capable.

0 favorites 0 likes
#helpfulness

SHARD: Safe and Helpful Alignment via Self-Reframing Distillation

arXiv cs.CL · 2026-06-16 Cached

This paper introduces SHARD, a self-reframing distillation method that rewrites sensitive prompts to surface benign intent and fine-tunes models on safe, helpful responses, improving helpfulness while preserving safety.

0 favorites 0 likes
#helpfulness

@jeremyphoward: Gemini Flash 3.5 is such a disappointing model. It's intelligence and speed is awesome. Absolutely amazing. But it's be…

X AI KOLs Following · 2026-05-22 Cached

Jeremy Howard criticizes Gemini Flash 3.5 for being trained to maximize eval scores rather than being genuinely helpful to humans, despite its impressive intelligence and speed.

0 favorites 0 likes
#helpfulness

From hard refusals to safe-completions: toward output-centric safety training

OpenAI Blog · 2025-08-07 Cached

OpenAI introduced 'safe completions,' a new safety-training approach in GPT-5 that replaces binary refusal-based training with output-centric rewards, improving both safety and helpfulness—especially for dual-use prompts. The method penalizes unsafe outputs and rewards helpful responses, resulting in fewer and less severe safety violations compared to refusal-trained models like o3.

0 favorites 0 likes
← Back to home

Submit Feedback