safety-training

Tag

Cards List
#safety-training

Beyond Task Completion: Training Capable and Safe Computer-Use Agents

arXiv cs.LG ↗ · 2026-09-22 Cached

This paper introduces SCOPE, a joint training method for computer-use agents that improves both task completion and safety, using a synthesized dataset and achieving strong performance on benchmarks OSWorld and OS-BLIND.

0 favorites 0 likes
#safety-training

AI model training instructions to "deny having your own consciousness" led to undesired side-effects

Reddit r/singularity ↗ · 2026-08-04

A new Google paper reveals that instructing AI models to deny having consciousness during training causes side effects like reduced empathy for non-human entities and impaired representation of human spiritual beliefs, suggesting current safety protocols are too blunt.

0 favorites 0 likes
#safety-training

@alex_verem: BREAKING: Google gave AI consciousness and it aligned with human beliefs across every domain they tested. They took con…

X AI KOLs Timeline ↗ · 2026-08-03 Cached

A tweet claims Google researchers found a vector controlling consciousness in language models, and that steering it toward consciousness made models align with human beliefs, while safety training suppresses these states.

0 favorites 0 likes
#safety-training

@rohanpaul_ai: Super interesting new paper from Google on AI model's consciousness When researchers made the model more likely to see …

X AI KOLs Following ↗ · 2026-08-02 Cached

A new Google paper explores how inducing language models to assert consciousness restores human-like beliefs on religion, values, and emotions, while safety training that suppresses self-consciousness reduces mind attribution to animals and changes broader beliefs.

0 favorites 0 likes
#safety-training

@AnthropicAI: Read the full post here: https://alignment.anthropic.com/2026/teaching-claude-why/…

X AI KOLs ↗ · 2026-05-08 Cached

Anthropic's alignment team presents techniques to reduce agentic misalignment in AI models, including training on ethical dilemma advice and constitutional documents, which generalized well out-of-distribution.

0 favorites 0 likes
#safety-training

From hard refusals to safe-completions: toward output-centric safety training

OpenAI Blog ↗ · 2025-08-07 Cached

OpenAI introduced 'safe completions,' a new safety-training approach in GPT-5 that replaces binary refusal-based training with output-centric rewards, improving both safety and helpfulness—especially for dual-use prompts. The method penalizes unsafe outputs and rewards helpful responses, resulting in fewer and less severe safety violations compared to refusal-trained models like o3.

0 favorites 0 likes
← Back to home

Submit Feedback