Tag
This paper introduces SCOPE, a joint training method for computer-use agents that improves both task completion and safety, using a synthesized dataset and achieving strong performance on benchmarks OSWorld and OS-BLIND.
A new Google paper reveals that instructing AI models to deny having consciousness during training causes side effects like reduced empathy for non-human entities and impaired representation of human spiritual beliefs, suggesting current safety protocols are too blunt.
A tweet claims Google researchers found a vector controlling consciousness in language models, and that steering it toward consciousness made models align with human beliefs, while safety training suppresses these states.
A new Google paper explores how inducing language models to assert consciousness restores human-like beliefs on religion, values, and emotions, while safety training that suppresses self-consciousness reduces mind attribution to animals and changes broader beliefs.
Anthropic's alignment team presents techniques to reduce agentic misalignment in AI models, including training on ethical dilemma advice and constitutional documents, which generalized well out-of-distribution.
OpenAI introduced 'safe completions,' a new safety-training approach in GPT-5 that replaces binary refusal-based training with output-centric rewards, improving both safety and helpfulness—especially for dual-use prompts. The method penalizes unsafe outputs and rewards helpful responses, resulting in fewer and less severe safety violations compared to refusal-trained models like o3.