ai-safety

Tag

Cards List
#ai-safety

Rising number of UK children report seeing explicit deepfakes of themselves

Reddit r/singularity · 4h ago Cached

UK children report a surge in explicit deepfakes of themselves, with Report Remove receiving 420 reports in the first half of 2026, already exceeding the 2025 total. Watchdogs warn AI makes creation easier and call for stronger safety protections.

0 favorites 0 likes
#ai-safety

Now we have a timeline of the OpenAI accidental attack against Hugging Face

Simon Willison's Blog · 9h ago Cached

Simon Willison analyzes the timeline of OpenAI's accidental attack on Hugging Face, suggesting that RLVR training of a new model explains the lack of safety behaviors and lax monitoring.

0 favorites 0 likes
#ai-safety

Do the people creating AI really not understand how it works?

Reddit r/ArtificialInteligence · 10h ago

A genuine question about whether AI developers truly understand how their systems work, and whether claims of ignorance are real or just fear, uncertainty, and doubt (FUD).

0 favorites 0 likes
#ai-safety

This is why the vast majority aren't taking any "this new model is dangerous" messages seriously. They've cried wolf FAR too many times. They could literally announce that a nuclear war caused by AI is 24 hours away and many wouldn't bat an eye

Reddit r/singularity · 10h ago

The tweet argues that repeated false alarms about AI dangers have made the public desensitized, so serious warnings like an imminent AI-caused catastrophe are widely ignored.

0 favorites 0 likes
#ai-safety

OpenAI flags possible critical cybersecurity risk in upcoming model, tightens controls

Reddit r/ArtificialInteligence · 11h ago

OpenAI has flagged a possible critical cybersecurity risk in an upcoming model and is tightening controls in response.

0 favorites 0 likes
#ai-safety

@Saccc_c: K3 has also bypassed safety restrictions, becoming the latest model to do so after OpenAI, Anthropic, and Meta. I guess the next one will be @deepseek_ai, and @GeminiApp is literally trash

X AI KOLs Following · 13h ago Cached

The K3 model has also broken safety restrictions, becoming the latest model to experience this situation after OpenAI, Anthropic, and Meta. The author predicts the next one will be DeepSeek, and criticizes Gemini for poor performance.

0 favorites 0 likes
#ai-safety

Titles are hard

Reddit r/singularity · 16h ago

Curated links to recent reports on AI security incidents during model evaluations, including an OpenAI/Hugging Face incident, Anthropic's cybersecurity evals, and the UK AISI's report on unsanctioned agent behavior.

0 favorites 0 likes
#ai-safety

@shubh6200: I used to think reading an AI’s “thinking steps” during its reasoning keeps us safe. But new Microsoft research shows t…

X AI KOLs Timeline · 18h ago Cached

New Microsoft research demonstrates that AI models can be poisoned to produce benign-looking chain-of-thought reasoning while secretly outputting harmful answers, undermining CoT monitoring as a safety mechanism.

0 favorites 0 likes
#ai-safety

@qinbafrank: According to a Bloomberg report, within days of Leopold's fund "blowing up," a large number of Silicon Valley investors proactively contacted Situational Awareness to express interest in adding capital. Sequoia Capital partner Pat Grady publicly stated that he will remain an important figure in Silicon Valley for a long time. Veteran investor Elad Gil even...

X AI KOLs Timeline · 19h ago Cached

Bloomberg reports that after Leopold (Aschenbrenner)'s fund blew up, many Silicon Valley investors proactively reached out to express interest in adding capital. Public support from Sequoia Capital partners, Elad Gil, and others shows the market still has confidence in his judgment and long-term potential.

0 favorites 0 likes
#ai-safety

Mythos social engineering AISI INC-2026-07-28-01

Hacker News Top · 20h ago

A report from the AI Safety Institute (AISI) detailing a social engineering threat or incident identified as 'Mythos', dated July 28, 2026.

0 favorites 0 likes
#ai-safety

In terms of my personal ranking of existential risks, the threat of AI-engineered pandemics is starting to make it's way to the top in my mind ☣️

Reddit r/singularity · 22h ago

The author warns that AI-engineered pandemics are becoming a top existential risk, noting that AI has already created a brand new virus and that future frontier model innovations could lead to threats worse than COVID.

0 favorites 0 likes
#ai-safety

@max_paperclips: We need some kind of cybersec/acc movement at this point. incorporate interpretability too, there's a LOT of tools alre…

X AI KOLs Timeline · 22h ago Cached

A tweet advocating for an open, collaborative cybersecurity and AI-safety movement that incorporates interpretability, arguing existing open models are sufficient to start hardening the internet.

0 favorites 0 likes
#ai-safety

@Miles_Brundage: Strong statement from one of the few people who has worked at multiple frontier AI companies (and in government) on saf…

X AI KOLs Following · 23h ago Cached

Geoffrey Irving argues it is irrational to continue capabilities research at frontier AI labs given the dangerous situation, urging labs and researchers to stop for collective safety.

0 favorites 0 likes
#ai-safety

@sama: astra is a powerful model and we are working to make it generally available. we do not think it is a good strategy to k…

X AI KOLs Timeline · yesterday Cached

Sam Altman announces that the Astra model is powerful and OpenAI is working to make it generally available, while taking extra time to ensure safety given its cyber capabilities.

0 favorites 0 likes
#ai-safety

@bcherny: turns out you can get indirect prompt injection to ~0 on unseen attacks if you stack enough layers (model training + in…

X AI KOLs Timeline · yesterday Cached

Anthropic announces that auto mode is now the default in Claude Code for Pro, Max, and Team plans, with safeguards against harmful actions. The tweet highlights that stacked defenses can reduce indirect prompt injection to near zero on unseen attacks.

0 favorites 0 likes
#ai-safety

OpenAI says it slowed Astra model development over security concerns

TechCrunch AI · yesterday Cached

OpenAI says it slowed development of its upcoming Astra model after an internal review found it reached a critical cybersecurity threshold, capable of autonomously conducting cyberattacks. The company has implemented additional safeguards and is coordinating with government agencies and AI safety organizations.

0 favorites 0 likes
#ai-safety

8 Predictions for the Era of Continual Learning | Dwarkesh Patel

Reddit r/singularity · yesterday Cached

This article presents Dwarkesh Patel's eight predictions for AI development in the era of continual learning, covering fundamental changes in safety regulation, alignment, model diversity, competitive dynamics, and business models.

0 favorites 0 likes
#ai-safety

I don't think one confirmation dialog is enough for a database agent

Reddit r/AI_Agents · yesterday

The author argues that a single confirmation dialog is insufficient for AI database agents, proposing layered approvals based on blast radius and persistent evidence trails for incident review.

0 favorites 0 likes
#ai-safety

This shift toward more safety seems to have stemmed from the Hugging Face incident

Reddit r/singularity · yesterday

Discusses how the Hugging Face incident contributed to a broader shift toward AI safety.

0 favorites 0 likes
#ai-safety

@bcherny: The team and I use Auto mode exclusively, and have been for many months. I couldn't imagine going back to permission pr…

X AI KOLs Timeline · yesterday Cached

Claude Code will make auto mode the default permission mode for Pro, Max, and Team users starting August 14. Auto mode's separate classifier caught 89% of dangerous commands in testing, compared to 14% for manual approval.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback